A method for classifying input images representing particles in a sample
The method efficiently classifies bacterial images using t-SNE and pre-trained neural networks to assess antibiotic susceptibility, overcoming the limitations of traditional antibiograms and deep learning's resource intensity.
Patent Information
- Application Number
- JP2023524073
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-20
- Filing Date
- 2021-10-19
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2041-10-19
AI Technical Summary
Existing methods for determining antibiotic susceptibility of bacteria, such as antibiograms, are lengthy, complex, and resource-intensive, often requiring chemical markers that are cytotoxic and limit repeated observations, and deep learning-based solutions are computationally intensive.
A method using data processing means to extract feature maps from input images of biological particles, reduce variables with the t-SNE algorithm, and perform unsupervised classification, utilizing pre-trained convolutional neural networks and k-nearest neighbor algorithms to classify images of bacteria effectively and efficiently.
Enables rapid, non-destructive classification of bacterial images with reduced computational resources, allowing early assessment of antibiotic efficacy without the need for extensive training data and manual annotation.
Smart Images

Figure 0007798256000003 
Figure 0007798256000004 
Figure 0007798256000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of optical acquisition of biological particles. The biological particles may be, for example, microorganisms such as bacteria, fungi or yeasts. They may also be cells, multicellular organisms or any other type of particle matter, such as pollutants or dust.
[0002] The present invention is particularly advantageously applicable to the analysis of the state of biological particles, for example with the aim of determining the metabolic state of bacteria after the application of antibiotics. By means of the present invention, for example, an antibiogram can be carried out for bacteria. [Background technology]
[0003] Antibiograms are laboratory techniques aimed at testing the phenotype of bacterial strains against one or more antibiotics. Traditionally, antibiograms are performed by culturing a sample containing bacteria and antibiotics.
[0004] European Patent Application No. 2 603 601 describes a method for performing an antibiogram, which involves visualizing the state of bacteria after an incubation period in the presence of an antibiotic. To visualize the bacteria, the bacteria are labeled with a fluorescent marker to reveal their structure. The fluorescence of the marker can then be measured to determine whether the antibiotic has effectively acted on the bacteria.
[0005] The traditional process for determining an effective antibiotic against a given bacterial strain involves collecting a sample containing that strain (e.g., from a patient, animal, food batch, etc.) and then sending the sample to an analytical center. Upon receiving the sample, the analytical center first cultures the bacterial strain to obtain at least one colony of the strain, which takes 24 to 72 hours. From this colony, the center then prepares several samples containing different antibiotics and / or different concentrations of antibiotic, and incubates the samples again. After a new incubation period, also requiring 24 to 72 hours, each sample is manually analyzed to determine whether the antibiotic was effective. The results are then returned to the practitioner so that the practitioner can apply the most effective antibiotic and / or antibiotic concentration.
[0006] However, the labeling process is particularly lengthy and complex to implement, and these chemical markers have cytotoxic effects on bacteria. Therefore, this visualization method does not allow for repeated observation of bacteria during culture. Consequently, to ensure reliable measurements, bacteria must be cultured for a sufficiently long time, such as 24 to 72 hours. Other methods for visualizing biological particles use microscopes, which allow for nondestructive measurement of samples.
[0007] Digital holographic microscopy (DHM) is an imaging technique that can overcome the depth-of-field limitations of conventional optical microscopy. Schematically, it involves recording a hologram formed by the interference between light waves diffracted by the observed object and a spatially coherent reference wave. This technique is described in a review article by Myung K. Kim entitled "Principles and techniques of digital holography microscopy," published in SPIE Reviews Vol. 1, No. 1, January 2010.
[0008] Recently, it has been proposed to use digital holographic microscopy to identify microorganisms in an automated manner. Thus, International Application WO 2017 / 207184 describes a method for particle acquisition that combines a simple defocused acquisition with digital focus reconstruction to allow the observation of biological particles while limiting acquisition times.
[0009] Typically, this solution allows for the detection of structural modifications to bacteria in the presence of antibiotics after just 10 minutes of incubation, and for sensitivity (detection of the presence or absence of division or patterns indicative of division) after 2 hours, unlike the conventional processes described above, which can take several days. Specifically, because the measurement is non-destructive, it is possible to perform the analysis very early in the culture process without risking sample destruction and therefore extending the analysis time.
[0010] To visualize particle behavior, e.g., its migration speed or its cell division process, it is even possible to track particles over several successive images (as the particles are intact after the initial analysis) to form a film representing the particle's progression over time.
[0011] It will therefore be appreciated that this visualization method gives excellent results, as these images or this film would be difficult to interpret on their own if it were desired to reach a conclusion regarding, for example, the susceptibility of bacteria to antibiotics present in the sample.
[0012] Various techniques have been proposed, ranging from simply counting bacteria over time to so-called morphological analyses that aim to detect specific "configurations" via image analysis. For example, when bacteria prepare to divide, two poles appear in the distribution well before the division itself, resulting in the division of the distribution into two distinct segments.
[0013] Combining these two techniques to assess antibiotic efficacy has been proposed in the paper [Choi et al., 2014]. However, as the authors emphasize, their approach requires a very fine calibration of a certain number of thresholds, which strongly depends on the nature of the morphological changes induced by the antibiotic.
[0014] More recently, a deep learning-based approach was described in the paper [Yu et al., 2018]. The authors propose using a convolutional neural network (CNN) to extract morphological features and features related to bacterial movement. However, this solution has proven to be very intensive in terms of computing resources, requiring a large database of training images to train the CNN.
[0015] The technical problem aimed at by the present invention is therefore to be able to provide a more effective and less resource-intensive solution for classifying images of biological particles. Summary of the Invention
[0016] According to a first aspect, the present invention relates to a method for classifying at least one input image representative of a target particle in a sample, characterized in that it comprises the implementation, by data processing means of a client, of the following steps: (b) extracting a feature map of said target particle from the input image; (c) A step of reducing the number of variables in the extracted feature map by the t-SNE algorithm; (d) performing unsupervised classification of the input image according to the feature map having a reduced number of variables;
[0017] According to advantageous but non-limiting features:
[0018] The particles are represented in a uniform manner in the input image and in each of the base images, in particular, they are centered and aligned with a given direction.
[0019] The method includes the step (a) of subtracting the input image from a global image of the sample to represent the target particles in the uniform manner.
[0020] Step (a) includes segmenting the overall image to detect the target particles in the sample, and then cropping the input image to the detected target particles.
[0021] Step (a) includes obtaining the overall image from an intensity image of the sample, the image being obtained by an observation device.
[0022] The feature map is a vector of numerical coefficients each associated with one base image of a set of base images each representing a reference particle, and step (a) includes determining the numerical coefficients such that a linear combination of the base images weighted by the coefficients approximates a representation of the target particle in the input image.
[0023] The feature map of the target particle is extracted in step (b) by a convolutional neural network pre-trained on a public image database.
[0024] Step (c) includes defining an embedding space for each feature map of the training database of already classified feature maps of particles in the sample and for the extracted feature map by the t-SNE algorithm, the feature map having a reduced number of variables that is the result of embedding the extracted feature map into the embedding space.
[0025] Step (c) involves implementing a k-nearest neighbor algorithm in the embedding space.
[0026] The method is for classifying a sequence of input images representing said target particles in a sample over time, and step (b) comprises concatenating the extracted feature maps of each input image in said sequence.
[0027] According to a second aspect, there is provided a system for classifying at least one input image representing a target particle in a sample, comprising at least one client comprising data processing means, characterized in that said data processing means are configured to perform: - extracting a feature map of said target particle through analysis of at least one input image; - Reducing the number of variables in the feature map by the t-SNE algorithm; - performing an unsupervised classification of said input image according to said feature map having a reduced number of variables.
[0028] According to an advantageous but non-limiting feature, the system further comprises a device for observing said target particles in the sample.
[0029] According to third and fourth aspects, there are provided a computer program product comprising code instructions for performing the method according to the first aspect for classifying at least one input image representing target particles in a sample, and a storage medium readable by a computing device, the computer program product comprising code instructions for performing the method according to the first aspect for classifying at least one input image representing target particles in a sample. [Brief explanation of the drawings]
[0030] Other characteristics and advantages of the invention will become apparent from the following description of preferred embodiments, which description will be given with reference to the accompanying drawings. [Figure 1] 1 is a schematic diagram of an architecture for implementing the method according to the invention; [Figure 2a]1 shows an example of an apparatus for observing particles in a sample, which is used in a preferred embodiment of the method according to the present invention. [Figure 3a] 3 illustrates the acquisition of an input image in an embodiment of the method according to the invention; [Figure 3b] 2 illustrates the acquisition of an input image in a preferred embodiment of the method according to the invention; [Figure 4] 1 illustrates the steps of a preferred embodiment of the method according to the present invention. [Figure 5a] 2 shows an example of a dictionary of base images used in a preferred embodiment of the method according to the invention; [Figure 5b] 1 shows an example of the extraction of feature vectors and matrices in a preferred embodiment of the method according to the invention. [Figure 6] 1 shows an example of a convolutional neural network architecture used in a preferred embodiment of the method according to the invention. [Figure 7] 1 represents an example of a t-SNE embedding used in a preferred embodiment of the method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] architecture The present invention relates to a method for classifying at least one input image representing particles 11a-11f present in a sample 12, called target particles. It should be noted that the method may be performed in parallel on all or some of the particles 11a-11f present in the sample 12, each of which in turn is considered a target particle.
[0032] As will be appreciated, the method may include one or more machine learning components, in particular one or more classifiers including convolutional neural networks (CNNs).
[0033] The input or training data are of image type and represent target particles 11a-11f in the sample 12 (in other words, they are images of the sample in which the target particles are visible). As will be appreciated, a sequence of images of the same target particle 11a-11f (or, if desired, multiple sequences of images of particles 11a-11f of the sample 12, if multiple particles are considered) can be provided as input.
[0034] The sample 12 consists of a liquid such as water, a buffer, a culture medium, or a reactive medium (which may or may not contain antibiotics) in which the particles 11a-11f to be observed are located.
[0035] As a variant, the sample 12 may take the form of a solid medium, preferably translucent, such as agar, in which the particles 11a-11f are located. The sample 12 may also be a gaseous medium. The particles 11a-11f may be located inside the medium or on the surface of the sample 12.
[0036] Particles 11a-11f may be microorganisms such as bacteria, fungi, or yeast. They may also be cells, multicellular organisms, or any other type of particle matter, such as pollutants or dust. In the remaining description, we consider a preferred example in which the particles are bacteria (and, as will be understood, sample 12 incorporates antibiotics). The size of observed particles 11a-11f varies between 500 nm and several hundred μm, or even several millimeters.
[0037] The "classification" of an input image (or a sequence of input images) consists in determining at least one class from a set of possible classes describing the image. For example, in the case of bacteria-type particles, a binary classification may be employed, i.e., two possible classes may be employed, indicating "division" or "no division," evidencing the presence or absence of resistance to antibiotics, respectively. Although the example of a binary classification of the effects of antibiotics on the target particles 11a-11f above is primarily described, the present invention is not limited to any one particular type of classification.
[0038] The method is implemented in an architecture as shown in Figure 1 by a Server 1 and a Client 2. The Server 1 is the device to be trained (which implements the training method) and the Client 2 is the user device (which implements the classification method), e.g., a doctor's or hospital terminal.
[0039] Preferably, the server 1 is a remote device and the client 2 is a mass-market device, in particular a desktop computer, laptop computer, etc., although it is entirely possible to combine the two devices 1, 2. The client device 2 is advantageously connected to the observation apparatus 10 so as to be able to directly acquire the input image (or, as will be seen below, "raw" acquired data such as a full image of the sample 12, or even an electromagnetic matrix), typically for immediate processing. Alternatively, the input image is loaded onto the client device 2.
[0040] In any case, each device 1, 2 is typically a remote computer device connected to a local network or a wide area network such as the Internet for the purpose of exchanging data. Each comprises data processing means 3, 20 of a processor type and data storage means 4, 21 such as computer memory, e.g., flash memory or a hard disk. The client 2 typically comprises a user interface 22, such as a screen, which allows interaction.
[0041] The server 1 advantageously stores a training database, i.e., a set of images of particles 11a-11f in various conditions (see below) and / or a set of already classified feature maps (e.g., associated with labels "with division" or "without division" indicating susceptibility or resistance to antibiotics). It should be noted that the training data is possibly associated with labels defining the test conditions, e.g., for bacterial cultures, indicating "strain," "antibiotic conditions," "time," etc.
[0042] acquisition As described above, the method can directly take as input any image of the target particles 11a-11f obtained by any method. However, the method preferably starts with step (a) of obtaining the input image from data provided by the observation device 10.
[0043] In a known manner, the skilled person can use, in particular, the DHM technique (DHM stands for digital holographic microscopy) as described in International Application WO 2017 / 207184. In particular, an intensity image of the sample 12 can be acquired in which the target particles are not in focus (the image is considered to be "out of focus") but which can be processed by data processing means (which are, for example, integrated into either the device 10 or the device 20 of the client 2, see below); such an image is called a hologram. It will be understood that a hologram "represents" all particles 11a-11f in the sample in a specific way.
[0044] FIG. 2 shows an example of an apparatus 10 for observing particles 11a-11f present in a sample 12. The sample 12 is placed between a spatially and temporally coherent light source 15 (e.g., a laser) or a quasi-coherent light source 15 (e.g., a light-emitting diode, a laser diode) and a digital sensor 16 sensitive in the spectral range of the light source. Preferably, the light source 15 has a narrow spectral width, e.g., narrower than 200 nm, narrower than 100 nm, or even narrower than 25 nm. In the following, reference is made to the central emission wavelength of the light source, e.g., in the visible range. The light source 15 emits a coherent signal Sn toward the first surface 13 of the sample, and the signal is transmitted by a waveguide, e.g., an optical fiber.
[0045] The sample 12 (typically described as a culture medium) is contained within an analysis chamber bounded vertically by a lower slide and an upper slide, e.g., a conventional microscope slide. The analysis chamber is bounded laterally by an adhesive or any other sealing material. The lower and upper slides are transparent to the wavelength of the light source 15, and the sample and chamber, for example, allow more than 50% of the light source's wavelength to pass through at normal incidence on the lower slide.
[0046] Preferably, particles 11a-11f are located in sample 12 adjacent to the upper slide. For this purpose, the bottom surface of the upper slide contains a ligand that allows particle attachment, such as a polycation (e.g., poly-L-lysine) in the context of microorganisms. This allows particles to be contained at a thickness equal to or close to the depth of field of the optical system, i.e., less than 1 mm (e.g., a tube lens), preferably less than 100 μm (e.g., a microscope objective lens). Nevertheless, particles 11a-11f can move within sample 12.
[0047] Preferably, the apparatus comprises an optical system 23, consisting of, for example, a microscope objective and a tube lens, located in air at a fixed distance from the sample. The optical system 23 optionally comprises a filter, which may be located in front of the objective or between the objective and the tube lens. The optical system 23 is characterized by its optical axis, its object plane (also called the focal plane), which is located at a certain distance from the objective, and its image plane, which is conjugate to the object plane by the optical system. In other words, an object located in the object plane corresponds to a sharp image of this object in the image plane, also called the focal plane. The optical properties of the system 23 are fixed (e.g., fixed focal length optics). The object plane and the image plane are perpendicular to the optical axis.
[0048] The image sensor 16 is located in or close to the focal plane facing the second surface 14 of the sample. The sensor, e.g., a CCD or CMOS sensor, comprises a periodic two-dimensional array of elementary sensing sites and associated electronics for adjusting the exposure time and zeroing the sites in a manner known per se. The signal output by an elementary site depends on the amount of radiation in that spectral region incident on that site during the exposure time. This signal is then converted, for example by the associated electronics, into image points, or "pixels," of a digital image. The sensor thus generates a digital image in the form of a matrix with C columns and L rows. Each pixel of this matrix, with coordinates (c, l), corresponds, in a manner known per se, to a position in the focal plane of the optical system 23 with Cartesian coordinates (x(c, l), y(c, l)), e.g., the position of the center of a rectangular elementary sensing site.
[0049] The pitch and fill factor of the periodic array are selected to satisfy the Nyquist criterion for the size of the particles being observed, thereby defining at least two pixels per particle. Thus, the image sensor 16 acquires a transmission image of the sample in the spectral range of the light source.
[0050] The image acquired by the image sensor 16 contains holographic information insofar as it results from the interference between the waves diffracted by the particles 11a-11f and a reference wave that passed through the sample without interacting with it. As mentioned above, in the context of a CMOS or CCD sensor, the acquired digital image is an intensity image, and therefore it is clear that here phase information is encoded in this intensity image.
[0051] Alternatively, the coherent signal Sn generated by the source 15 can be split into two components, for example by a semitransparent plate, where the first component acts as a reference wave and the second component is diffracted by the sample 12, and the image in the image plane of the optical system 23 results from the interference between the diffracted wave and the reference wave.
[0052] Referring to FIG. 3a, in step (a), it is possible to reconstruct at least one overall image of the sample 12 from a hologram and then extract the input image from the overall image of the sample.
[0053] Specifically, it will be understood that the target particles 11a-11f must be represented in a uniform manner in the input image, in particular, centered and aligned in a predetermined direction (e.g., horizontally). The input image must also have a standardized size (it is also desirable that only the target particles 11a-11f are visible in the input image). Therefore, the input image is called a "thumbnail," and its size may be defined to be, for example, 250 x 250 pixels. In the case of a sequence of input images, for example, one image is taken every minute during a 120-minute time interval, and the sequence therefore forms a 3D "stack" of size 250 x 250 x 120.
[0054] The overall image is reconstructed by the data processing means of the device 10 or by the data processing means 20 of the client 2 as explained.
[0055] Typically, a series of complex matrices, called "electromagnetic matrices", are constructed (for each given acquisition time), which model the wavefront of a light wave propagating along the optical axis for multiple deviations relative to the focal plane of the optical system 23, in particular deviations located within the sample, based on an intensity image (hologram) of the sample 12.
[0056] These matrices can be projected into real space (e.g., via a Hermitian norm) to form a stack of whole images at various focal lengths.
[0057] From there, it is possible to determine the average focal length (and select the corresponding overall image or recalculate it from the hologram), or even to determine the optimal focal length for the target particle (and again select the corresponding overall image or recalculate it from the hologram).
[0058] In either case, with reference to Figure 3b, step (a) advantageously comprises segmenting and then cropping the one or more overall images to detect the target particles in the sample, in particular the input image may be extracted from the overall image of the sample so as to represent the target particles in the uniform manner.
[0059] Generally, segmentation allows for the detection of all particles of interest while removing artifacts such as filaments or microcolonies to improve one or more overall images. One of the detected particles is then selected as the target particle and a corresponding thumbnail is extracted. As described, this can be done for all detected particles.
[0060] Segmentation can be performed in any known manner. In the example of Figure 3b, fine segmentation is first performed to remove artifacts, and then coarse segmentation is performed to detect particles 11a-11f. Any segmentation technique known to those skilled in the art can be used.
[0061] If it is desired to acquire a sequence of input images of a target particle 11a-11f, tracking techniques may be used to track any movement of the particle from one global image to the next.
[0062] Note that, as seen on the right side of Figure 3a, all input images acquired over time for a given sample (for multiple particles, or even all particles, of sample 12) can be pooled to form a corpus describing sample 12 (in other words, a corpus describing the experiment), which is then copied, among other things, to storage means 21 of client 2. This is at the "field" level, as opposed to the "particle" level. For example, if particles 11a-11f are bacteria and sample 12 contains (or does not contain) an antibiotic, this description corpus would contain all information about the growth, morphology, internal structure, and optical properties of these bacteria across the entire field of acquisition. As will be appreciated, this description corpus can be transmitted to server 1 for integration into the training database described above.
[0063] Feature Extraction 4, the method is particularly noteworthy in that it does not attempt to classify the input image directly, but rather the step (b) of extracting a feature map from the input image is carried out separately from the step (d) of classifying the input image according to said feature map, with the step (c) of reducing the number of variables of the feature map by means of the t-SNE algorithm occurring between these two steps. More precisely, in step (c), an embedding of the feature map, called a "t-SNE embedding", is constructed, which has fewer variables than the number of variables of the extracted feature map, advantageously only two or three variables.
[0064] In the remainder of this document, a distinction is made between the number of "dimensions" of feature maps in the geometric sense, i.e., the number of independent directions they extend (e.g., vectors are objects of dimension 1, while our feature maps are at least dimension 2, advantageously dimension 3, and sometimes dimension 4), and the number of "variables" of these feature maps, i.e., their size in each dimension, i.e., the number of independent degrees of freedom (which actually corresponds to the concept of dimension in vector spaces; more precisely, a set of feature maps with a given number of variables forms a vector space of dimension equal to the number of variables, and similarly for sets of t-SNE embeddings). Thus, step (c) is sometimes called a "dimensionality reduction" step, insofar as a first high-dimensional vector space (feature map space) is mapped to a second, lower-dimensional vector space (2D or 3D space), although in fact it is the number of variables that is reduced.
[0065] Therefore, two examples are described below in which the feature maps extracted at the end of step (b) are respectively a 2D object (i.e., an object-matrix of dimension 2) with a size of 60x25 and thus 1500 variables, and a 3D object (i.e., an object of dimension 3) with a size of 7x7x512 and thus 25088 variables. In these two examples, the number of variables is reduced to 2 or 3.
[0066] As will be appreciated, each step may involve an independent learning mechanism that may (but is not necessarily) be automatic, and therefore the above training database of server 1 may include particle images and feature maps that are not necessarily already classified.
[0067] Therefore, the main step (b) is the extraction, by the data processing means 20 of the client 2, of a feature map of said target particles, ie of "coding" the target particles.
[0068] Those skilled in the art will now appreciate that the t-SNE algorithm of step (c) can ingeniously obtain a "simplified" version of the feature map, which is very easy to handle, and so any technique for extracting feature maps may be used, including techniques capable of generating a large number of feature maps with high dimensionality (three, or even four).
[0069] We then describe several techniques that, in particular, allow for obtaining high-semantic-level feature maps without requiring large amounts of computational power or annotated databases.
[0070] Thus, when a sequence of input images is provided, step (b) advantageously involves extracting one feature map per input image, which may be combined into a single feature map called the "profile" of the target particle. More precisely, since the maps all have the same size and form a sequence of maps, it is sufficient to concatenate them in the order of the input images to obtain a "deep" feature map. In such cases, reducing the number of variables per t-SNE is even more advantageous.
[0071] Alternatively or additionally, feature maps corresponding to multiple input images associated with multiple particles 11a-11f of the sample 12 may be summed.
[0072] According to a first embodiment of step (b), the feature map is simply a feature vector, wherein said features are numerical coefficients each associated with one base image of a set of base images each representing a reference particle, such that a linear combination of said base images weighted by said coefficients approximates a representation of said particle in the input image.
[0073] This is called "sparse coding". These base images are called "atoms", and the set of atoms is called a "dictionary". The idea behind sparse coding is to represent any input image as a linear combination of these atoms, by analogy with the words in the dictionary. More precisely, for a dictionary D of size p and a feature vector α, also of size p, the best approximation Dα of the input image x is found. In other words, if the optimal vector (the sparse code of the input image x) is α*, then step (b) consists in solving a minimization problem of a function with λ as a regularization parameter (this allows making a compromise between the quality of the approximation and the sparsity of the vector, i.e., containing as few atoms as possible). For example, the constrained minimization problem can be written as follows:
number
[0074] It can also be expressed as a variational formulation problem:
number
[0075] The coefficients advantageously have values in the interval [0,1] (which is simpler than for R), and it will be appreciated that, in general, due to the "sparse" nature of the coding, the majority of the coefficients will have a value of 0. The atoms associated with the non-zero coefficients are called activation atoms.
[0076] Naturally, the base image is a thumbnail comparable to the input image, i.e. the reference particles are represented therein in the same uniform manner as in the input image, in particular centered and aligned in the above-mentioned predetermined direction, and the base image advantageously has the same size as the input image (e.g. 250x250).
[0077] Thus, Figure 5a shows an example dictionary of 36 basis images (for the antibiotic cefpodoxime and the bacterium Escherichia coli).
[0078] The reference images (atoms) may be predefined, but preferably the method comprises a step (b0) of learning from a training database, in which the reference images (i.e. dictionary images) are learned in particular by the data processing means 3 of the server 1, so that the method does not require human intervention at any point.
[0079] This learning method, called "dictionary learning" because it involves learning a dictionary, is unsupervised insofar as it requires that images in the training database be annotated, and is therefore extremely simple to implement. In particular, it will be appreciated that manually annotating thousands of images would be very time-consuming and prohibitively expensive.
[0080] The idea is simply to provide thumbnails representing particles 11a-11f in various conditions in a training database, and based on that, find an atom that allows any thumbnail to be represented as easily as possible.
[0081] If a sequence of input images is provided, step (b) advantageously comprises extracting one feature vector per input image, as explained, whose feature maps can be combined into a feature matrix called the "profile" of the target particle. More precisely, since the vectors all have the same size (number of atoms) and form a sequence of vectors, it is sufficient to juxtapose them in the order of the input images to obtain a sparse two-dimensional code (coding the spatiotemporal information, and therefore the two dimensions).
[0082] Figure 5b shows another example of extracting feature vectors, this time using a dictionary of 25 atoms. Shown is the entire image obtained at a given time T1, along with various extracted input images (corresponding to detected particles). Thus, the image representing the second target particle can be approximated as 0.33 × atom13 + 0.21 × atom2 + 0.16 × atom9 (i.e., the vector (0;0.21;0;0;0;0;0;0;0;0.16;0;0;0;0;0.33;0;0;0;0;0;0;0;0;0;0;0;0).
[0083] The summed vector, called the "cumulative histogram," is shown in the center. Advantageously, the coefficients are normalized so that their sum equals 1. The summed matrix (summed over 60 minutes), called the "activation profile," is shown on the right and can be seen to have a size of 60x25.
[0084] It will be appreciated that this activation profile is a high-level feature map that represents (over time) the sample 12. According to a second embodiment of step (b), a convolutional neural network (CNN) is used to extract the feature map. It will be recalled that CNNs are particularly suited to vision-related tasks. In general, CNNs can directly classify input images (i.e., perform steps (b) and (d) simultaneously).
[0085] Here, by separating step (b) and step (d), the use of CNN can be limited to feature extraction, and in this step (b), a CNN that has been pre-trained on a public image database, i.e., a CNN that has already been independently trained, can be used alone. This is called "transfer learning."
[0086] In other words, the CNN does not need to be trained or retrained on a training database of images of particles 11a-11f, and therefore may not be annotated. In particular, it will be appreciated that manually annotating thousands of images would be very time consuming and very expensive.
[0087] Specifically, to perform the task of feature extraction, it is sufficient for the CNN to distinguish, i.e., identify, differences between images, including public image databases that have no relation to the current input image. Advantageously, the CNN is an image classification network insofar as such a network is known to operate on feature maps that are particularly discriminative with respect to image classes, and is therefore particularly suited in the present context of particles 11a-11f to be classified, even if this is not the task for which the CNN was originally trained. It will be appreciated that even image detection, recognition, or segmentation networks are a particular case of classification networks, since they actually perform another task in addition to the task of classification (of the entire image or of an object within the image) (such as determining the coordinates of the bounding box of the classified object in the case of a detection network, or generating a segmentation mask in the case of a segmentation network).
[0088] Regarding public training image databases, for example, the well-known public database ImageNet can potentially be used, which contains over 1.5 million annotated images and can be used to achieve supervised learning of almost any image processing CNN (for tasks such as classification, recognition, etc.).
[0089] Therefore, it is advantageous to use "off-the-shelf" CNNs that do not even need to be trained. Various classification CNNs are known that are pre-trained on the ImageNet database (i.e., can be obtained with their parameters initialized to correct values as a result of training on ImageNet), such as the VGG model (VGG stands for Visual Geometry Group), e.g., the VGG-16 model, AlexNet, Inception, or even ResNet. Figure 6 represents the VGG-16 architecture (with 16 layers).
[0090] Generally, a CNN consists of two parts: - The first feature extraction sub-network, which usually includes a series of blocks consisting of convolutional layers and activation layers (using, for example, a ReLU function) to increase the depth of the feature map, terminated by a pooling layer that allows reducing the size of the feature map (reducing the input dimensionality—typically by a factor of two). Thus, in the example of Figure 6, VGG-16 has 16 layers divided into five blocks, as described. The first block receives as input an input image (224x224 spatial size, with three channels corresponding to the RGB characteristics of the image), and includes two convolutional + ReLu sequences (one convolutional layer and one ReLu activation layer) that increase the depth to 64, followed by a max-pooling layer (global average pooling can also be used). The output is a feature map of size 112x112x64 (the first two dimensions are spatial dimensions, the third is depth, and therefore each spatial dimension is divided by 2). The second block has the same architecture as the first block, generating a feature map of size 112x112x128 (double depth) at the output of the final convolution + ReLU sequence, and generating a feature map of size 56x56x128 as the output of the max pooling layer. The third block now has three convolution + ReLU sequences, generating a feature map of size 56x56x256 (double depth) from the final convolution + ReLU sequence, and generating a feature map of size 28x28x256 as the output of the max pooling layer. The fourth and fifth blocks have the same architecture as the third block, and successively generate feature maps of size 14x14x512 and 7x7x512 (depth no longer increasing) as the output. This feature map is the "final" map. It should be understood that there are no limitations on map size at any level, and the above sizes are merely examples. - A feature processing second sub-network, in particular a classifier if the CNN is a classification network. This sub-network receives as input the final feature map generated by the first sub-network and returns the expected result, e.g., the class of the input image if the CNN performs classification. This second sub-network typically includes one or more fully connected (FC) layers and a final activation layer employing, for example, a softmax function (in the case of VGG-16). Both sub-networks are generally trained simultaneously in a supervised manner.
[0091] Therefore, in this second embodiment, step (b) is preferably performed by the feature extraction sub-network of said pre-trained convolutional neural network, i.e., the first part as highlighted in Figure 6 for the example of VGG-16.
[0092] More precisely, the pre-trained CNNs (such as VGG-16) are not intended to deliver any feature maps; they are simply for internal use. By "truncating" the pre-trained CNNs, i.e., using only the first sub-network layer, the final feature maps containing the "deepest" information are obtained as output.
[0093] It will be appreciated that it is entirely possible to use a feature extraction sub-network that terminates before the layer where the final feature map is generated, for example, using only blocks 1-3 instead of blocks 1-5. This provides more comprehensive but less in-depth information.
[0094] It should be noted that when a sequence of input images is fed, instead of extracting one feature map per input image, it is possible to combine the maps into a single feature map (by concatenating the maps in the order of the input images, so as to obtain a "deep" feature map). Therefore, it is possible to directly use so-called 3D CNNs, which can be fed with the entire sequence of input images, and in this case there is no need to work image by image.
[0095] To do this, step (b) involves pre-concatenating the input images of the sequence into a three-dimensional or 3D stack, and then directly extracting feature maps of the target particles 11a-11f from the 3D stack by a 3D CNN.
[0096] The 3D stack is processed by the 3D CNN as a single, one-channel, 3D object (e.g., if the input image is 250x250 in size and one image is acquired every minute for 120 minutes, it is 250x250x120 in size - the first two dimensions are traditionally spatial dimensions (i.e., the size of the input image) and the third is the "temporal" dimension (time of acquisition)), rather than as a multi-channel 2D object (e.g., as used with RGB images), so the output feature map is 4-dimensional.
[0097] The present 3D CNN uses at least one 3D convolutional layer that models the spatiotemporal dependencies of various input images.
[0098] A 3D convolutional layer is a convolutional layer that can apply a 4D filter and therefore operate on multiple channels of an already 3D stack, i.e., a 4D feature map. In other words, a 3D convolutional layer applies a 4D filter to a 4D input feature map to produce a 4D output feature map. The fourth and final dimension, like any feature map, is the semantic depth.
[0099] These layers differ from traditional convolutional layers, which can only operate on 3D feature maps that represent multiple channels of a 2D object (image).
[0100] The concept of 3D convolution may seem counterintuitive, but it is simply a generalization of the concept of a convolutional layer, which specifies that multiple "filters" of depth equal to the number of input channels (i.e., the depth of the input feature map) are applied by scanning them across all dimensions of the input (in 2D of the image), with the number of filters defining the output depth.
[0101] Thus, our 3D convolution applies 4D filters of depth equal to the number of channels in the 3D input stack and scans these filters across the entire volume of the 3D stack, thus across not only the two spatial dimensions but also the time dimension, i.e., across three dimensions (hence the name 3D convolution). Thus, in effect, we get one 3D stack per filter, i.e., a 4D feature map. In a traditional convolution layer, using a large number of filters certainly increases the semantic depth (number of channels) of the output, but the output will always be a 3D feature map.
[0102] Reducing the number of variables The feature maps obtained in step (b) (especially when an image sequence is input) may have a very large number of variables (thousands or tens of thousands), making direct classification complicated.
[0103] Thus, in step (c), the use of the t-SNE algorithm has two important advantages: - By using a low-dimensional space (sometimes called the embedding space or visualization space), preferably a two-dimensional space, the data can be visualized and manipulated much more easily and intuitively than in the original space of the feature maps. In particular, step (c) allows for unsupervised classification of the input images, i.e. without the need to train a classifier.
[0104] The trick is to construct a t-SNE embedding of the entire training database, i.e., it is possible to define the embedding space depending on the training database.
[0105] In other words, thanks to the t-SNE algorithm, the feature maps of the input image and each feature map of the training database can be represented by a two-variable or three-variable embedding in the same embedding space, so that two feature maps that are close (far apart) in the original space are close (far apart) in the embedding space, respectively.
[0106] Specifically, the t-SNE algorithm (t-SNE stands for t-distributed stochastic neighbor embedding) is a nonlinear method for achieving dimensionality reduction for data visualization, allowing a set of points in a high-dimensional space to be represented in a two- or three-dimensional space, allowing the data to be visualized using a scatter plot. The t-SNE algorithm attempts to find an optimal configuration (the t-SNE embedding mentioned above) in terms of point proximity according to information-theoretic criteria.
[0107] The t-SNE algorithm is based on a probabilistic interpretation of proximity. For pairs of points in the original space, a probability distribution is defined such that points that are close to each other have a high probability of being selected, and points that are far apart have a low probability of being selected. A similar probability distribution is defined for the embedding space. The t-SNE algorithm consists in matching two probability densities by minimizing the Kullback-Leibler divergence between the two distributions with respect to the locations of the points on the map.
[0108] The t-SNE algorithm can be implemented both at the particle level (target particles 11a-11f for individual particles for which maps are available in the training database) and at the field level (for the entire sample 12 - in the case of multiple input images representing multiple particles 11a-11f), especially in the case of single images rather than stacks.
[0109] Note that t-SNE embedding can be achieved efficiently, particularly through implementation in Python, and thus can be performed in real time. To speed up computations and reduce memory footprint, it is also possible to first go through a step of linear dimensionality reduction (e.g., PCA—Principal Component Analysis) before computing the t-SNE embedding of the training database and the input image in question. In this case, the PCA embedding of the training database can be stored in memory, after which it remains only to complete the embedding using the feature maps of the input image in question.
[0110] classification In step (c), the input image is classified in an unsupervised manner according to a feature map with a reduced number of variables, i.e., its t-SNE embedding.
[0111] It will be appreciated that any technique that allows for descriptive analysis of the t-SNE embedding space can be used. In particular, all the information of the training database is already contained within it, and therefore, it is sufficient to look at the spatial organization of this embedding space to reach conclusions about classification.
[0112] The easiest is to use the k-NN method (k-NN stands for k-nearest neighbors).
[0113] The idea is to look at the neighborhood of a point corresponding to the feature map of one or more input images in question and look at their classification. For example, if the neighborhood is classified as "no division," we can assume that the input image in question should also be classified as "no division." Note that the neighborhood considered may be limited, for example, depending on the strain, antibiotic, etc. Figure 7 shows two example t-SNE embeddings obtained for an E. coli strain in response to various concentrations of cefpodoxime. In the top example, two blocks are clearly visible, visually indicating the existence of a minimum inhibitory concentration (MIC) beyond which morphology and, therefore, cell division are affected. Vectors near the top are classified as "division," while vectors near the bottom are classified as "no division." In the bottom example, we can see that only the highest concentration stands out (and therefore appears to have an antibiotic effect).
[0114] computer program products According to a second and third aspect, the present invention relates to a computer program product comprising code instructions for executing (in particular on a data processing means 3, 20 of a server 1 and / or a client 2) a method for classifying at least one input image representing target particles 11a-11f in a sample 12, as well as storage means readable by a computing device (memory 4, 21 of the server 1 and / or client 2) on which this computer program product is stored.
Claims
1. 1. A method for classifying at least one input image representing a target particle in a sample, comprising: (b) extracting a feature map of the target particle from the input image; (c) constructing an embedding of the extracted feature maps to reduce the number of variables by a t-SNE algorithm; (d) performing unsupervised classification of the input image by analyzing an embedding space for the feature map with a reduced number of variables; A method comprising:
2. 2. The method of claim 1, wherein the particles are represented in a uniform manner in the input image and in each elementary image, in particular, centered and aligned in a predetermined direction in the input image and in each elementary image.
3. 3. The method of claim 2, comprising the step of: (a) subtracting the input image from a global image of the sample to represent the target particles in the uniform manner.
4. 4. The method of claim 3, wherein step (a) comprises segmenting the overall image to detect the target particles in the sample, and then cropping the input image to the detected target particles.
5. 5. The method of claim 3 or 4, wherein step (a) comprises obtaining the overall image from an intensity image of the sample, the image being obtained by a viewing device.
6. 6. The method of claim 1, wherein the feature map is a vector of numerical coefficients each associated with one base image of a set of base images each representing a reference particle, and step (b) comprises determining the numerical coefficients such that a linear combination of the base images weighted by the coefficients approximates a representation of the target particle in the input image.
7. 6. The method of claim 1, wherein the feature map of the target particle is extracted in step (b) by a convolutional neural network pre-trained on a public image database.
8. 8. The method of claim 1, wherein step (c) comprises defining an embedding space for each feature map of a training database of already classified feature maps of particles in a sample by the t-SNE algorithm and for the extracted feature map, wherein the feature map with the reduced number of variables is a result of embedding the extracted feature map into the embedding space.
9. The method of claim 8 , wherein step (d) comprises implementing a k-nearest neighbor algorithm in the embedding space.
10. 10. The method of claim 1, wherein step (b) comprises concatenating in three dimensions the extracted feature maps of each input image in the sequence.
11. 1. A system for classifying at least one input image representing a target particle in a sample, comprising at least one client comprising data processing means, said data processing means comprising: - extracting a feature map of said target particle through analysis of said at least one input image; - constructing an embedding of said feature maps to reduce the number of variables by the t-SNE algorithm; - performing unsupervised classification of the input image by analyzing the embedding space for the feature map with a reduced number of variables; A system configured to perform the steps of:
12. The system of claim 11 , further comprising a device for observing the target particles in the sample.
13. 11. A computer program product comprising code instructions for performing the method of any one of claims 1 to 10 for classifying at least one input image representing a target particle in a sample when said program is run on a computer.
14. 11. A storage medium readable by a computer device, the storage medium comprising a computer program product comprising code instructions for performing the method of any one of claims 1 to 10 for classifying at least one input image representing target particles in a sample.
Citation Information
Patent Citations
Object tracking method, object tracking device, and program
JP2018026108A