A method for classifying a sequence of input images representing particles in a sample over time.

A CNN-based method processes a sequence of images as a 3D stack to classify bacterial responses to antibiotics efficiently, overcoming the limitations of traditional antibiogram methods by reducing processing time and computational demands.

JP7845610B2Active Publication Date: 2026-04-14BIOMERIEUX SA +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BIOMERIEUX SA
Filing Date
2021-10-19
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing methods for analyzing the metabolic state of bacteria after antibiotic application, such as creating antibiograms, are lengthy, complex, and require chemical markers with cytotoxic effects, limiting observation times and interpretation of image analysis.

Method used

A method using a convolutional neural network (CNN) to classify a sequence of input images by concatenating them into a three-dimensional stack and processing them directly with a 3D convolutional layer, activation, pooling, and fully connected layers, without intermediate feature extraction.

Benefits of technology

Enables efficient and less intensive classification of bacterial images, allowing early analysis of antibiotic susceptibility without damaging the sample, reducing processing time and computational intensity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007845610000001
    Figure 0007845610000001
  • Figure 0007845610000002
    Figure 0007845610000002
  • Figure 0007845610000003
    Figure 0007845610000003
Patent Text Reader

Abstract

The present invention relates to a method for classifying a sequence of input images representing target particles (11a-11f) in a sample (12) over time, characterized in that the method comprises the following steps executed by a data processing means (20) of a client (2): (b) concatenating the input images in the sequence as a three-dimensional stack; and (c) directly classifying the three-dimensional stack using a convolutional neural network (CNN).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optical acquisition of biological particles. The biological particles can be, for example, microorganisms such as bacteria, fungi, or yeast. It can also include any other particles of the type of contaminating particles including cells, multicellular organisms, or dust.

[0002] The present invention is particularly advantageously applicable, for example, to analyzing the state of biological particles in order to know the metabolic state of bacteria after the application of antibiotics. According to the present invention, for example, an antibiogram of bacteria can be created.

Background Art

[0003] An antibiogram is an experimental technique aimed at testing the phenotype of a bacterial strain against one or more antibiotics (multiple possible). Conventionally, an antibiogram is created by culturing a sample containing bacteria and an antibiotic.

[0004] European Patent Application No. 2603601 describes a method for creating an antibiogram by visualizing the state of bacteria after an incubation period in the presence of an antibiotic. To visualize the bacteria, the bacteria are marked with a fluorescent marker that can reveal their structure. Then, by measuring the fluorescence of the marker, it can be determined whether the antibiotic has effectively acted on the bacteria.

[0005] The conventional process for determining an effective antibiotic against a bacterial strain involves taking a sample containing the strain (e.g., from a patient, animal, or food batch) and then sending the sample to an analytical laboratory. Upon receiving the sample, the analytical laboratory first begins culturing the bacterial strain to obtain at least one colony, with the incubation period ranging from 24 to 72 hours. From this colony, several samples containing various antibiotics and / or varying concentrations of the antibiotic are then prepared and incubated again. After another incubation period, also ranging from 24 to 72 hours, each sample is manually analyzed to determine whether the antibiotic was effective. The results are then returned to the practitioner to determine which antibiotic and / or the most effective antibiotic concentration to apply.

[0006] However, the marking process is particularly long and complex to carry out, and these chemical markers have cytotoxic effects on bacteria. Consequently, this visualization mode does not allow observation of bacteria at certain moments during bacterial culture; therefore, a sufficiently long culture time of about 24 to 72 hours is required to ensure the reliability of the measurement. Another method for visualizing biological particles is to use a microscope, which allows for non-destructive measurement of the sample.

[0007] Digital holographic microscopy (DHM) is an imaging technique that overcomes the limitations of depth of field in conventional optical microscopy. In general terms, it involves recording a hologram formed by the interference between light waves diffracted by the object being observed and a reference wave with spatial coherence. This technique is described in the review article "Principles and techniques of digital holographic microscopy" by Myung K. Kim, published in SPIE Reviews, Vol. 1, No. 1, January 2010.

[0008] In recent years, the use of digital holographic microscopy has been proposed for the automated identification of microorganisms. Accordingly, international application WO2017 / 207184 describes a method for acquiring particles that integrates simple, unfocused acquisition associated with digital reconstruction of focus, enabling observation of biological particles while limiting acquisition time.

[0009] Typically, this solution allows for the detection of structural modifications to bacteria in the presence of antibiotics after only about 10 minutes of incubation, unlike the conventional process described above which can take several days, and the detection of its sensitivity (detection of the presence or absence of division or patterns indicating division) at the end of 2 hours. In fact, since the measurement is non-destructive, it is possible to perform the analysis very early in the culture process without risking damaging the sample and thus extending the analysis time.

[0010] To visualize the behavior of particles, such as their migration speed or their cell division process, it is even possible to track particles with multiple sequential images to form a film representing the evolution of the particles over time (since the particles do not change after the initial analysis).

[0011] Therefore, it is understood that this visualization method yields excellent results. For example, when it is intended to reach conclusions, especially automatically, regarding the susceptibility of bacteria to antibiotics present in a sample, it is difficult to interpret these images or this film on their own.

[0012] Various techniques have been proposed, ranging from simply counting bacteria over time to "morphological" analyses aimed at detecting specific "configurations" through image analysis. For example, when bacteria are preparing to divide, two poles appear in the distribution well before the division itself, resulting in two distinct parts of the distribution.

[0013] The paper "A rapid antimicrobial susceptibility test based on single-cell morphological analysis" by Choi J., Yoo J., Lee M. et al. (2014), Science Translational Medicine, 6(267), https: / / doi.org / 10.1126 / scitranslmed.3009650, proposes combining two techniques to evaluate the effectiveness of antibiotics. However, as the authors emphasize, their approach requires very precise calibration of a certain number of thresholds, which are highly dependent on the nature of the morphological changes caused by the antibiotic.

[0014] More recently, a deep learning-based approach is described in the paper "Phenotypic Antimicrobial Susceptibility Testing with Deep Learning Video Microscopy" by Yu H., Jing W., Iriya R. et al. (2018), Analytical Chemistry, 90(10), 6314-6322, https: / / doi.org / 10.1021 / acs.analchem.8b01128. The authors propose using a convolutional neural network (CNN) to extract morphological features and features related to bacterial movement. However, this solution has proven to be highly computationally intensive on the one hand, and requires a large base of training images to train the CNN.

[0015] Therefore, the technical problem addressed by the present invention is to provide a more efficient and less intensive solution for classifying images of biological particles. [Overview of the Initiative]

[0016] According to a first aspect, the present invention relates to a method for classifying a sequence of input images representing target particles in a sample, the method comprising the following steps performed by a client data processing means: (b) A step of concatenating the above input images of the sequence in the form of a three-dimensional stack; (c) A step of directly classifying the three-dimensional stack by a convolutional neural network (CNN), wherein the CNN comprises a series of convolutional blocks consisting of a “3D” convolutional layer that applies a four-dimensional filter to a four-dimensional input feature map to generate a four-dimensional output feature map, and an activation layer and a 3D pooling layer, then a flattening layer, and finally one or more fully connected layers.

[0017] According to advantageous and non-limiting features:

[0018] The particles are represented uniformly in each input image, particularly positioned at the center and aligned in a predetermined direction.

[0019] This method includes step (a) extracting an overall image of the sample from each input image so as to show the target particles in the uniform manner described above.

[0020] Step (a) includes segmenting the overall image for each input image so as to detect the target particles in the sample, and then cropping the input image over the detected target particles.

[0021] Step (a) includes obtaining the overall image from the intensity image of the sample (12) acquired by the observation device.

[0022] The above 3D stack has two spatial and temporal dimensions, and the above filter and feature map have three first dimensions: the above spatial and temporal dimensions, and a fourth dimension: semantic depth.

[0023] The filter of the 3D convolutional layer has a depth equal to the depth of the input feature map, and the output feature map has a depth equal to the number of filters of the 3D convolutional layer.

[0024] The CNN includes only two convolutional blocks.

[0025] The method includes step (a0) of learning the parameters of the classifier from a learning base of a previously classified sequence of images of particles in the sample by data processing means of a server.

[0026] According to a second aspect, a system for classifying a sequence of input images representing target particles in a sample over time is proposed, comprising at least one client equipped with data processing means, the data processing means being configured to perform the following: - concatenating the input images of the sequence in the form of a three-dimensional stack; - directly classifying the three-dimensional stack by a convolutional neural network (CNN), the CNN comprising a "3D" convolutional layer that applies a four-dimensional filter to a four-dimensional input feature map to generate a four-dimensional output feature map, an activation layer and a 3D pooling layer, then a flattening layer, and finally one or more fully connected layers (if any), and classifying by a series of convolutional blocks.

[0027] According to an advantageous and non-limiting feature, the system further comprises a device for observing the target particles in the sample.

[0028] According to third and fourth aspects, a computer program product is proposed that includes code instructions for performing the method according to the first aspect for classifying a sequence of input images representing target particles in a sample, and storage means is proposed that can be read by an item of computer equipment that includes code instructions for performing the method according to the first aspect for classifying a sequence of input images representing target particles in a sample.

Brief Description of the Drawings

[0029] Further features and advantages of the present invention will become apparent from the following description of the preferred embodiments, which will be provided with reference to the accompanying drawings. [Figure 1] It is a diagram of an architecture for implementing the method according to the present invention. [Figure 2] An example of an apparatus for observing particles in a sample, which is used in a preferred embodiment of the method according to the present invention, is shown. [Figure 3a] Obtaining an input image in one embodiment of the method according to the present invention is shown. [Figure 3b] Obtaining an input image in a preferred embodiment of the method according to the present invention is shown. [Figure 4] Steps of a preferred embodiment of the method according to the present invention are shown. [Figure 5] An example of a convolutional neural network architecture used in a preferred embodiment of the method according to the present invention is shown.

Modes for Carrying Out the Invention

[0030] architecture The present invention relates to a method for classifying a sequence of input images representing particles 11a to 11f present in a sample 12, which are called target particles. It should be noted that this method can be implemented simultaneously for all or some of the particles 11a to 11f present in the sample 12, each being considered a target particle in turn.

[0031] As will be understood, this method includes a machine learning component, particularly a convolutional neural network (CNN).

[0032] The input or training data is of image type and represents target particles 11a-11f in sample 12 (in other words, it includes images of the sample in which the target particles are visible). The above sequence consists of multiple input images of the same target particles 11a-11f over time. As can be understood, if several particles are considered as necessary, it is possible to have multiple sequences of images of particles 11a-11f of sample 12 as input.

[0033] Sample 12 consists of a liquid such as water, buffer, culture medium, or reactive medium (which may or may not contain antibiotics) in which the particles 11a to 11f to be observed are located.

[0034] In an alternative embodiment, sample 12 may be in the form of a solid medium, preferably translucent, such as agar, on which particles 11a to 11f are located. Sample 12 may also be a gaseous medium. Particles 11a to 11f may be located inside the medium or on the surface of sample 12.

[0035] Particles 11a–11f may be microorganisms such as bacteria, fungi, or yeast. This may also include any other particles of the contaminating particle type, including cells, multicellular organisms, or dust. Throughout the remainder of the description, a preferred example is used in which the particles are bacteria (and, as understood, sample 12 contains antibiotics). The sizes of the observed particles 11a–11f vary between 500 nm and several hundred μm, and even several millimeters.

[0036] The “classification” of an input image sequence involves determining at least one class from a set of possible descriptive classes for the image. For example, in the case of bacterial-type particles, there may be a binary classification, i.e., two possible classes of effect: “with fission” or “without fission,” indicating resistance to antibiotics or non-resistance to antibiotics, respectively. The present invention is not limited to any particular type of classification, although examples of binary classification of the effect of antibiotics on the target particles 11a to 11f described above will be primarily explained.

[0037] This method is implemented by Server 1 and Client 2 within the architecture shown in Figure 1. Server 1 is a learning device (which implements the learning method), and Client 2 is an item of an operating device (which implements the classification method), such as a doctor's or hospital's terminal.

[0038] While it is entirely possible to combine two devices 1 and 2, preferably, server 1 is a remote device and client 2 is a consumer device, particularly a desktop computer, portable device, etc. Client device 2 is typically connected to observation device 10 to be advantageous for direct processing, so that it can directly acquire the input image (or, as can be seen below, "raw" acquired data such as a whole image of sample 12, or even an electromagnetic matrix), or alternatively, the input image is loaded into client device 2.

[0039] In any case, each item of equipment 1 and 2 is typically a remote computing device linked to a local network or a wide area network such as the Internet for exchanging data. Each comprises a processor-type data processing means 3 and 20, and a data storage means 4 and 21 such as computer memory, for example, flash memory or a hard disk. Client 2 typically includes a user interface 22 such as a screen for interaction.

[0040] Server 1 advantageously stores a training database, i.e., a set of sequences of images of particles 11a-11f already classified under various conditions (see below) (associated with labels indicating, for example, "fission present" or "non-fission" indicating susceptibility or resistance to antibiotics). Note that the training data may be associated with labels defining test conditions, such as "strain," "antibiotic condition," and "time" for bacterial cultures.

[0041] acquisition As explained, even if the method can directly take any image of target particles 11a to 11f obtained by any method as input, the method preferably begins with step (a) of acquiring an input image from data supplied by the observation device 10.

[0042] In known ways, those skilled in the art can use, in particular, digital holographic microscopy (DHM) techniques, such as those described in international application WO2017 / 207184. Specifically, an intensity image of the sample 12, called a hologram, can be obtained, which is not in focus on the target particles (considered as an "out of focus" image), and can be processed by data processing means (which are, for example, integrated into the apparatus 10 or the apparatus 20 of client 2, see below). It is understood that the hologram "represents" all the particles 11a-11f in the sample in a particular manner.

[0043] Figure 2 shows an example of an apparatus 10 for observing particles 11a to 11f present in a sample 12. The sample 12 is placed between a light source 15 that is spatially and temporally coherent (e.g., a laser) or a pseudo-coherent light source 15 (e.g., a light-emitting diode, a laser diode) and a digital sensor 16 that is sensitive in the spectral region of the light source. Preferably, the light source 15 has a low spectral width, e.g., less than 200 nm, less than 100 nm, and even less than 25 nm. Throughout the remainder of the description, for example, the central emission wavelength of the light source is referred to in the visible region. The light source 15 is oriented toward a first surface 13 of the sample and emits a coherent signal Sn that is transmitted by a waveguide, for example, an optical fiber.

[0044] The sample 12 (typically a culture medium, as described) is housed in an analysis chamber that is vertically separated by a lower slide and an upper slide, for example, conventional microscope slides. The analysis chamber is laterally separated by an adhesive or any other sealing material. The lower and upper slides are transparent to the wavelength of the light source 15, and the sample and chamber allow, for example, more than 50% of the wavelength of the light source to pass through under perpendicular incidence to the lower slide.

[0045] Preferably, particles 11a-11f are placed in the sample 12 on the upper slide. For this purpose, the underside of the upper slide contains a ligand for attaching the particles, for example, a polycation (e.g., poly-L-lysine) in the context of microorganisms. This allows the particles to be contained with a thickness equal to or close to the depth of field of the optical system, i.e., less than 1 mm thick (e.g., a tube lens), preferably less than 100 μm thick (e.g., a microscope objective lens). Nevertheless, particles 11a-11f can move within the sample 12.

[0046] Preferably, the apparatus comprises an optical system 23, for example, formed by a microscope objective lens and a tube lens, and positioned in air at a certain distance from the sample. The optical system 23 optionally includes a filter that can be positioned in front of the objective lens or between the objective lens and the tube lens. The optical system 23 is characterized by its optical axis, its object plane (also called the focusing plane) at a certain distance from the objective lens, and its image plane which is conjugate to the object plane by the optical system. In other words, an object located in the object plane has a corresponding sharp image of this object in the image plane, also called the focal plane. The optical properties of the system 23 are fixed (e.g., a fixed-focus optical system). The object plane and the image plane are perpendicular to the optical axis.

[0047] The image sensor 16 is positioned in or near the focal plane, facing the second surface 14 of the sample. The sensor, for example, a CCD or CMOS sensor, comprises a periodic two-dimensional array of sensing base areas and proximity electronic equipment that adjusts the exposure time and area reset in a manner known to itself. The output signal of the base areas depends on the amount of radiation in a spectral region incident on the area during the exposure time. This signal is then converted into image points, or "pixels," of a digital image by, for example, the proximity electronic equipment. Thus, the sensor generates a digital image in the form of a matrix having C columns and L rows. Each pixel of this matrix with coordinates (c,l) corresponds in a manner known to itself to a position in Cartesian coordinates (x(c,I), y(c,I)) in the focal plane of the optical system 23, for example, to the center of a rectangular sensing base area.

[0048] The pitch and fill factor of the periodic array are selected to conform to the Nyquist-Shannon criterion with respect to the size of the observed particles, thereby defining at least two pixels per particle. Thus, the image sensor 16 acquires a transmitted image of the sample in the spectral region of the light source.

[0049] The image acquired by the image sensor 16 contains holographic information insofar as it arises from the interference between waves diffracted by particles 11a-11f and a reference wave that passed through the sample without interacting with the sample. Clearly, as described above, in the context of a CMOS or CCD sensor, the acquired digital image is an intensity image, and therefore, in this case, it is understood that phase information is encoded in this intensity image.

[0050] Alternatively, the coherent signal Sn generated from the light source 15 can be split into two components, for example, by a translucent slide. The first component then acts as a reference wave, and the second component is diffracted by the sample 12, with the image in the image plane of the optical system 23 arising from the interference between the diffracted wave and the reference wave.

[0051] Referring to Figure 3a, in step (a), it is possible to reconstruct multiple overall images of sample 12 from the hologram, and then extract each input image from the overall images of the sample.

[0052] In fact, it will be understood that the target particles 11a-11f need to be represented uniformly in each input image, and in particular, they need to be centered and aligned in a predetermined direction (e.g., horizontally). The input images also need to have a standardized size (it is desirable that only the target particles 11a-11f are visible in the input images). Therefore, the input images are called "thumbnails" and can be defined with a size of, for example, 250 x 250 pixels. As long as a sequence of input images is desired, for example, one image can be taken every minute over a time interval of 120 minutes, thus obtaining a sequence of 120 input images.

[0053] The reconstruction of each overall image is performed by the data processing means of the device 10 or the data processing means 20 of the client 2, as described above.

[0054] Typically, a series of complex matrices called "electromagnetic matrices" are constructed (for the moment of acquisition) to model the wavefronts propagating along the optical axis for multiple deviations of the optical system 23 relative to the focal plane, particularly deviations located within the sample, based on the intensity image of the sample 12 (hologram).

[0055] These matrices can be projected into real space (e.g., via the Hermitian standard) to form a stack of the overall image at various focal lengths.

[0056] From the above, it is possible to determine the average focal length (and select the corresponding overall image, or recalculate it from the hologram), or to determine the optimal focal length for the target particle (and again select the corresponding overall image, or recalculate it from the hologram).

[0057] In any case, referring to Figure 3b, step (a) advantageously includes segmenting the whole image to detect the target particles in the sample, and then cropping it. In particular, each input image can be extracted from one of the whole images of the sample so as to represent the target particles in the uniform manner described above.

[0058] Generally, segmentation allows for the detection of all particles of interest by removing any artifacts, such as filaments or microcolonies, to improve one or more overall images. One of the detected particles is then selected as the target particle, and a corresponding thumbnail is extracted. As described, this process can be performed for all detected particles.

[0059] Segmentation can be performed using any known method. In the example in Figure 3b, fine segmentation is performed to remove artifacts, followed by coarse segmentation to detect particles 11a-11f in this case. Those skilled in the art can use any known segmentation technique.

[0060] To obtain a sequence of input images for target particles 11a to 11f, a tracking technique can be employed to track any movement of particles from one overall image to the next.

[0061] As seen on the right side of Figure 3a, all input images acquired for the sample (over time for some, or even all, particles of sample 12) are pooled to form a description base for sample 12 (in other words, an experimental description base), which should be noted in particular to be copied to the storage means 21 of client 2. The "field" level is referred to in contrast to the "particle" level. For example, if particles 11a-11f are bacteria and sample 12 contains or does not contain antibiotics, this description base would include all information about the growth, morphology, internal structure, and optical properties of these bacteria across the entire field of acquisition. As understood, this description base can be sent to server 1 for inclusion in the learning base described above.

[0062] Lamination As you can see, this method differs in that it can act directly on a sequence of input images without requesting or working on each image individually, or extracting feature maps in an intermediate way. Furthermore, it turns out that a very simple and lightweight CNN is sufficient to perform reliable and efficient classification.

[0063] Referring to Figure 4, the method includes step (b) concatenating the above input images of the sequence in the form of a three-dimensional stack, or in other words, a 3D "stack". More specifically, since all the input images have the same size and form a sequence of matrices, in order to obtain a three-dimensional stack, it is simply necessary to stack them in the order of the input images.

[0064] Therefore, just as an RGB image is a two-dimensional object with three channels, as can be seen below, even if this method processes this stack in a very original way as a single three-dimensional object with a single channel (for example, if there is an input image of size 250x250 and one image is acquired every minute for 120 minutes, it will have a size of 250x250x120), this three-dimensional stack can be considered an image with the same number of channels as the moment of acquisition. The first two dimensions are conventionally spatial dimensions (i.e., the size of the input image), and the third dimension is the "time" dimension (the moment of acquisition).

[0065] Preferably, step (b) includes downsampling the three-dimensional stack, i.e., reducing the size of the input.

[0066] The downsampling described above can be performed on the time dimension and / or spatial dimension of the stack, preferably both.

[0067] especially: - In terms of the time dimension, this can be reduced by dividing the acquisition period into n intervals and selecting n+1 images from the sequence corresponding to the end of these intervals, for example, keeping only 5 images from a sequence of 120 images, in short, by taking one image every 120 / 4 = 30 minutes (in particular, the images acquired at the end of 1, 30, 60, 90, and 120 minutes). However, it is still possible to select images in a non-uniform manner (for example, by selecting more images at the beginning of the acquisition period than at the end). - In terms of spatial dimensions, they can be reduced using any image downsampling technique with a given sampling coefficient, e.g., a coefficient of 2, on each axis (in order to maintain the ratio).

[0068] By implementing the two downsampling methods mentioned above, we can transition from a stack of size 250x250x120 to a stack of size 125x125x5 (the stack size is reduced to approximately 1 / 100th).

[0069] It should be noted that this downsampling can actually be done before generating the stack (by selecting and modifying the input images).

[0070] classification

[0071] In step (c), the 3D stack described above is directly classified by a suitable convolutional neural network called a "3D CNN" due to its ability to process the 3D objects that constitute the stack. In fact, as mentioned above, it is important to understand that the 3D stack is processed by the CNN as a single 3D object (i.e., with a single channel), rather than as a 2D object with several channels (as in the case of an RGB image, for example).

[0072] The terms direct classification or "end-to-end" are understood to mean that there is no separate extraction of at least one feature map of the target particles 11a-11f, and the CNN naturally has internal states in the form of feature maps, but these are never sent outside the CNN, and the CNN has the classification result as its sole output.

[0073] As a reminder, CNNs are generally well-suited for visual tasks, more specifically, for image classification. Generally, CNNs use multiple convolutional layers, and this 3D CNN uses at least one 3D convolutional layer that models the spatiotemporal dependencies between various input images.

[0074] A "3D convolutional layer" is understood to mean a convolutional layer that applies a 4D filter, and therefore can act on multiple channels of a stack that is already 3D, in other words, a 4D feature map. In other words, a 3D convolutional layer applies a 4D filter to a 4D input feature map to produce a 4D output feature map. The fourth and last dimension, as with any feature map, is the semantic depth.

[0075] This should be distinguished from conventional convolutional layers, which can only operate on 3D feature maps that represent some channels of a 2D object (image).

[0076] While this concept of 3D convolution may seem counterintuitive, it is a generalization of the concept of a convolutional layer, which expects only that multiple “filters” with a depth equal to the number of channels in the input (i.e., the depth of the input feature map) are applied by scanning them across all dimensions of the input (in 2D of the image), and the number of filters defines the output depth.

[0077] Therefore, the 3D convolution referred to herein applies a 4D filter with a depth equal to the number of channels in the 3D stack in the input, and scans these filters across the entire volume of the 3D stack, and thus across the time dimension as well as the two spatial dimensions, i.e., in 3D (hence the name 3D convolution). Thus, one 3D stack, i.e., a 4D feature map, is obtained for each filter. In conventional convolutional layers, the semantic depth (number of channels) in the output can certainly be increased by using a large number of filters, but a 3D feature map will always exist.

[0078] While 3D convolutional layers remain heavier and require greater computational power, it is understood that, as explained, a very simple architecture (and far simpler than known CNNs such as VGG616) is sufficient. Figure 5 shows the architecture of this embodiment of the 3D CNN.

[0079] Traditionally, this architecture advantageously includes a set of "convolutional" blocks consisting of 3D convolutional layers, activation layers (e.g., ReLU functions) to increase the depth of the feature map, and 3D pooling layers that allow the size of the feature map to be reduced (generally by half). It is worth noting that two convolutional blocks are sufficient, and therefore very preferably, this 3D CNN includes only two convolutional blocks.

[0080] A "3D pooling layer" is understood to mean a layer that can act on a 4D feature map, which already has one or more channels in a stack that is 3D in terms of 3D convolution. In other words, the size reduction is performed across all dimensions of the 3D stack, i.e., the first three dimensions of the 4D feature map.

[0081] Throughout the remainder of this specification, a clear distinction is made between the number of "dimensions" of feature maps in the geometric direction, i.e., the number of independent directions in which these cards extend (for example, a vector is a one-dimensional object, an image is two-dimensional, and this feature map is four-dimensional), and the number of "variables" of these feature maps, in other words, the size in each dimension, i.e., the number of independent degrees of freedom (which actually corresponds to the concept of dimension in a vector space, and more specifically, a set of feature maps with a given number of variables constitutes a vector space with dimensions equal to the number of these variables).

[0082] Therefore, in the example in Figure 5, the 3D CNN starts with six layers distributed across two blocks, as described. The first block includes a convolution + ReLU sequence (a first 3D convolutional layer and an activation layer with a ReLU function) that takes a 3D stack with a single channel as input (thus forming an object of size 125 × 125 × 5 × 1 when favorably downsampled as proposed above) and increases the depth to 30, followed by a max pooling layer (it is also possible to use overall average pooling) with a 62 × 62 × 2 × 30 feature map as output (the 3D pooling layer acts as described for three dimensions as well as two spatial dimensions, and therefore involves division by 2 including the time dimension).

[0083] In the illustrated example, the first 3D convolutional layer uses 30 filters with dimensions of 3 × 3 × 3 × 1, and therefore requires ((3 * 3 * 3 * 1) + 1) * 30 = 570 parameters.

[0084] The second block has the same architecture as the first block, and generates a 62×62×2×60 (twice the depth) feature map as output from a new convolution + ReLU set (a second 3D convolutional layer and an activation layer with a ReLU function), and generates a 12×12×1×60 feature map as output from the max pooling layer (note that in this case the spatial size is reduced to one-fifth, but the temporal size is always reduced to one-half).

[0085] In this case, the second 3D convolutional layer uses 60 filters with dimensions of 3 × 3 × 3 × 30, and therefore requires ((3 * 3 * 3 * 30) + 1) * 60 = 32,460 parameters.

[0086] At the output of the final convolutional block (in this case, the second), the 3D CNN advantageously includes a "flattening" layer that transforms the "final" feature map (containing the "deepest" information) at the output of this block into a vector (a one-dimensional object). Thus, for example, a 12×12×1×60 feature map transitions into a vector of size 12*12*1*60=8,640. It should be understood that there are no restrictions on the size of any map / filter at any level, and the sizes mentioned above are merely examples.

[0087] Finally, in the conventional method, upon completion, there is one or more fully connected layers (FCs, or "dense" layers as shown in Figure 5), and optionally a final activation layer, such as a softmax layer. In the illustrated example, the first layer FC transforms a vector of size 8,640 into a smaller vector of size 100 (which requires (8,640+1)*100=864,100 parameters), and the second layer FC transforms a vector of size 8,640 into a final vector of size 2 (which requires (100+1)*2=202 parameters).

[0088] Preferably, the 3D CNN consists of (i.e., strictly includes) a sequence of convolutional blocks, then flattening layers, and finally one or more fully connected layers.

[0089] This final part of the 3D CNN returns the expected result, in this case, the class of the input image sequence (a vector of size 2 corresponds to the binary result).

[0090] Therefore, the total number of parameters is less than 900,000, which is remarkably low for a CNN (generally tens of millions of parameters), especially considering the fact that the input data is already a large sequence of images. Thus, this 3D CNN can be used by multiple clients, including those with moderate computing resources.

[0091] Preferably, the method may include a step (a0) in which the data processing means 3 of server 1 learns the parameters of a 3D CNN from a training base. In fact, this step is typically performed quite far upstream, in particular by remote server 1. As described, the training base may include a certain amount of training data, in particular a sequence of images associated with their classes (e.g., "split" or "non-split" for binary classification).

[0092] Training a 3D CNN can be done using conventional methods. The learning cost function can consist of attachments to conventional "cross-entropy" data that are minimized via a gradient descent algorithm.

[0093] In all embodiments, the learned parameters of the CNN can be stored in the data storage means 21 of the client 2 for use in classification, if necessary. Note that the same CNN can be embedded in multiple clients 2, and only one training stage is required.

[0094] Computer program products According to second and third aspects, the present invention relates to a computer program product comprising code instructions for executing a method (particularly on data processing means 3, 20 of server 1 and / or client 2) for classifying an input image sequence representing target particles 11a to 11f in a sample 12, and storage means readable by an item of computer equipment (memories 4, 21 of server 1 and / or client 2) in which the computer program product is found.

Claims

1. A method for classifying a sequence of input images representing target particles in a sample over time, wherein the client's data processing means... (b) A step of concatenating the input images of the sequence in the form of a three-dimensional stack, (c) A step of directly classifying the three-dimensional stack by a convolutional neural network (CNN), wherein the CNN comprises a series of convolutional blocks consisting of a “3D” convolutional layer that applies a four-dimensional filter to a four-dimensional input feature map to generate a four-dimensional output feature map, and an activation layer and a 3D pooling layer, then a flattening layer, and finally one or more fully connected layers. A method characterized by including the implementation of

2. The method according to claim 1, wherein the target particles are represented in a uniform manner in each input image, and in particular are centrally located and aligned in a predetermined direction.

3. The method according to claim 2, comprising step (a) extracting an overall image of the sample from each input image so as to show the target particles in the uniform manner.

4. The method according to claim 3, wherein the extraction step (a) comprises, for each input image, segmenting the overall image so as to detect the target particles in the sample, and then cropping the input image over the detected target particles.

5. The method according to claim 3 or 4, wherein step (a) includes obtaining the overall image from the intensity image of the sample obtained by the observation device.

6. The method according to any one of claims 1 to 5, wherein the three-dimensional stack has two spatial dimensions and a time dimension, and the four-dimensional filter and feature map have three first dimensions: the spatial dimensions and the time dimension, and a fourth dimension: semantic depth.

7. The method according to claim 6, wherein the filters of the 3D convolutional layer have a depth equal to the depth of the 4D input feature map, and the 4D output feature map has a depth equal to the number of filters of the 3D convolutional layer.

8. The method according to any one of claims 1 to 7, wherein the CNN comprises only two convolutional blocks.

9. The method according to any one of claims 1 to 8, comprising the step (a0) of learning the parameters of the CNN from a learning base of previously classified sequences of images of particles in the sample by a data processing means of a server.

10. A system for classifying a sequence of input images representing target particles in a sample over time, comprising at least one client equipped with data processing means, wherein the data processing means is - The input images of the sequence are concatenated in the form of a three-dimensional stack, - Classifying the three-dimensional stack directly using a convolutional neural network (CNN), wherein the CNN consists of a series of convolutional blocks comprising a "3D" convolutional layer that applies a four-dimensional filter to a four-dimensional input feature map to generate a four-dimensional output feature map, and an activation layer and a 3D pooling layer, followed by a flattening layer, and finally one or more fully connected layers. A system characterized by being configured to carry out the following.

11. The system according to claim 10, further comprising an apparatus for observing the target particles in the sample.

12. A computer program, which, when executed on a computer, includes code instructions for performing the method according to any one of claims 1 to 9 for classifying a sequence of input images representing target particles in a sample over time.

13. A storage means readable by an item of computer equipment, which stores a computer program including code instructions for performing on a computer the method according to any one of claims 1 to 9 for classifying a sequence of input images representing target particles in a sample over time.

Citation Information

Patent Citations

  • Systems and Methods For Analyzing Perfusion-Weighted Medical Imaging Using Deep Neural Networks

    US20180374213A1