Method for detecting, classifying, and evaluating the motility of sperm cells in an image or video
The method employs super-resolution imaging and neural networks for automated sperm detection and classification, addressing the limitations of existing sperm motility assessment methods by enhancing resolution and reducing human intervention, especially in conditions with low sperm count and high debris.
Patent Information
- Application Number
- JP2024577066
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-20
- Filing Date
- 2023-06-26
- Publication Date
- 2025-07-17
AI Technical Summary
Existing sperm motility assessment methods suffer from limited spatial and temporal resolution, require human intervention, and are subjective, especially in cases of low sperm count and high debris, such as azoospermia.
A method using super-resolution imaging and neural networks for automated sperm detection and classification, capable of functioning under low magnification and low resolution conditions, with features like superpixel resolution, background removal, and latent space feature determination.
Enables accurate, automated, and repeatable sperm motility evaluation, even in adverse conditions, by improving resolution and reducing human intervention, particularly effective in cases of low sperm count and high debris.
Smart Images

Figure 2025522820000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of automated visual inspection means for sperm detection and classification, including the provision of hardware and methods for motility analysis.
Background Art
[0002] Motility per se is the ability of sperm to properly traverse the female reproductive tract and reach the egg. Methods for assessing sperm motility generally attempt to indirectly assess the reproductive survival rate of a sperm population, and sperm motility generally correlates with the fertilization success rate. For the automated or semi-automated assessment of sperm motility, several different devices and methods have been developed, including typical sperm analysis, time-lapse microscopy, frame-by-frame playback video microscopy, spectrophotometry, stroboscopy, and various computer analysis methods.
[0003] Computer motility analysis can provide objective measurements of sperm motility characteristics obtained from the tracking of a large number of sperm. Such measurements include the percentage of motile sperm, the percentage of progressively motile sperm (i.e., above preset threshold points for speed and curvature of movement so as to correlate with conventional manual assessment), the amplitude of lateral head displacement during progressive movement, and measurements of linear and curvilinear velocity.
[0004] Typical sperm analysis evaluates the percentage of motile sperm and a finer "motility grade" from grade A (rapid progressive - swimming forward rapidly in a nearly straight line) to grade D for immotile sperm. The motility of rapidly progressive sperm is generally considered the most reliable measure of sperm movement for predicting the fertilizing ability of a semen sample, but the "rapidity" here is a subjective assessment. These evaluations have conventionally been performed manually by a skilled technician using a microscope, usually equipped with phase contrast optics and a warming stage.
[0005] Such evaluation of sperm motility generally includes the total sperm motility rate (the proportion of sperm showing motility of any form), the progressive sperm motility rate (the proportion of sperm showing rapid straight-line movement), and sperm velocity on an arbitrary scale from 0 [immotile] to 4 [rapid motility]. For example, a motility of 75 / 70(4) indicates that 75% of the sperm are motile, 70% of the sperm have progressive motility, and they are moving "rapidly" across the microscopic field of view.
[0006] However, all of these methods have various drawbacks, including limited spatial and temporal resolution, the use of old-fashioned motion evaluation generally limited to the measurement of average linear velocity (when measured), and the need for varying degrees of human intervention and analysis.
Summary of the Invention
[0007] The present invention includes systems and methods adapted to quantitatively and repeatedly evaluate sperm motility, which provide several improvements over the prior art with respect to both the resolution of the measurements obtained and the nature of the motility parameters that can be evaluated.
[0008] In particular, the present invention, firstly, provides means and methods for achieving super-resolution for evaluating position and motion at the sub-pixel level, and secondly, models the motion of motile sperm by fitting several motility parameters to the observed motion. This latter enables a better characterization of sperm motion by introducing quantitative measurements of linearity and curvature as well as "higher-order" motion, as described in detail hereinafter.
[0009] A further method for automatically detecting sperm cells (or other types of cells or particles) is disclosed that uses a neural network adapted to function under both good imaging conditions and bad imaging conditions such as low magnification and low resolution. Detection is the first step necessary for further automated analysis including motility assessment. For such purposes, a training stage is used that can employ both automatically generated training data and more standard human-annotated training data, and upon completion of this stage, the network can automatically detect sperm cells (stationary or dynamic) within a given field of view, possibly in real time. The training stage can generally use methods such as backpropagation that employ a probabilistic approach to training in batches.
[0010] The method is particularly adapted to address cases of azoospermia where the number of sperm cells is very low (most of which are not swimming) and there is a lot of debris within the imaging field of view. With this method, much of this debris can be automatically digitally removed to facilitate subsequent analysis.
[0011] For classification purposes, a further method for determining the latent space features or hidden features of a neural network from static images (morphological features) and / or videos (dynamic features), and an unsupervised method for generating such features and enabling training on unlabeled data are disclosed. The use of latent space features enables a new standardization of sperm image and video analysis by indicating what the most important static and dynamic features of sperm analysis are.
[0012] Using the method of the present invention, not only can the quality of sperm cells be automatically sorted, but sperm cells within an image or video can also be automatically detected and distinguished from debris.
[0013] The above embodiments of the present invention have been described and illustrated in conjunction with systems and methods, but these are for illustrative purposes only and are not limiting. Further, just as all specific references can, but do not have to, embody a specific method / system, ultimately, such teachings are applicable to all representations regardless of the use of a specific embodiment.
[0014] Embodiments and features of the present invention are described herein in conjunction with the following drawings.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0016] The present invention will be understood from the following detailed description of preferred embodiments, which are for illustrative purposes and not limiting. For the sake of brevity, some well-known features, methods, systems, procedures, components, circuits, etc. are not described in detail.
[0017] Hardware Configuration The present invention can use a standard biological cell imaging configuration (for example, including one or more phase contrast or other types of microscopes, a sample stage including x-y axis control and optionally z-axis control, any heated sample stage, and a video camera). In principle, a mobile phone including appropriate magnifying means may be used instead of a dedicated imaging configuration. Alternatively, images or videos taken using such a method can be analyzed by software or other means of the present invention.
[0018] Superpixel resolution The present invention provides means and methods for achieving super-resolution in order to evaluate position and movement at the sub-pixel level, enabling subsequent image processing steps described below to be performed with higher accuracy than is possible by other means.
[0019] The present invention generally uses a method for super-resolution, which is a technique for achieving a higher resolution than the camera hardware by using a plurality of lower-resolution images together. This can be achieved by one or more means. Hereinafter, a method for increasing the number of images acquired at a given time will be described, and then several methods for combining the information obtained from these images to achieve superpixel resolution will be explained.
[0020] Movement of the sample stage To cope with the movement of the sample stage (e.g., when the stage is moving to analyze a new area of the sample), batch motion cancellation can be used. This can be done using continuous estimation of the motion vectors across the entire field of view. For this motion, the respective motion vectors of the individually tracked moving objects can be subtracted to find the "absolute motion" of these objects. This feature enables continuous tracking and continuous accurate motion evaluation of sperm cells even while the operator is moving the microscope stage. The motion vectors can be calculated as the highest amplitude percentile motion vectors detected in the sampling space within the field of view, which is divided, for example, into equal rectangles of 1 / 10 of the width × 1 / 10 of the length of the field of view (or other ratios of the field of view to the size of each individual cell). Utilizing a fixed pattern on the sample dish (e.g., a grid, a ruler, or other easily identifiable and trackable objects) is also included in the provision of the present invention, based on which the basic motion vector of the field of view is estimated. Such subtraction should be adjusted to the "frequent" average motion vector percentile with a minimum threshold energy.
[0021] The alternative embodiment averages the motion of the tracked object and, from this value, calculates the average motion of the slide when such motion is found to have sufficient energy (e.g., measured with respect to the kinetic energy averaged over several time frames). This approach may have limitations when the motion of the sperm cells is spatially restricted (e.g., by the edge of the water droplet in which the sperm cells are swimming). Further, such calculations may be biased when there are significant effects of chemotaxis or thermotaxis, i.e., when there is a chemical or thermal stimulus to the individual sperm cells and the sperm cells swim selectively towards the source of the stimulus, in which case the motion vectors of the field of view cannot be accurately estimated. However, since chemotaxis and / or thermotaxis can be relevant parameters for healthy sperm, these cases may have special utility. Accordingly, providing a chemical gradient and / or a thermal gradient in the sample dish and observing the motion of the sperm under the influence of these gradients is included in the present disclosure. By observing both the motion with and without the gradient, the selective motion under the gradient can be understood. Thus, the sample dish may initially be neutral, and after motion has been observed, a thermal or chemical gradient is introduced and new motion is observed, and the difference and thus the contribution of the gradient can be found.
[0022] A further method is to employ an x-y (and optionally z) stage that includes either, or both, an electric, shaft encoder, linear encoder, or other means for position reading. In this case, the movement of the stage is known from either, or both, the shaft encoder and the operation of the motor, and thus can be described by the above algorithms. In the case of superpixel resolution, the measured / known movement can be used as an initial input to an algorithm that is adapted to further improve these initial estimates of the stage movement. Further methods that can be considered for this purpose are the 8-parameter projective motion model, block matching, and Horn-Schunck optical flow estimation (see “Subpixel Motion Estimation for Super-Resolution Image Sequence Enhancement,” Journal of Visual Communication and Image Representation Volume 9, Issue 1, March 1998, Pages 38-50).
[0023] Recording and background removal A method of recording and / or background removal can be employed as a first step before superpixel and other subsequent operations.
[0024] When the camera is arranged in a substantially fixed manner, various types of background removal can be performed to increase the resolution. The background image can be evaluated by using a moving average over time, or, for example, a "median image" can be used, in which case the background image is calculated for each pixel as the median (not the average value) for that pixel over some history. This tends to more completely remove motion artifacts due to, for example, the movement of sperm within the image than in the case of taking an average value image. This background, however it is calculated, can then be subtracted from a given frame to more clearly show only the moving objects. Since not all elements of the moving object necessarily move between frames (in other words, when motion occurs between regions having the same pixel value), rather than marking only the pixels that move only between frames as indicating motion, for example, any pixel that has moved within the last N frames can be marked. A segmentation method can be employed to identify the entire moving object (in this case, usually a moving sperm), for example, based on a conditional random field, a neural network, or other segmentation methods.
[0025] When the camera is not fixed relative to the sample stage, various methods of image recording can be employed to record consecutive images such that consecutive frames result in a "matching" situation with respect to the background, reducing the effect of relative movement between the camera and the sample stage. The recording can be achieved, for example, by using image features such as SURF features or SIFT features (well-known to those versed in the field of computer vision) in two frames to be recorded and calculating a homography that transforms the coordinate system between the images using the features common to both frames.
[0026] Filtering of irrelevant objects To avoid tracking objects that appear to be moving within the input frame but are irrelevant to the measurements being made, several techniques can be used.
[0027] Object exclusion based on magnification scale and size can be used. For example, a micropipette with a surface area substantially larger than that of a typical sperm cell head can be excluded from the list of moving objects to be tracked. This is to prevent "wasting" expensive computer resources when processing the motion of irrelevant objects. Therefore, segmentation techniques (e.g., based on conditional random fields, neural networks, or other segmentation means) can be used to segment such objects and remove the objects from the frames in which they appear. An example of an image containing both sperm and a pipette is shown in Figure 3, where, as described above, the much larger pipette can be excluded from further consideration. In Figure 3, sperm cell 301 and capillary head 302 can be seen.
[0028] Object exclusion based on morphology can also be used. For example, tissue debris can usually be removed even if it is on the same scale as sperm cells. This can be done by morphology-based means such as curvature-based measurements and moment-based measurements, or by neural network methods. Debris has no "tail" (as expected of mature sperm cells), so it is morphologically different from sperm cells and can be excluded from tracking based on this. An example of an image containing both sperm and a pipette is shown in Figure 3, where, as described above, the morphologically different pipette can be excluded from further consideration.
[0029] Another way to exclude an object is based on the estimated average speed, i.e., whether this speed is above or below a certain threshold. Again, this is to prevent "wasting" expensive computer resources on irrelevant objects. For example, an operator can move the pipette (manually) at a speed much higher than the maximum speed of sperm cells, and in this case, this can be used as indicating an irrelevant target.
[0030] Adaptive noise reduction thresholding (e.g., using a constant false alarm rate (CFAR) method) and overall vibration damping can also be employed using the following steps. 1. Perform a spatial (i.e., grid) calculation of the statistical values of "noise" within each grid region. 2. Define a "low-pass filter" criterion. This determines the accumulation of historical values for each pixel so as to increase the sensitivity of the detection mechanism when the "duty cycle" of the movement is low. 3. Define a "threshold percentile". Exceeding this is considered to be an over-threshold. 4. Define an N / M "persistence criterion" and consider only the pixels that exceed the "threshold percentile" value within each grid as "persistent" targets as opposed to pseudo-noise.
[0031] All of these methods can be improved by the aforementioned background subtraction methods, such as using the median image as an indicator of the stationary background to be removed.
[0032] Luminance calibration Calibration steps can be performed at the beginning of the process (e.g., before or after recording / background removal) to remove the effects of variable lighting. Fluorescence, LEDs, or other ambient lighting leaking into the image may have periodic hills and valleys effects, also known as the "beat" phenomenon, which increases the brightness of one frame and decreases the brightness of other frames. Similarly, fluctuations in the power supply of microscope lighting can also cause such effects. To address this situation, steps of automatic photometric equalization can be performed, where, for example, a "reference frame" is calculated as the average over many frames (in some cases, a moving average that gradually changes over time), and the photometric (or brightness, contrast, histogram, white balance, entropy, or other parameters) of subsequent frames is adjusted to match the "reference frame". The purpose of such automatic calibration is to equalize the photometric differences between input frames before combining the input frames in subsequent steps. Similarly, automatic detection and removal of hot pixels and dead columns can be performed to remove the effects of CCD defects.
[0033] Increase in the number of images To obtain a large number of frames with minimal movement in the observation image, the camera frame rate can be operated as high as possible, for example, in the "burst mode" available in some cameras or by using a camera with a particularly high frame rate. As is understood, the higher frame rate itself may involve a greater noise trade-off unless the system uses a larger aperture and / or brighter lighting to achieve the same image brightness as that achieved at a lower frame rate.
[0034] Second, when color information is not of interest (in the case of sperm motility tracking, this color information may actually be redundant), each of the R, G, and B planes of a given color image can be used separately as intensity images to provide three images for each acquired frame. Such an approach may make it possible to increase the signal-to-noise ratio of the target image.
[0035] Combination of images Since there is a set of recorded and calibrated images or a "stacked" set, various means can be employed to combine these images.
[0036] Averaging The first method of superpixel resolution is to use the averaging of a set of images to achieve superpixel resolution. This is the simplest method and uses the average value of all pixels within the stack calculated for each pixel.
[0037] Median The second method is to employ the median (which is also useful for background subtraction). This method uses the median value of the pixels within the set calculated for each pixel of the image.
[0038] Maximum value In this method, the maximum value of all pixels within the stack is calculated for each pixel. This can be useful for debugging purposes to show all defects in all calibration images.
[0039] Kappa-sigma clipping This method is used to iteratively reject deviant pixels and utilizes two parameters, namely the number of iterations and the standard deviation multiplier (kappa). For each iteration, the average value and standard deviation (sigma) of the pixels within the stack are calculated. Each pixel having a value furthest from the average value greater than kappa*sigma is rejected. The average value of the remaining pixels within the stack is calculated for each pixel.
[0040] Median kappa-sigma clipping This method is similar to the kappa-sigma clipping method, but instead of rejecting the outlier pixel values, they are replaced with the median value.
[0041] Adaptive weighted average This method calculates the robust average obtained by iteratively weighting each pixel with respect to its deviation from the mean value as a ratio of the standard deviation (see The Techniques of Least Squares and Stellar Photometry with CCDs - Peter B. Stetson 1989).
[0042] Entropy-weighted average (high dynamic range) This method is based on the research of Jenkin and Lesperance in Germany (see Entropy-Based image merging - 2005), and it is used to stack a set of images into a final photo while maintaining the best dynamic range for each pixel.
[0043] Another method to achieve super-resolution uses a burst of CFA raw images with small offsets. As described in Handheld Multi-Frame Super-Resolution (https: / / arxiv.org / abs / 1905.03277), these frames are then aligned and combined to form a single image with red, green, and blue values at all pixel sites, which serves both to improve the image resolution and increase the signal-to-noise ratio.
[0044] The new image obtained by any of the above methods can be implemented as an image with a higher bit depth at the same resolution (e.g., by addition instead of averaging), or can have the same bit depth but lower noise due to averaging, or can have a higher spatial resolution at the same bit depth.
[0045] By using such a superpixel method as described above, the method can handle situations where the difference in the general position of sperm cells between frames, for example, is, for example, 0.3 pixels (in the original image without improved image quality), and on average more than 60% of the "naive" frames show no movement. Nevertheless, assuming that the motion detection process subtracts the "moving tails" of 10 recent frames, then even an average motion of less than 0.1 pixel per frame may be identifiable.
[0046] Parameterization of motion The present invention models the motion of motile sperm by fitting a curve based on several motion parameters to the observed motion. This latter enables a better characterization of sperm motion by introducing quantitative measures of linearity and curvature as well as "higher-order" motion, as described in the detailed description. These methods may enable both the estimation of trajectory parameters and the handling of sperm cell "collisions" by characterizing the motion.
[0047] In one approach, the best estimator can be selected from among a small number of estimators over a small number of motion models. The concept of "best" here can be quantified, for example, by using a cost function over the estimation error of each particular fitting model. For example, such a cost function can use the average rms distance between the fitted position and the observed position over a number of frames or over a time frame.
[0048] Motion estimator Kalman filtering can be used as a motion estimator. The position to be estimated may be, for example, the center of the sperm cell head. This position can be extrapolated forward based on polynomial curve fitting, sine curve fitting, or other suitable functions that fit the observed motion. A polynomial or other estimator can be calculated by using the linear velocity, rotational velocity, and acceleration, as well as the slalom-like trajectory. Alternatively, the motion of the sperm may be modeled as a sine curve, cardioid, nth-degree polynomial, etc. Whatever the type of motion used for modeling, the parameters of this motion can be estimated using a Kalman filter to estimate the subsequent or previous position of the sperm cell.
[0049] The model shown in Figure 1 shows three types of trajectories superimposed on each other. The thickest trajectory 101 is circular and has a radius R0, an angular velocity ω0, and a starting angle θ0.
[0050] The intermediate trajectory 102 is a sine curve and is superimposed on the above-mentioned zero-order circular motion and has an amplitude R1, an angular velocity ω1, and a starting angle θ1.
[0051] The finest trajectory 103 is also a sine curve and is superimposed on the above-mentioned zero-order and first-order trajectories and has an amplitude R2 (shown as "B" in the image), an angular velocity ω2, and a starting angle θ2.
[0052] Generally, the polar coordinate representation of the position of an individual sperm cell can be approximated by the following formula. R(t)=R0+R1*sin(ω1*t+θ1)+R2*Sin(ω2*t+θ2) θ(t)=θ(t)+ω0*t
[0053] Using Kalman filtering over at least a third degree of freedom (i.e., position, velocity, acceleration, acceleration derivative), the above parameterized R0, R1, R2, θ 0、 θ1, θ 2、 ω 0、 ω1, ω2, or parameterization R(t)=R0+V*t+1 / 2a2*t Various parameters such as R0, V, and a of
[0054] When the motion is fitted, the motility function can be calculated from the fitted motion parameters. For example, for the parameterization of a circle and a sine, the motility function is M = A / R0 + B|R 1,THRESH -R1| + C|ω 1,THRESH -ω1| + D|R 2,THRESH -R2| + E|ω 2,THRESH -ω2| may be.
[0055] Similarly, for the parameterization of the second velocity and acceleration, M = A / R0 + B|V THRESH -V| + C|a THRESH -a| can be the motility function. In both cases, the motility function defined here is better (indicating healthier sperm) the smaller the value of M. Alternatively, a bounded or unbounded motility function that increases in the case of healthier sperm may be defined. Furthermore, the use of a multi-valued motility function is also included in the provision of the present invention. For example, two values M R = A / R0 + B|R 1、THRESH -R1| + D|R 2、THRESH -R2| and M ω = C|ω 1、THRESH -ω1| + E|ω 2、THRESH -ω2| are calculated to show different health estimation values for the curvature M R of the path and the velocity M ω of the motion.
[0056] Overcoming sperm "collisions" (i.e., blocking of sperm trajectories) For the purpose of tracking individual sperm, the method described above uses Kalman filtering, i.e., the characterization of the movement by feed-forward of the position of each individual sperm cell head, based on various possible parameterizations. As described above, these parameterizations can include linear velocity, rotational velocity and acceleration, as well as slalom-shaped or sinusoidal trajectories, and a Kalman filter is used to feed-forward or estimate the subsequent position of the sperm cell and maintain the tracking of each individual sperm cell head.
[0057] An alternative approach is to use a temporal neural network such as a recurrent neural network (RNN) adapted for the purpose of motion prediction. The input to such an RNN can be a series of images, and the output can be the next one or more predicted frames. Alternatively, the input can be a series of images and labels for each sperm, and the output can be a set of the next predicted positions for each labeled sperm. A more detailed description of these methods is given below.
[0058] Determination of length scale Automatic scaling can be used by introducing a known spatial pattern, such as a grid or a ruler. Also, for example, by measuring the average length of the sperm cell tail, which is known to be about 50 μm, or by measuring a standard object introduced into the field of view, such as a pipette of known shape and dimensions, automatic scaling can be employed. If the objective magnification and the sample distance are known, the length scale can be calculated directly from this information. A further "natural ruler" may simply be the known average size of the sperm itself. Since this size is known (e.g., for a given donor population or an individual donor), the length scale can be determined with a certain level of accuracy based only on the size of the sperm, and the accuracy of the length scale improves with the number of visible sperm.
[0059] Pattern recognition for TESA / TESE Next, a pattern recognition method for identifying non-motile sperm cells in TESE / TESA tissue samples will be described. Identification of viable but stationary sperm cells in TESE / TESA tissue samples can be achieved using the vibration pattern of the sperm tail. In many cases, sperm cells in TESE / TESA tissue samples move their tails but do not appear to swim. This is because the sperm cells can be trapped in tissue debris or simply do not have enough energy. By applying the temporal Fourier transform to individual pixels or to individual spatial cells within a grid of the field of view (e.g., a 10×10 or 50×50 grid), regions with substantial energy content within a typical tail vibration frequency range of sperm cells, such as a frequency range of 6 - 14 Hz, can be identified. Subsequently, the system uses the presence of signals exceeding a threshold (with respect to energy) in specific spatial cells within the grid of the field of view to provide visual cues to the operator for these specific spatial cells. Thus, the operator positions the stage at the center of these cells, changes the microscope to a higher magnification, and checks whether motile sperm cells are placed on the stage. Automatically performing this operation of positioning the stage at the center by a computer-driven x-y (and optionally z) stage is included in the provision of the present invention. The tail vibration frequencies of both motile and stationary cells can be determined by the motion parameterization as described above, which is specifically adapted for the purpose of tail vibration frequency, for example, by a preliminary step of subtracting the mass center motion of the sperm to obtain a sample that is stationary but still has a vibrating tail.
[0060] Implementing a scoring mechanism to distinguish sperm cells with different tail vibration frequencies as an indication of the quality of each individual sperm cell is included in the provision of this specification. FIG. 2 shows, for example, the correlation between frequency and motility.
[0061] Automated detection of moving sperm cells, stationary sperm cells with vibrating tails, and completely stationary sperm cells within an image or video can be achieved by detecting sperm cells (or other types of cells or particles) using a deep neural network (e.g., a convolutional neural network (CNN)). One embodiment of the method is shown in FIG. 5.
[0062] This method can use any of the preprocessing steps described above, but even without such preprocessing, the described method is adapted to function under adverse imaging conditions such as low magnification, low resolution, variable illumination, and noisy environments (occurring in TESE / TESA samples).
[0063] During training, the neural network receives images of verified sperm cells (“labeled examples”) under the same imaging conditions, and the weights of the network are learned. As described above, these labeled examples can be obtained by manual or automated labeling for training. For example, multiple position frames can be derived from a single initial frame by using a typical computer vision tracking algorithm. The aforementioned motion subtraction algorithms such as median image subtraction can also be used. This is because only sperm move in most samples, and thus the only object remaining after motion subtraction is the sperm cell. This sperm cell can be manually verified after being automatically labeled using (for example) typical computer vision morphological operations.
[0064] An example of a training stage is shown in flowchart 610 of FIG. 5. The initial input 601 includes images or videos of swimming sperm cells, sometimes under poor imaging conditions. Since these sperm are swimming, automatic detection, tracking, and tagging are easily performed by a conventional computer vision (CV) motion detector / tracker 602. Motion detection and tracking can include dense optical flow that estimates motion vectors for all pixels in a video frame, sparse optical flow or Kanade-Lucas-Tomashi (KLT) feature tracker that tracks the positions of a few feature points in an image, Kalman filtering that can be used to predict the position of a moving object based on previous motion information, algorithms such as Meanshift and Camshift that locate the maximum value of a density function, and other object detectors and trackers that can be used in this situation. A single object tracker can be combined to perform multiple object tracking by re-identification, as is often the case with the above algorithms.
[0065] These detections generate tagged sperm cell images or videos 603 suitable for training a deep neural network 603 adapted for sperm detection. The detections can be verified by other means such as manual verification. Once an appropriate number of tagged image or video sequences are generated, these image or video sequences can be used to train a neural network 604 such as a deep convolutional neural network.
[0066] Next, an inference stage is possible after the training stage is completed. In inference, the network can automatically detect stationary or dynamic sperm cells in the imaging field of view. This is because the network has learned the characteristics of sperm cell images under various expected conditions (low magnification, low resolution, variable illumination, etc., as described above).
[0067] An example of the inference stage is shown in the flowchart 620 of FIG. 5. Here, an untagged image or video sequence 605 (of the kind used in the training step of flowchart 610) is input into a pre-generated, trained deep neural network 606. This network outputs an inference 607 indicating the position of the stationary (or non-stationary) sperm cells. The output may relate to a bounding box, pixel-level segmentation, etc., determined by the nature of the training data.
[0068] Inferences for any of the networks of the present invention can be performed in real time (e.g., by dedicated hardware).
[0069] This method is useful in cases of azoospermia in TESE / TESA samples, where (to view the entire sample) sperm cells may be found at low magnification, the number of sperm cells is very small, and most of them are not swimming. Furthermore, often there is a lot of debris in the imaging field. In this case, for training, images of swimming sperm cells (which can be automatically detected) are taken from another healthy sample (not azoospermic), the network is trained, and then during inference, all sperm cells in the azoospermic sample can be detected.
[0070] Viscous medium test Evaluating the motility parameters measured in the same cells in an aqueous medium with the sperm sample placed in a more viscous medium is included in the provision of the present invention in the following steps. 1. Obtain a threshold value for the velocity of human sperm cells in an aqueous medium, indicating the score for the quality of their motility. 2. Obtain the viscosity-velocity curve of human sperm cells. 3. Obtain the concentration-viscosity relationship of a specific sperm cell suspension medium, e.g., polyvinylpyrrolidone (PVP). 4. For each operator session, obtain a specific type of IVF medium and its concentration (relative to water). 5. For each measured velocity of an individual human sperm cell in the medium, apply the concentration-viscosity conversion and subsequently the viscosity-velocity conversion to evaluate the instantaneous equivalent velocity of the human sperm cell in the aqueous medium. 6. Compare the converted velocity of an individual sperm cell with the threshold value to obtain a score for the quality of movement of the specific sperm cell.
[0071] In particular, the relationship between the sperm velocity in various commonly used concentrations of PVP solutions and the sperm velocity in water is brought about by the following approximate relationships. · For PVP 7%, V_who = V_pvp * 1.9, or · For PVP 10%, V_who = V_pvp * 2.4 or · For water, V_who = V_pvp Here, V_who is the velocity in water and V_pvp is the velocity in various PVP solutions.
[0072] Flowchart The flowchart of an exemplary embodiment is shown in a simplified form in FIG. 6. Any step is shown in a dashed frame and the necessary steps are shown in a solid frame.
[0073] First, obtain an image sequence 501. This input can include one or more images and, in principle, the pipeline can function on a per-image or per-sequence basis (e.g., as may be useful in conjunction with an RNN for tracking as described below). It may be found useful to include as input other information such as the magnification and illumination of the system, sample parameters such as temperature, pH, viscosity, PVP concentration, and, in some cases, other information that can provide a context useful for subsequent analysis.
[0074] As described above, next, the steps of sample stage motion removal 502, recording 503, background removal 504, length scale determination 505, luminance and camera calibration 506, and image combination for superpixel resolution 507 may be employed. The order of these operations is not necessarily as listed in this example, and it may be found that different orders are more useful in different scenarios.
[0075] Next, ideally, the step of sperm detection 508 is performed in an "instance-aware" manner such that each sperm can be distinguished from other sperm and thus tracked between frames. Next, the motion parameterization 509 of individual and / or collective motion can be performed in several of the ways described above. Generally, a set of measurements related to the motility of individual sperm is calculated. Next, any step of non-motile sperm detection and parameterization 510 can also be performed (optionally using an image that has not undergone background removal 504, which may tend to exclude stationary portions of the sperm of interest).
[0076] Next, the calculations of the individual sperm obtained can be combined into sample-level calculations 511, which include, for example, the average number of motile sperm per cc or cm 2 or the average of motility measurements, or an overall ranking of the sample regarding more detailed information such as a histogram of the population size versus motility level or other measurements.
[0077] Next, the measurements obtained from the sample can be presented 512 to the user in some form, for example, via a suitable GUI or, optionally, another web interface remotely managed via a network. Next, the sequence can be repeated when acquiring a new image or image sequence 501. In some embodiments, the sample-level calculations and presentation may be performed only once every N iterations of the loop.
[0078] Pattern recognition - supervised learning for detection, tracking, classification, and regression The above method can be used in combination with machine learning, computer vision, and artificial intelligence to better determine various parameters of interest.
[0079] For example, the systems and methods of the present invention can use images or videos that have undergone the preprocessing as described above or are used directly. One method of the present invention for detecting sperm cells (or other types of cells or particles) involves using various types of deep neural networks (e.g., convolutional neural networks (CNNs) and / or recurrent neural networks (RNNs)) for detection, tracking, regression, and classification. These methods can be adapted to function under poor imaging conditions such as low magnification and low resolution by the aforementioned preprocessing in addition to training using images or videos with poor imaging conditions.
[0080] The analysis can be performed with respect to detection (e.g., generation of bounding boxes around each detection example or pixel-level segmentation), classification (e.g., good / medium / bad or motility levels 1 - 5), regression (e.g., motility levels on a scale of 1 - 100), tracking (ideally determination of the trajectory over time in an "instance-aware" implementation), and other outputs of the neural network model.
[0081] A training stage is employed, during which one or more neural networks receive images of verified sperm cells along with continuous variables such as class labels and / or velocities ( "labeled examples") under a set of some imaging conditions, and the weights of the network are learned. Training is performed using labeled examples of the desired type of output.
[0082] Training is realized by backpropagation of errors, and the weights of the neural network can be gradually changed in a direction that tends to minimize the loss function, which can be implemented using cross-entropy (for class labels), mean squared error (for continuous variables), etc. The loss is calculated in small batches using so-called stochastic methods, avoiding the need to use the entire training cohort for each training step, thus increasing the learning speed.
[0083] The labeled examples used in the supervised learning scenario can be obtained by manual labeling or by automatic or semi-automatic labeling. For example, considering a video of motile sperm cells without other motile elements (a "sample that is easy to detect"), the motile elements can be detected using methods such as background subtraction or typical computer vision-based motion detection, and the remaining motile sperm cells can be training examples and can be used as training examples. Tagging can be realized with respect to bounding boxes, pixel-level segmentation, or other means. Bootstrap tagging can also be used in this situation, which generates a low-precision network using a small number of training examples to tag more images, and these images are then corrected and used to train a slightly better network, and its output is used to train a third generation, etc.
[0084] Methods for (automatically) generating images of sperm cells for training deep learning networks can be used for further training of neural network classifiers. These methods include the use of parametric 3D models, the use of adversarial generation models based on a set of real 3D modeled images, and the use of other methods that may be considered useful.
[0085] After the completion of the training phase, the network can be used to perform inferences for the automatic detection, classification, and parameter estimation (regression) of sperm cells in static or dynamic situations. As long as training data under quasi-optimal conditions (such as poor resolution, lighting, jitter, etc.) is provided, it can also be expected that the inference (performed by the trained network) will handle such images or videos smoothly under similar conditions.
[0086] The inference of the network can be performed in real time (e.g., by dedicated hardware) or offline, and in some cases, it is performed with an online connection (thereby, for example, by an application programming interface (API) that defines a protocol for requesting and receiving such inferences, such that an image / video is sent to a server, analyzed, and the results are sent back to the user).
[0087] The present invention also includes providing a hybrid mode of operation, in which within the so-called "operation" mode of such a system, an operator can indicate to the system the locations where suspicious sperm cells are identified. In that case, the system adds that location to the training set locally, i.e., in a way that acts only on that specific operator, or on a workstation basis, i.e., on that specific system, or on a site basis, i.e., for a plurality of such systems located in a given laboratory, and finally globally, i.e., such that all similar systems are updated similarly.
[0088] Tracking can be achieved by several means. The simplest involves the use of the aforementioned system adapted for still image localization. More advanced techniques utilize convolutional neural networks adapted to take in multiple frames as input, and / or convolutional nets combined with recurrent neural networks (RNNs). By means of any of these techniques, the system can process videos for tracking purposes. Tracking is ideally performed in an instance-aware manner so as to track a given sperm even in cases of occlusion, interference, and overlap. Training in the case of videos can be done using tagged or annotated video sequences, which themselves may be semi-automatically generated using the still image identification / localization means described above, and which often involve instance tagging, which may be automatic in many cases, including correction by a human in cases of occlusion, interference, or overlap.
[0089] Next, using a network capable of processing image sequences, the dynamic characteristics of sperm can be measured with respect to classification (giving "class-level" outputs such as non-motile, low-motile, normal-motile, and highly-motile) or regression (giving, for example, a continuous score of motility in any unit or velocity unit, or measurements of physical characteristics such as power output, morphology, or frequency). In the latter case of measuring motility, velocity, or other dynamic characteristics, the output of the still image localization method described above can be used in conjunction with knowledge of the time difference between frames (and optionally the movement of the sample stage) to calculate velocity or other motility parameters. Further motility parameters that can be targeted include straightness of path, efficiency of swimming, forward velocity compared to "wobble" velocity, velocity in various media, functions of chemical gradients or other (e.g., temperature or pressure) gradients, and the like. Providing outputs regarding static or dynamic parameters of sperm behavior and morphology is included in the provision of the present invention.
[0090] Similar classifications regarding cell morphology can be made, for example, classifying into groups of "ideal", "acceptable", "poor", and "abnormal". Classification / regression in combination with some outputs is also included in the present invention's offering. For example, a single network can be trained to give outputs regarding both morphological classes and velocities.
[0091] Unsupervised techniques Utilizing unsupervised techniques of machine learning for the purpose of clustering or determining efficient latent representations that can then be used for classification is also included in the present invention's offering. For example, an autoencoder or another method for unsupervised feature learning can be used. In the case of an autoencoder, convolutional layers are used to reduce the information content of the input image or video sequence of sperm to a bottleneck where the network attempts to reproduce the original input. The point of this exercise is for the "bottleneck layer" to contain a compact representation or "latent space" representation of the input, and then this representation can be used to compare various inputs and show clusters of similar inputs, which can reflect the actual situation more faithfully than more or less any set of classes, and then can be manually annotated or determined by human approval.
[0092] For example, when an autoencoder is trained on static images or video sequences of sperm, it may be found that after training, the sperm images naturally cluster into three types, which in visual inspection seem to be a normal type and two different types of abnormal morphology or motility. Similarly, autoencoding of video sequences can reveal several distinct types of swimming ones, and this finding can be used in both clinical and research settings. Such a network, once generated, can be used for classification using cluster-based means such as the measured distance to the cluster center or similar means that will be apparent to those skilled in the art. Similarly, the compact representation of such a network can be fine-tuned with annotated data to overcome the situation of sparse training data.
[0093] Specific Application - Azoospermia In the case of azoospermia, the number of sperm cells is very low (most of them are not swimming). Therefore, wide fields of view and low magnifications are generally used (in some cases, to enable viewing a larger number of cells at once). In azoospermia, a biological sample is obtained by a semi-surgical testicular sperm extraction procedure called TESE or TESA, in which some amount of testicular tissue is excised or aspirated and then made into a cell suspension. As a result, most of the non-fluid content of such a sample is non-sperm cells and cell debris surrounding the sperm cells, and in most cases, these are either non-motile or, if motile, are stationary due to being trapped in tissue debris.
[0094] In this case, for the supervised learning scenario, in order to train the network, images of swimming sperm cells (which can be automatically detected by the use of motion detection, as will be apparent to those skilled in the art) can be taken, for example, from healthy samples. During inference, this trained network can then detect all sperm cells in an azoospermia sample even under poor imaging conditions, since the network has learned the morphology of the sperm cells.
[0095] Additional Method Here, a method for automatically detecting and classifying sperm cells based on both conventional features and hidden features detected by a neural network is described. These features can be extracted from still images (morphological features) and / or videos (dynamic features).
[0096] During training, the deep neural network uses training images (or videos) of sperm cells with labels (such as "normal" vs. "abnormal", etc., and / or continuous variables for regression) and learns features that characterize these inputs by calibrating the weights of the network.
[0097] For training, this labeling can be created manually (e.g., using the expertise of one or more embryologists), by a video object motion detection method, by imaging sperm cells that have passed through a sperm concentration assay (such as a swim-up assay, a microfluidic assay, a DNA fragmentation assay, a hyaluronic acid binding assay, etc.), or even by collecting sperm cells that have naturally reached the egg region within a female's body and are thus assumed to be "good" sperm cells. As mentioned above, classification is just one possibility; for example, regression is another possible approach for the network output (in this case, the input is an image labeled with one or more continuous values such as sperm quality on a scale, swimming speed, etc., rather than class labels).
[0098] The network encodes each sperm cell into a latent space vector that represents all the features (conventional and hidden features) of the sperm being examined, for example, by an autoencoder that has received appropriate training.
[0099] During inference, the network receives a new sperm cell image or video and provides an appropriate output (e.g., classifies those images or videos, or in the case of a network trained for regression, provides one or more continuous outputs).
[0100] Next, the differences between sperm cells can be quantified by the mathematical distance between the latent space vectors. The latent space can be, for example, the second-to-last layer of a neural network, before a regression or class-labeling layer.
[0101] This process can be used to automatically sort the quality of sperm cells, automatically detect sperm cells within an image or video, and distinguish them from debris.
[0102] To identify the specific features detected by a good cell network, the generative network can be used, which can take a latent space vector and use it "backwards" to generate an image of a good sperm cell. Subsequently, by varying the values of the latent space vector, the changing features in the generated image can be visually identified.
[0103] This can lead to a new standardization of sperm image and video analysis, indicating what the most important static and dynamic features of sperm analysis are. A similar approach is to find the best correlation between the latent space elements and the network output.
[0104] Overcoming sperm "collisions" (i.e., sperm trajectory blockages) As described above, one approach to enabling instance-aware tracking (which substantially overcomes common problems of tracking such as occlusion) is to use a temporal neural network such as a recurrent neural network (RNN) adapted for the purpose of motion prediction. The input to such an RNN (or an LSTM, GRU, or other network adapted for such purposes) may be a series of images, and the output may be the positions of a set of detected objects (uniquely identified) for each frame. As shown in FIG. 4, this input can pass through a preliminary CNN stage. Alternative inputs may be a series of images and labels for each sperm (e.g., instance labels, class labels, and / or continuous variable measurements for regression), and the output may be a set of predicted next positions and labels for each labeled sperm. Another possibility is to provide an output regarding a predicted motion vector, where each sperm cell is detected by the network and a predicted motion vector is assigned. This motion vector may be a vector that tends to generate the predicted position of the next frame of the video sequence (e.g.). Higher-order motion can also be extracted from such a network, along with (e.g.) outputs of position, velocity, and acceleration for each instance detection. If the motion of the sperm cells is parameterized (e.g., by considering the approximate sinusoidal shape of the position of the cells over time, which is caused by the swimming motion of healthy sperm cells), the network can be trained to output the coefficients of such parameterization.
[0105] The above description and illustration of embodiments of the present invention are presented for purposes of example. This is not exhaustive and does not limit the present invention in any way to the above description.
[0106] The terms defined above and used in the claims should be construed according to this definition.
[0107] The reference signs in the claims are used for facilitating the reading of the claims and are not part of the claims. These reference signs shall not be construed as limiting the claims in any way.
[0108] All features disclosed in this specification, including the claims, the abstract and the drawings, as well as all steps of the disclosed method or process, can be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification, including the claims, the abstract and the drawings, can be replaced by alternative features that perform the same, equivalent, or similar function, unless expressly stated otherwise.
Claims
1. A method for determining sperm motility, comprising: a. capturing a sequence of images of a sperm sample; b. tracking the position of each sperm in the image; c. fitting a function to the tracking for each sperm; d. calculating one or more motility functions based on the function; wherein one or more quantitative motility parameters are objectively measured.
2. The method according to claim 1, wherein the step of capturing a sequence of images is realized by a video camera observing a field of view through a microscope.
3. The method according to claim 1, wherein the step of tracking the position is performed using sub-pixel position accuracy.
4. The method according to claim 1, wherein the tracking is performed using a convolutional neural network.
5. The method according to claim 4, further comprising performing the tracking using a recurrent neural network.
6. The method according to claim 4, wherein the neural network is trained with artificial data generated by means selected from the group consisting of GAN, 2D model, and 3D model.
7. The method according to claim 3, wherein the tracking is performed using median image subtraction.
8. The method according to claim 1, wherein the fitting function is a set of equations
9. R(t) = R 0 + R 1 * sin(ω 1 * t + θ 1 ) + R 2 * Sin(ω 2 * t + θ 2 ) θ(t) = θ(t) + ω 0 * t
10. The method according to claim 4, wherein the motility function is defined by wherein the fitting function is R(t) = R 0 + V * t + 1 / 2 a 2 * t, the method according to claim 1.
11. The method according to claim 5, wherein the motility function is defined by M = A / R 0 + B | R 1,THRESH - R 1 | + C | ω 1,THRESH - ω 1 | + D | R 2,THRESH - R 2 | + +E|ω 2,THRESH -ω 2 |
12. The method according to claim 1, wherein the step of fitting is realized using a minimization method.
13. M = A / R 0 + B|V THRESH - V|+ C|a THRESH - a| The method according to claim 1, wherein the step of fitting is realized using a Kalman filter.
14. The method according to claim 1, further comprising eliminating the influence of the movement of the sample stage by means selected from the group consisting of determining the movement of the sample stage using a position encoder, determining the movement of the sample stage using an average particle velocity, and determining the movement of the sample stage using a known fixed pattern of the sample stage.
15. The method according to claim 1, further comprising measuring the velocity of the sperm in a polyvinylpyrrolidone solution to obtain the velocity of the sperm in water according to V_who = k(c) * V_pvp estimated based on the relationship, where V_who is the velocity of the sperm in water, V_pvp is the velocity of the sperm in the PVP solution, and k(c) is a constant according to the concentration of the polyvinylpyrrolidone solution, the method according to claim 1.
16. A method for analyzing sperm, a. obtaining training data composed of a set of labeled images or image sequences; b. training a neural network by backpropagation using the labels; c. predicting the label of an input image using the trained neural network and a method including.
17. The method according to claim 16, wherein the neural network includes a convolutional neural network incorporated into a recurrent neural network, and the labeled image sequence includes future instance labels and the position of the center of mass of the sperm head.
18. The step of obtaining training data a. obtaining a video sequence including sperm; b. performing any step of background removal; c. performing a motion detection step of generating a bounding box around the moving sperm and thereby automatically generating training data for the neural network from the video sequence, the method according to claim 16.
19. A method for identifying sperm cells according to claim 16, further performing hybrid identification of potential sperm cells in a given image, combining neural network inference with any manual inference by an operator and visual selection of such cells, in which case the system enhances its training data set to include the cells selected by the operator.
20. The method according to claim 19, wherein the neural network is retrained after manual inference.
21. The method according to claim 19, wherein the retraining is performed on an individual operator basis, in a work system, in a laboratory (i.e., all systems belong to the same laboratory), or on a regional or global basis.
22. The method according to claim 18, wherein the manual inference includes localization, identification, classification, and regression.
23. A method for generating a sperm image or video sequence from the amplified latent vectors of a trained network.
24. The method according to claim 23, wherein GAN is implemented to generate the image or video sequence.
25. The method according to claim 16, wherein the step of obtaining training data is realized by generating the training data using GAN.
26. A method for analyzing a sperm image sequence using unsupervised learning and a set of training data image sequences, comprising: a. training an autoencoder having a bottleneck layer with the training data image sequence, wherein the output of the bottleneck layer is useful as a latent representation of the image sequence; b. identifying the clustering of the latent representation in the training data image sequence with respect to a discontinuous number of population clusters; c. identifying the image sequence to be analyzed with respect to the membership of one or more of the clusters. The method includes.
27. A method for sperm analysis useful in the case of stationary sperm cells such as those occurring in azoospermia, comprising: a. In a preliminary stage, i. obtaining a video sequence of motile sperm; ii. tagging moving objects in the video sequence using a conventional motion detection algorithm; iii. using the tagged objects as training data for training a convolutional neural network; b. In an inference stage, inferring the positions of both stationary and moving sperm cells using the convolutional neural network. The method includes, where widely available training data of motile sperm cells is utilized to train a network adapted to detect sparse stationary sperm cells that may occur in azoospermia samples.
28. The method according to claim 27, wherein the step of obtaining a video sequence of motile sperm uses normal semen that is not azoospermic.
29. Furthermore, superpixel resolution is obtained by averaging a plurality of frames of the video sequence and using the averaged frames as an input to the convolutional neural network in both the training step and the inference step. The method according to claim 27.
Citation Information
Cited By
Information processing device, method, and program for tracking sperm
JP7840606B1