Path consistency self-supervised learning living cell high-resolution tracking method and system

Through the path consistency self-supervised learning method, the deep learning model is trained using unlabeled microscope videos, which solves the problems of high manual labeling cost and occlusion re-identification in living cell tracking, and realizes efficient and accurate morphology-vitality comprehensive evaluation and automated screening.

CN120689369APending Publication Date: 2025-09-23SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510846550.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies rely on manually labeled data in live cell tracking, which is costly and difficult to handle re-identification after long-term occlusion or wandering of the field of view. They are also unable to achieve comprehensive morphology-vitality assessment in high-resolution images, making the screening process subjective and inefficient.

Method used

A path consistency self-supervised learning method is adopted to train a deep learning embedding model using unlabeled high-resolution microscopy videos. The model is optimized using path consistency loss and entropy minimization terms to achieve robust appearance feature extraction of living cells, and is combined with the Hungarian algorithm for efficient matching and evaluation.

Benefits of technology

It completely gets rid of the dependence on manual labeling, improves the generalization ability and tracking robustness of the model, realizes long-term tracking and integrated morphology-vitality evaluation in high-density scenarios, and improves screening efficiency and standardization level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689369A_ABST
    Figure CN120689369A_ABST
Patent Text Reader

Abstract

The invention discloses a path consistency self-supervised learning living cell high-resolution tracking method, which comprises the following steps: S1, acquiring an unlabeled living cell microscope video, and intercepting a video slice from the video; s2, adopting a target detector to process the slices; s3, selecting a starting frame and an ending frame from the slices, determining any living cell as a query living cell from the starting frame, and generating a plurality of observation paths by randomly skipping frames in a middle frame sequence between the starting frame and the ending frame; s4, extracting feature vectors of the living cells in the path by using a deep learning embedding model, and calculating association probability distribution of all the living cells from query of the living cells to an end frame; and S5, constructing a path consistency loss and iterative optimization model. The invention further discloses a path consistency self-supervised learning living cell high-resolution tracking system. Based on path consistency, a deep learning embedding model is trained in a self-supervision mode, and a tracking system is constructed based on the deep learning embedding model, so that form-activity integrated intelligent optimization of living cells is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of biomedical image processing, artificial intelligence and assisted reproductive technology, and in particular to a method and system for high-resolution tracking of living cells through self-supervised learning based on path consistency. Background Art

[0002] In assisted reproductive technology (ART), selecting sperm with excellent morphology and high motility is a prerequisite for successful treatments such as in vitro fertilization (IVF) and intracytoplasmic sperm injection (ICSI). Currently, clinical screening relies primarily on visual observation and manual manipulation by embryologists. This process has numerous drawbacks, including high subjectivity and a lack of standardized standards; low efficiency and labor intensity; and limitations in resolution and separation of assessments to achieve a comprehensive assessment of "morphology and motility."

[0003] To overcome these shortcomings, the academic community has begun researching AI-based multi-object tracking (MOT) algorithms for sperm tracking. However, sperm tracking is a more challenging scenario than conventional MOT tasks, and existing technologies face the following bottlenecks in this application: 1. Tracking algorithms struggle to adapt to complex scenarios: Sperm samples are characterized by high density, high similarity, high velocity, and a high intersection rate. Tracking methods based on motion models (such as Kalman filtering) are prone to losing track of the target when sperm trajectories suddenly change or collisions occur. Tracking methods based on appearance features struggle to establish stable associations when faced with hundreds or thousands of sperm that are nearly indistinguishable in appearance. The accuracy of existing algorithms plummets, especially when frequent intersections and overlaps (i.e., occlusions) occur. 2. They rely heavily on manually annotated data: Current state-of-the-art supervised learning MOT models require large-scale, finely annotated training data (i.e., each sperm is labeled with a bounding box and unique ID in each frame of the video). For the sperm tracking task, completing such labeling is labor-intensive and costly, which fundamentally limits the application and promotion of supervised learning methods; 3. Limitations of existing self-supervised / unsupervised methods: In order to get rid of the dependence on labeling, some self-supervised methods have been proposed. For example, the method based on "cycle consistency" tracks the target from frame A to frame B and then back to frame A, and requires it to be able to associate back to itself for learning. However, such methods are usually only effective within a very short time span (such as 1-2 frames). For fast-moving objects such as sperm, once they are occluded for a long time or temporarily out of view, the simple cycle consistency assumption will fail, resulting in the model being unable to learn robust long-term tracking capabilities.

[0004] Therefore, technicians in this field are committed to developing a path consistency self-supervised learning living cell high-resolution tracking method and system. Summary of the Invention

[0005] In view of the above-mentioned defects of the prior art, the present invention at least solves the following technical problems:

[0006] Solve the problem that existing technologies for tracking living cells rely on manually labeled data, resulting in large labeling workload and high costs; solve the problem that existing algorithms are unable to handle the re-identification of living cells after they are blocked or out of view for a long time; solve the problem that existing technologies can no longer combine high-resolution images in a living state to achieve a comprehensive "morphology-vitality" assessment; solve the problem that the clinical screening process is subjective and inefficient.

[0007] To achieve the above objectives, the present invention discloses a method for high-resolution tracking of living cells using self-supervised learning based on path consistency, comprising the following steps:

[0008] S1: Obtain unlabeled high-resolution microscopy videos of living cells as training data, and extract video slices from the videos;

[0009] S2: processing each frame of the video slice using an object detector to locate and crop a bounding box image of the living cell;

[0010] S3: selecting a start frame and an end frame from the video slice, determining any live cell from the start frame as a query live cell, and generating multiple observation paths by randomly skipping frames in a sequence of intermediate frames between the start frame and the end frame;

[0011] S4: extracting the feature vectors of the living cells in the observation path using a deep learning embedding model, and calculating the association probability distribution from the query living cell to all living cells in the end frame;

[0012] S5: Constructing path consistency loss and iteratively optimizing the deep learning embedding model;

[0013] Furthermore, the deep learning embedding model is based on the ResNet-34 architecture;

[0014] Furthermore, the calculation of the association probability distribution includes: aggregating feature similarities of adjacent frames by matrix multiplication to derive the association probability between the query living cell and all living cells in the end frame;

[0015] Furthermore, the random frame skipping adopts different strategies, including skipping all intermediate frames, skipping odd frames, and randomly skipping several intermediate frames;

[0016] Furthermore, the path consistency loss includes a consistency term and an entropy minimization term;

[0017] Furthermore, the optimization of the deep learning embedding model also includes regularization loss;

[0018] Furthermore, the regularization loss includes one-to-one loss to ensure matching uniqueness and bidirectional consistency loss for forward-backward tracking consistency;

[0019] The present invention also discloses a path consistency self-supervised learning living cell high-resolution tracking system, which includes an image acquisition module, a core processing module, and a display and control module connected in sequence:

[0020] The image acquisition module is used to capture high-resolution dynamic video streams of living cells;

[0021] The core processing module includes a real-time detection unit, a feature embedding unit, a tracking association unit, a trajectory analysis and optimization unit, and a result visualization and interaction unit. The core processing module realizes living cell tracking by executing the tracking method;

[0022] The display and control module is used to provide a human-computer interaction interface to display real-time screening results;

[0023] Furthermore, the real-time detection unit is used to locate live cells in the video stream and crop the bounding box;

[0024] The feature embedding unit is used to call the trained deep learning embedding model to generate a living cell feature vector;

[0025] The tracking association unit assigns the live cells of the current frame to an existing track, creates a new track, or terminates the track based on the distance between the feature vectors through an efficient matching algorithm to handle the re-identification of live cells after they are occluded;

[0026] The trajectory analysis and optimization unit is used to calculate cell viability parameters and morphological parameters for comprehensive evaluation;

[0027] The result visualization and interaction unit is used to overlay and display the living cell ID, trajectory and optimization results;

[0028] Furthermore, the efficient matching algorithm is the Hungarian algorithm.

[0029] Based on "path consistency", the present invention trains a deep learning embedding model in a self-supervised manner that can extract robust appearance features of living cells, and based on this, constructs a system that can achieve accurate, long-term living cell tracking and optimization in high-resolution, high-density scenarios, so as to realize the intelligent optimization of the morphology and vitality of living cells such as sperm.

[0030] In general, the above technical solutions conceived by the present invention have at least the following advantages compared with the prior art:

[0031] Beneficial effects:

[0032] 1. Completely eliminate the reliance on expensive, time-consuming, and impractical manual data annotation. This makes it possible to use massive, easily accessible unlabeled microscope videos to train highly accurate and robust tracking models, greatly lowering the application threshold of the technology and improving the generalization ability of the model.

[0033] 2. Greatly improved tracking robustness in scenarios with high density and high crossing rates of living cells, such as sperm. Even if a sperm is completely obscured by other sperm for a long time, or temporarily swims out of the high-power microscope's field of view and then returns, the system can re-identify and associate the sperm with the correct trajectory based on learned stable appearance features, ensuring tracking continuity and accuracy.

[0034] 3. This technology achieves the integrated "morphological-motility" assessment and screening of individual sperm in vivo, enabling the selection of high-quality sperm that are both "well-rounded" and "excellent," potentially significantly improving the clinical outcomes of assisted reproduction.

[0035] 4. Transforming the traditional subjective, tedious, and inefficient manual screening process into an objective, accurate, and efficient automated process. This significantly reduces the workload of doctors, improves screening efficiency and standardization, and provides powerful technical support for precision medicine in the field of assisted reproduction. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of the method architecture of the present invention;

[0037] Figure 2 Schematic diagram of the deep learning embedding model training process of the present invention;

[0038] Figure 3 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0039] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0040] In the drawings, components with identical structures are denoted by the same reference numerals, and components with similar structures or functions are denoted by similar reference numerals. The size and thickness of each component shown in the drawings are arbitrary and are not limited by the present invention. For clarity, the thickness of components in some places in the drawings is appropriately exaggerated.

[0041] The present invention discloses a method for high-resolution tracking of living cells using self-supervised learning of path consistency. Figure 1 As shown, the method includes the following steps:

[0042] S1: Obtain unlabeled high-resolution microscopy videos of living cells as training data, and extract video slices from the videos;

[0043] S2: processing each frame of the video slice using an object detector to locate and crop a bounding box image of the living cell;

[0044] S3: selecting a start frame and an end frame from the video slice, determining any live cell from the start frame as a query live cell, and generating multiple observation paths by randomly skipping frames in a sequence of intermediate frames between the start frame and the end frame;

[0045] S4: extracting the feature vectors of the living cells in the observation path using a deep learning embedding model, and calculating the association probability distribution from the query living cell to all living cells in the end frame;

[0046] S5: Constructing a path consistency loss and iteratively optimizing the deep learning embedding model to extract robust appearance features of the living cells.

[0047] The present invention discloses a real-time, self-supervised multi-target tracking method and system for high-resolution living cell (especially human sperm) microscopy video. The tracking method includes a deep learning embedded model training process such as Figure 2 As shown, the process is fully automated and does not require any manual identity (ID) labeling.

[0048] Example 1

[0049] 1. Data preparation: Collect more than 500 anonymous, unlabeled live sperm high-resolution (e.g., 1000x objective) microscope videos from multiple reproductive centers as a training dataset.

[0050] 2. Training process:

[0051] Step 1: Data Input and Object Detection

[0052] High-resolution (e.g., 600-6000x) unlabeled live sperm microscopy videos are used as input data. An off-the-shelf object detector (off-the-shelf detector), such as the YOLO series model, is used to process each frame of the video to locate and crop the bounding box images of all sperm objects.

[0053] A video slice of 60 frames in length is randomly extracted from the input dataset, and the YoloV5 model is used as the object detector to detect all 180 sperm in the first frame of the video.

[0054] Step 2: Generate multiple observation paths

[0055] A start frame ts and an end frame te are selected from the video slice. For any query sperm oi,ts in the start frame ts, the system simulates and generates multiple different “observation paths” (π1,π2,…,π NFor example, path 1 observes all intermediate frames, path 2 may skip odd frames, and path 3 randomly skips a few frames. This simulates the process of tracking the same target under different observation conditions (such as different sampling rates).

[0056] One of the sperm is randomly selected as the query object oit1, and the end frame is set to the 60th frame. By randomly skipping frames with different strategies between the 2nd frame and the 59th frame, 30 different observation paths are generated.

[0057] Step 3: Calculation of cross-path association probability

[0058] For each generated observation path πk, the model calculates the probability distribution qk of association between the query sperm oi,ts and all sperm objects oj,te in the ending frame te. This calculation is performed using a deep learning embedding model. The model extracts the appearance and spatial features of each sperm image, forming a feature vector. The model then derives the final association probability from oi,ts to each object in oj,te by aggregating the similarities between sperm feature vectors across all adjacent frames along the path.

[0059] A deep learning embedding model based on ResNet-34 was used to extract a 128-dimensional feature vector for all sperm images along each path. The similarities between adjacent frames were aggregated through matrix multiplication to calculate the final association probability distribution qk for each path.

[0060] Step 4: Path Consistency Loss (PCL) calculation and model optimization

[0061] Regardless of how the observation path changes, the true identity of the query sperm oi,ts is unique and certain in the end frame te. Therefore, the association probability distribution finally calculated for all paths should be highly consistent.

[0062] Based on this, the present invention designs a path consistency loss, which mainly consists of two parts:

[0063] 1) Consistency term: penalize different paths (π1,π2,…,π N ) computed from the associated probability distributions (q1,q2,…,qN) (e.g., using KL divergence to measure the distance between each qk and the mean q^ of all distributions).

[0064] 2) Entropy minimization term: Encourages the minimization of the entropy of each qk, that is, drives the probability distribution toward a "one-hot" vector to ensure the uniqueness and certainty of the association results.

[0065] Step 5: Model training

[0066] The deep learning embedding model is iteratively optimized using the standard back-propagation algorithm, using PCL and some auxiliary regularization losses (e.g., one-to-one loss to ensure the matching is one-to-one, and bidirectional consistency loss to ensure forward-backward tracking consistency).

[0067] After training, the model was able to extract robust, distinguishable sperm appearance features that were resistant to long-term occlusion. The deep learning embedding model was trained using the Adam optimizer with a learning rate of 0.0001 on four NVIDIA A100 GPUs until the model loss on the validation set stopped decreasing.

[0068] The present invention also discloses a path consistency self-supervised learning living cell high-resolution tracking system, such as Figure 3 As shown, it includes connecting the image acquisition module, core processing module and display and control module in sequence:

[0069] The image acquisition module is used to capture high-resolution dynamic video streams of living cells;

[0070] The core processing module includes a real-time detection unit, a feature embedding unit, a tracking association unit, a trajectory analysis and optimization unit, and a result visualization and interaction unit. The core processing module realizes live cell tracking by executing the aforementioned path consistency self-supervised learning live cell high-resolution tracking method;

[0071] The display and control module is used to provide a human-computer interaction interface and display real-time screening results.

[0072] Example 2

[0073] like Figure 3 The path consistency self-supervised learning high-resolution living cell tracking system shown in the figure realizes a fully automated process from image acquisition to final optimization. The system includes:

[0074] 1. Image acquisition module: It consists of a high-resolution microscope and a high-speed industrial camera connected to it, and is used to capture the dynamic video stream of living cell samples (human sperm is used as an example in this embodiment) in real time.

[0075] 2. Core processing module: This is a high-performance computing device (such as a workstation or server with a GPU) that deploys the core algorithm software of the present invention. It is mainly composed of the following units:

[0076] 2.1 Real-time detection unit: Receives video frames from the module, runs the object detection algorithm, accurately locates all sperm in each frame in real time, and outputs their bounding box coordinates;

[0077] 2.2 Feature Embedding Unit: Each sperm image output by the detection unit is input into the deep embedding model trained through the above self-supervised process, generating a vector for each sperm in real time that can accurately represent its fine morphological and appearance features;

[0078] 2.3 Tracking Association Unit: This unit is the core of tracking. It receives the feature vectors of all sperm in the current frame and matches them with the existing tracking tracks in the system. It uses the distance between feature vectors (such as cosine distance) as the association basis, and uses efficient matching algorithms (such as the Hungarian algorithm) to assign sperm in the current frame to existing tracks (i.e., "update tracks"), create new tracks (for sperm that have just entered the field of view), or terminate tracks (for sperm that have left the field of view). Based on the model's powerful long-distance matching capabilities, this unit can effectively handle the situation where sperm reappears after being blocked;

[0079] 2.4 Trajectory Analysis and Optimization Unit: Quantitatively analyzes all continuously tracked sperm trajectories. It can calculate a series of key performance indicators (KPIs) for each sperm, including but not limited to:

[0080] 2.4.1 Vitality parameters: average path velocity (VAP), curve linear velocity (VCL), straight line linear velocity (VSL), linearity (LIN), etc.

[0081] 2.4.2 Morphological parameters: Based on high-resolution images, analyze the head size / shape, acrosome integrity, mid-section length / curvature, etc.

[0082] Each sperm is scored in real time based on clinically pre-defined criteria (e.g., speed, linearity, and morphological normality).

[0083] 2.5 Results Visualization and Interaction Unit: This unit visually overlays tracking and analysis results onto the video screen in real time. Each tracked sperm is assigned a unique ID and a trailing trajectory line. Crucially, the sperm or sperm currently ranked as "optimal" by the optimization unit are highlighted with a prominent indicator (e.g., a green highlight box), providing the operator with clear, intuitive visual guidance.

[0084] 3. Display and Control Module: This module provides a human-computer interface for users (clinicians, embryologists, etc.). This module allows users to view real-time tracking and screening video, view detailed parameters of each sperm, and fine-tune optimal parameters or perform final manual confirmation as needed.

[0085] Example 3

[0086] Through Figure 3 The path consistency self-supervised learning high-resolution living cell tracking system shown in the figure has the following system operation examples:

[0087] 1. Start the system: On the clinical ICSI operating table, the embryologist prepares living cells (human sperm is used as an example in this embodiment) and starts the system of the present invention.

[0088] 2. Real-time acquisition and processing: The camera of the image acquisition module captures video stream at a rate of 50 frames per second and transmits it to the core processing module in real time.

[0089] 3. Initialization: The real-time detection unit detects 210 sperm in the first frame. The feature embedding unit immediately uses the trained model to generate feature vectors for these 210 sperm. The tracking association unit then initializes 210 new tracking trajectories based on these vectors.

[0090] 4. Real-time tracking and association: For each subsequent frame, the real-time monitoring unit, feature embedding unit and tracking association unit work together. Suppose that in the 150th frame, the sperm with ID "S_007" overlaps with the sperm with ID "S_102", causing "S_007" to be completely blocked for 15 frames (ie 0.3 seconds). In the 165th frame, "S_007" reappears. At this time, although its position has changed significantly, the tracking association unit successfully reassociates it back to the "S_007" track with 98% confidence by calculating the distance between the feature vector of its new appearance and the last feature vector of all "lost" tracks in the system, thereby avoiding tracking interruption.

[0091] 5. Intelligent Analysis and Optimization: The trajectory analysis and optimization unit continuously analyzes all active trajectories. It calculated the average velocity of "S_007" to be 165 μm / s, with a linearity of 0.92. Analysis of its high-resolution images confirmed normal head and acrosome morphology. Meanwhile, another trajectory, "S_088," had a velocity of 180 μm / s, but analysis revealed vacuoles in its head. Based on the pre-defined comprehensive scoring criteria of "high activity + excellent morphology," "S_007" received the highest score of 96.

[0092] 6. Result Presentation and Interaction: The result visualization and interaction unit highlights "S_007" in a prominent green box on the display and control module's screen. Its ID, score "96," and key parameters, including "VAP: 165, LIN: 0.92, Morph: Normal," are displayed alongside the unit. Embryologists use this clear guide to precisely and quickly manipulate the micromanipulator to capture the sperm for subsequent ICSI procedures.

[0093] Through the above-mentioned implementation methods, the present invention applies cutting-edge self-supervised artificial intelligence technology to extremely challenging clinical medical scenarios, solves many pain points of existing technologies, and demonstrates huge clinical application value and broad market prospects.

[0094] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible without inventive effort by those skilled in the art. Therefore, any technical solution that can be derived by one skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A method for high-resolution tracking of living cells using self-supervised learning of path consistency, characterized by: The method comprises the following steps: S1: Obtain unlabeled high-resolution microscopy videos of living cells as training data, and extract video slices from the videos; S2: processing each frame of the video slice using an object detector to locate and crop a bounding box image of the living cell; S3: selecting a start frame and an end frame from the video slice, determining any live cell from the start frame as a query live cell, and generating multiple observation paths by randomly skipping frames in a sequence of intermediate frames between the start frame and the end frame; S4: extracting the feature vectors of the living cells in the observation path using a deep learning embedding model, and calculating the association probability distribution from the query living cell to all living cells in the end frame; S5: Construct path consistency loss and iteratively optimize the deep learning embedding model.

2. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 1, characterized in that: The deep learning embedding model is based on the ResNet-34 architecture.

3. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 2, characterized in that: The calculation of the association probability distribution includes: aggregating feature similarities of adjacent frames through matrix multiplication, and deriving the association probability between the query living cell and all living cells in the end frame.

4. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 1, characterized in that: The random frame skipping adopts different strategies, including skipping all intermediate frames, skipping odd frames, and randomly skipping several intermediate frames.

5. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 1, characterized in that: The path consistency loss includes a consistency term and an entropy minimization term.

6. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 1, characterized in that: The optimization of the deep learning embedding model also includes regularization loss.

7. The method for high-resolution tracking of living cells using path consistency self-supervised learning according to claim 1, characterized in that: The regularization loss includes one-to-one loss to ensure matching uniqueness and bidirectional consistency loss for forward-backward tracking consistency.

8. A path consistency self-supervised learning living cell high-resolution tracking system, characterized by: It includes connecting the image acquisition module, core processing module and display and control module in sequence: The image acquisition module is used to capture high-resolution dynamic video streams of living cells; The core processing module includes a real-time detection unit, a feature embedding unit, a tracking association unit, a trajectory analysis and optimization unit, and a result visualization and interaction unit. The core processing module realizes living cell tracking by executing any tracking method according to claims 1-7; The display and control module is used to provide a human-computer interaction interface and display real-time screening results.

9. The path consistency self-supervised learning living cell high-resolution tracking system according to claim 8, characterized in that: The real-time detection unit is used to locate living cells in the video stream and crop the bounding box; The feature embedding unit is used to call the deep learning embedding model trained as claimed in any one of claims 1 to 7 to generate a living cell feature vector; The tracking association unit assigns the live cells of the current frame to an existing track, creates a new track, or terminates the track based on the distance between the feature vectors through an efficient matching algorithm to handle the re-identification of live cells after they are occluded; The trajectory analysis and optimization unit is used to calculate cell viability parameters and morphological parameters for comprehensive evaluation; The result visualization and interaction unit is used to overlay and display the living cell ID, trajectory and preferred results.

10. The path consistency self-supervised learning living cell high-resolution tracking system according to claim 9, characterized in that: The efficient matching algorithm is the Hungarian algorithm.