High-quality embryo screening method based on embryo image and cfRNA multi-modal fusion
By integrating embryo images and cfRNA information through a multimodal fusion embryo screening method, and employing deep learning and a cross-modal fusion architecture, the invasiveness and single-modality issues of existing technologies are resolved. This achieves high-precision and high-stability embryo quality assessment and improves the success rate of ART pregnancy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-27
AI Technical Summary
Existing embryo screening technologies suffer from invasive damage, incomplete information from a single modality, low efficiency of multimodal fusion methods, and significant subjective interference, making it difficult to achieve high-precision and high-stability embryo quality assessment.
A non-invasive multimodal embryo screening system is adopted, which acquires morphological and molecular biological data in parallel through an embryo image feature extraction module and a molecular detection module. The feature data is integrated by a deep learning multimodal fusion analysis module. Combined with a cross-modal fusion architecture and attention mechanism, a comprehensive and objective assessment of embryo quality is achieved.
This approach enables a deep and integrated assessment of embryo quality, improves the success rate of ART clinical pregnancy, overcomes the limitations of subjectivity and single-modality in traditional methods, and enhances the accuracy and stability of the assessment.
Smart Images

Figure CN121747709A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of biomedical engineering, and particularly relates to a high-quality embryo screening method based on embryo image and cfRNA multi-modal fusion, and particularly relates to an embryo quality screening method by fusing embryo morphological image and cfRNA data in culture medium through artificial intelligence technology, so as to realize non-invasive and precise quality assessment and screening of in vitro fertilization embryos. BACKGROUND
[0002] At present, about 186 million people worldwide are facing fertility problems, and ART, as an important means to help infertile couples achieve pregnancy, its clinical pregnancy success rate is still only 30%-50%. Improving the success rate of ART has become a key problem to be solved in this field, and the screening of high-quality embryos is particularly important. At present, the traditional embryo morphological scoring is a commonly used screening method, which is based on morphological indicators such as blastocyst expansion, inner cell mass and trophoblast ectoderm for evaluation. However, this method is highly subjective, relies only on static observation at a specific time point, and is difficult to accurately reflect the dynamic development potential of the embryo, and also cannot effectively identify embryos with normal morphology but with genetic or epigenetic defects, so the prediction accuracy is limited.
[0003] The embryo time-lapse imaging technology (TLT) is introduced, which realizes continuous observation of the whole process of embryo development, and can identify abnormal kinetic parameters related to chromosomal abnormalities. However, embryo screening under TLT technology still highly depends on manual analysis, which is not only time-consuming and labor-intensive, but also easily interfered by human factors. On the other hand, preimplantation genetic testing for aneuploidy (PGT-A) as an invasive method, although it can directly assess the chromosomal state to reduce the risk of implantation, but the biopsy operation itself may cause damage to the embryo. At the same time, due to the widespread existence of embryo chimerism and self-repair mechanism, the diagnostic accuracy of PGT-A also faces challenges, which may lead to misjudgment.
[0004] In recent years, artificial intelligence (AI) technology has been applied to the field of embryo evaluation. Through deep learning algorithms, the morphological features and morphokinetic features of embryos are automatically analyzed to classify the quality and developmental stage of embryos, which to some extent reduces the interference of human factors and improves the objectivity and efficiency of evaluation. However, existing AI models still have technical shortcomings: they can only analyze based on single modal image information, are difficult to capture subtle morphological differences and abnormal information at the genetic level, and cannot achieve deep characterization of embryo quality. At the same time, non-invasive detection technology based on free RNA (cfRNA) in embryo culture medium has gradually emerged. This technology detects cfRNA secreted by embryos into the culture medium, indirectly reflecting the molecular state, gene expression profile, and developmental potential of embryos, providing a new perspective in the field of molecular biology for embryo evaluation. However, this technology can only provide single molecular modal information and cannot be combined with morphological features for comprehensive evaluation, making it difficult to fully and accurately judge embryo quality.
[0005] Although some studies have attempted to combine multi-modal information (such as morphology and metabolomics), there are still deficiencies in the field of embryo screening in terms of deep fusion of high-dimensional dynamic image sequences and ultra-low cfRNA transcriptome data. Existing solutions mostly stop at feature splicing or early fusion, making it difficult to address the semantic gap and dynamic contribution allocation problem. Specifically, existing fusion methods mostly use simple feature splicing or early fusion strategies, failing to address the semantic gap between image spatial features and gene expression vectors, and lacking a dynamic and interpretable weighting mechanism for the contribution of the two modalities, resulting in limited performance improvement and insufficient generalization ability of the fusion model. Currently, there is no deep fusion architecture in the public literature and patents that significantly improves the prediction performance and stability of ultra-low cfRNA transcriptome and high-dimensional dynamic embryo image sequences, making it difficult to meet the demand for high precision and high stability in the clinical setting.
[0006] In summary, existing embryo screening technologies have obvious limitations: traditional morphological scoring methods are highly subjective and provide single information; TLT technology relies on manual analysis and is inefficient; PGT-A technology is invasive and its diagnostic accuracy is disturbed; existing AI models and cfRNA detection technologies are limited by single modal information and cannot achieve fusion evaluation of embryo morphological features and molecular biology information. Therefore, there is an urgent need in the current assisted reproduction field to develop a non-invasive embryo screening system that can deeply fuse multi-dimensional information of embryos (especially dynamic morphology and cfRNA molecular features) and has high precision and stability, in order to break through the existing technical bottlenecks and improve the clinical pregnancy success rate of ART. SUMMARY
[0007] Purpose of the Invention: The purpose of this invention is to overcome the shortcomings of existing embryo screening technologies, such as invasive damage, incomplete information from a single modality, low efficiency of existing multimodal fusion methods, and significant subjective interference, and to provide a non-invasive, deeply integrated embryo screening method. This method integrates the temporal morphological characteristics of the embryo with the molecular biological information contained in the culture medium's cfRNA through an innovative cross-modal fusion architecture, achieving a comprehensive, objective, and accurate assessment of embryo quality. This improves the clinical pregnancy success rate of ART and provides a safer and more reliable embryo screening solution for individuals with fertility problems.
[0008] Technical Solution: The non-invasive multimodal embryo screening system provided by this invention mainly acquires morphological and molecular biological data of embryos in parallel through an embryo image feature extraction module and a molecular detection module; subsequently, a deep learning multimodal fusion analysis module integrates and analyzes the above feature data, and finally outputs reliable embryo quality assessment results, including the following steps:
[0009] (1) Collect embryo morphology images, time-difference imaging videos, and corresponding embryo culture medium samples from ART patients, and simultaneously record clinical information such as patient age, health status, and pregnancy outcome after embryo transfer;
[0010] (2) Preprocess the embryo morphological images and time-difference imaging videos obtained in step (1) to obtain standardized image input data;
[0011] (3) Perform cfRNA sequencing library construction and bioinformatics analysis on the embryo culture medium samples obtained in step (1) to obtain embryo molecular biological information;
[0012] (4) Input the embryo image features obtained in step (2) and the cfRNA features obtained in step (3) into a dual-branch deep network, and generate a fusion feature vector through a residual network containing a multi-attention module and a cross-modal fusion unit with an expert hybrid mechanism.
[0013] (5) Input the fusion features obtained in step (4) into the predictor. The trained deep model can output the embryo availability assessment result or the pregnancy success probability.
[0014] The specific sub-steps of step (2) include:
[0015] Screening for high-quality static images: Embryo images were collected on the fifth or sixth day after fertilization, before any intervention. Screening criteria included: adequate lighting and clear structure; clear boundary between the zona pellucida and trophoblast; the image contained only one complete embryo, with no instrument parts obstructing the view and very few fragments; and no text symbols obscuring key embryonic structures within the field of view.
[0016] Annotation and dataset construction of development video: The embryo development video is exported as AVI format, and the embryo is evaluated and annotated by experienced physicians combined with clinical information, which is divided into high quality and low quality two groups as the dataset for subsequent deep learning model training and verification;
[0017] Extraction of embryo region in image: In order to accurately obtain the embryo region, Hough circle detection algorithm is used to crop the image. This algorithm identifies circular targets through gradient distribution analysis, and its parameters follow the formula (x - a)² + (y - b)² = r², where (a, b) is the center coordinate and r is the radius, so as to realize the automatic positioning and extraction of ROI.
[0018] The specific sub-steps of step (3) include:
[0019] cfRNA sequencing library preparation is performed on the collected embryo culture medium samples, and transcriptome sequencing is completed;
[0020] Filter low-quality reads and repetitive sequences in sequencing data, and align high-quality sequences with human reference genome;
[0021] Extract valid reads from the generated BAM file, analyze the fragment length distribution, end motif pattern and variable splicing events of cfRNA;
[0022] Use Mann-Whitney U test to analyze the differential expression genes of cfRNA in different quality embryo culture media, and correct FDR by Benjamini-Hochberg method, and screen FDR < 0.01 and |Fold Change| > 2 genes as significantly differentially expressed genes;
[0023] Use Lasso regression algorithm to screen feature gene combination from differential genes to predict embryo quality, determine the optimal regularization parameter λ by ten-fold cross-validation, and use glmnet package to fit the model to reduce variable redundancy and establish a high-quality embryo prediction feature set.
[0024] The specific sub-steps of step (4) include:
[0025] Divide the preprocessed embryo image features and cfRNA feature data into training set and test set according to the ratio of 80%:20%, and ensure that the model is only trained and verified on the training set;
[0026] A double-branch feature extraction network based on convolution attention block is designed, and a hybrid expert system is introduced to dynamically fuse double-path information from images and genes. The fused features are output by the prediction module to output a binary classification result ("0" represents unusable embryo, and "1" represents usable embryo).
[0027] The model learns the deep features of the morphology and the genes through an end-to-end learning mode, and generates a prediction confidence for each embryo. Based on the prediction result and the true pregnancy outcome, a loss value is calculated, and the network weights are updated through back propagation using the Adam optimizer (learning rate set to 0.0001). After multiple iterations, the loss is continuously reduced to optimize the performance of the model.
[0028] The model weights after training are saved, and the prediction is performed on the test set that has not been seen before. The prediction result is compared with the true embryo availability label, and comprehensive evaluation of the model performance is performed by using indicators such as Accuracy, F1-score, Precision, Sensitivity, Specificity, and AUC.
[0029] The specific sub-steps of the step (5) include:
[0030] 5-fold stratified cross-validation is used to minimize the bias caused by random sampling of the data set;
[0031] In 5-fold cross-validation, the data set is randomly divided into 5 sub-samples of the same size and balanced in terms of the number of embryos of each class. Five independent embryo quality prediction models are trained from scratch using 4 sub-samples, and the remaining 5th sub-sample is used for validation.
[0032] Preferably, the static image acquisition satisfies: uniform light source, clear boundary between zona pellucida and trophoblast, single embryo presentation, no occlusion, no fragments, and is taken at 5-6 days after fertilization;
[0033] Preferably, the Hough circle detection algorithm is used to automatically extract the ROI region of the embryo, and the gradient distribution is used to locate the center and radius of the embryo, which satisfies: (x−a) 2 +(y−b) 2 =r 2 .
[0034] Preferably, the cfRNA sequencing library is sequenced on a sequencing platform after random primer reverse transcription, ribosome cDNA removal, and PCR amplification, and the data is filtered, aligned, de-duplicated, and fragment feature analyzed.
[0035] Preferably, the differential abundance analysis of cfRNA uses Mann-Whitney U test and FDR correction.
[0036] Preferably, Lasso regression is used to screen cfRNA feature genes that can predict embryo quality, and the optimal regularization parameter λ is determined through ten-fold cross-validation.
[0037] Preferably, the image branch and the cfRNA branch construct convolution feature extraction networks respectively, outputting deep feature vectors ϕ(X b ;W b ) and ϕ(X rna ; W rna ), and the image branch can be optionally a residual convolution network structure.
[0038] Preferably, the multi-modal feature extraction module fuses channel attention and spatial attention mechanisms: channel attention generates feature weights by combining maximum pooling and average pooling operations, and spatial attention focuses on the distribution of key regions in the image, and on this basis, the module further introduces a MoE expert mixing mechanism to dynamically and adaptively weight and fuse features of different modalities.
[0039] Preferably, the model is trained using an Adam optimizer with a learning rate of 0.0001, and the training set and test set are divided at a ratio of 80:20.
[0040] Preferably, the performance of the model is evaluated by Accuracy, F1-score, Precision, Sensitivity, Specificity, and the area under the receiver operating characteristic curve AUC, and is clinically verified by five-fold stratified cross-validation.
[0041] Compared with existing methods that rely only on morphological images or single cfRNA molecular signals for embryo quality evaluation, the present application has achieved significant breakthroughs in method system, information acquisition depth, cross-modal fusion capability, and model stability, and the specific technical advantages are as follows:
[0042] (1) Realize the deep fusion of "embryo morphology-molecular state" double-level cross-modal, solve the inherent ambiguity and uncertainty problem of single-modal method. Traditional morphological methods cannot identify embryos with normal morphology but abnormal transcriptional state; existing cfRNA methods lack developmental kinetics information. The present application deeply fuses the two complementary modalities, so that the morphological information complements the low spatial resolution of molecular detection, and the cfRNA molecular characteristics compensate for the recognition blind area of image models for "abnormal gene expression embryos";
[0043] (2) Introduce cross-modal attention mechanism to realize dynamic alignment and focusing of image spatial information and cfRNA feature vectors, and improve model reliability. Traditional multi-modal methods mostly use simple splicing fusion and cannot judge the correlation between morphological regions and molecular features. The present application realizes dynamic enhancement of key morphological regions and expression mode weighting of cfRNA features through attention mechanism, automatically suppresses weakly correlated features, and significantly improves the feature expression ability and decision reliability of the model;
[0044] (3) The adaptive weighted fusion of cross-modal information is realized by using a mixed expert model, and the model stability and generalization ability are improved. In view of the heterogeneity of different embryonic development stages and gene expression patterns, the model has a personalized fusion strategy for different samples through the MoE structure. The gating network automatically selects the best expert combination according to the input, effectively avoiding the overfitting problem of traditional models on heterogeneous data;
[0045] (4) The cfRNA feature screening system has multi-layer refining ability, and the feature is more indicative. Instead of simply using differentially expressed genes, the invention fuses fragment length distribution, terminal motif, and variable splicing event analysis, and uses Lasso regression for regularization screening, finally obtains a set of key cfRNA feature genes with refined quantity but strong embryonic development directionality, greatly improving the information quality of the molecular mode;
[0046] (5) It is completely non-invasive, only needs conventional blastocyst images and trace amount of culture solution, does not need biopsy, has high clinical applicability and safety, and meets the ethical requirements.
[0047] Advantages: Compared with the prior art, the present application has the following obvious advantages:
[0048] (1) The static morphological image and the time difference imaging dynamic sequence are comprehensively utilized, the shortcomings that the traditional morphological score only focuses on the static characteristics of a specific time point are overcome, and the dynamic process characteristics of embryonic development can be captured;
[0049] (2) Through cfRNA sequencing combined with multi-layer bioinformatics analysis, key information at the molecular level that cannot be detected by morphology can be mined;
[0050] (3) By introducing the attention mechanism and the expert mixed model (MoE) fusion architecture, the deep complementarity and adaptive fusion of morphology and molecular modalities are realized. Among them, the attention mechanism aligns the semantic gap between image space features and gene expression vectors, and MoE dynamically allocates fusion weights according to the heterogeneity of embryos, and the two cooperate to produce nonlinear gain, rather than simple superposition, so as to effectively distinguish complex situations such as normal morphology but abnormal genes, or normal genes but slow development, and significantly improve the accuracy and robustness of evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 The figure is a flow structure diagram of the present application.
[0052] Figure 2 The figure is a model structure diagram of the present application. A double-branch network is constructed based on a residual network, and CAM_black is introduced based on the residual block to make the network model pay attention to the main information of each modality, and MoE network is used to dynamically fuse the double-branch features to realize information complementation.
[0053] Figure 3 This is a schematic diagram of the CAM of the present invention, which uses newly constructed channel attention and spatial attention modules to capture the network's areas of interest in multimodal data at different levels.
[0054] Figure 4 The schematic diagram of the MoE structure shows that multiple expert networks and gating networks are constructed using conventional convolutions, and multimodal deep features are dynamically fused through the gating mechanism. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of protection of this invention.
[0056] Current embryo quality assessment methods have the following limitations: First, most methods focus on static images at a single time point during the blastocyst stage, failing to fully capture the dynamic temporal characteristics of the entire embryonic development process. Second, they typically only assess local structures such as the blastocoel, inner cell mass, or trophoblast independently, making it difficult to systematically integrate the overall morphological information of the embryo. Third, most existing technologies do not incorporate molecular biological data from the embryo culture medium, failing to reflect the embryonic developmental potential at the gene expression level. These limitations restrict the accuracy and reliability of current methods in predicting embryo quality.
[0057] See Figure 1 The diagram shown illustrates a multimodal fusion embryo quality assessment method based on embryo images and cell-free RNA information in culture medium, according to an embodiment of the present invention. The method includes the following steps:
[0058] 1. Embryo Image Acquisition and Preprocessing
[0059] (1) Embryo morphological image acquisition standards: Morphological image acquisition must meet the following quality control requirements: high-quality filters are used to ensure sufficient light and clear identification of embryonic structures; the boundaries between the zona pellucida and the trophoblast are clear; each microscopic image contains only a single embryo, with no interference from instrument components in the field of view, and very few or no embryonic fragments; the embryo is presented completely within the image area, with no text or labels obscuring key morphological structures. All images are acquired on the fifth or sixth day after fertilization and are taken before interventional procedures such as embryo biopsy or transfer;
[0060] (2) Embryo time-lapse imaging video processing: The collected embryo development time-lapse imaging videos were uniformly exported to standard AVI format using professional embryo imaging analysis software ESCOMIRITL Viewer to ensure consistency in video encoding and resolution. Embryo quality assessment and grouping were independently completed by at least two experienced senior embryologists. The assessment strictly followed the established clinical morphological scoring standard, the Gardner scoring system, and the key dynamic morphological indicators during embryo development were observed and recorded, including pronucleus appearance and disappearance time, blastomere number, size uniformity, fragment proportion, blastocyst formation time, inner cell mass, and trophoblast rating. After independent assessment, the embryologists would reach a consensus through discussion on cases with differences, and finally the embryos were clearly divided into high-quality embryo group and low-quality embryo group. All AVI format videos were subjected to uniform pretreatment procedures before inputting into the model, including frame rate standardization, image size normalization, and brightness and contrast calibration to eliminate technical variations.
[0061] 2 Embryo image feature extraction
[0062] Hough circle detection algorithm was used to crop embryo images to achieve ROI extraction. Hough circle detection algorithm is mainly based on Hough transform, which uses gradient distribution map to analyze the image to detect the position and parameters of circular objects. The calculation formula is represented as: (x - a) 2 + (y -b) 2 = r², where a and b represent the center coordinates, and r represents the radius of the circle.
[0063] 3 Embryo culture medium cfRNA sequencing library preparation
[0064] (1) Library construction and sequencing: Take 4 μL of embryo culture medium, use random primers for reverse transcription and PCR amplification; then use ZapR and R-Probes to remove ribosome cDNA, use 18 cycles for the second round of PCR to form a cDNA library. The constructed cDNA library was detected by Agilent@2100 bioanalyzer, including library fragment size, purity and concentration. The qualified library was sequenced on the Illumina HiSeq X10 PE150 platform.
[0065] (2) Bioinformatics analysis of embryo culture medium cfRNA sequencing data
[0066]
[0067] cfRNA fragment feature analysis: Extract reads aligned to the reference genome from the aligned bam files, and calculate the cfRNA fragment length distribution;
[0068] Variable splicing analysis: First, Salmon (version 1.5.2) was used to quantify the transcripts. Subsequently, based on the quantitative results, SUPPA2 (version 2.3) was used to systematically identify variable splicing events, including exon skipping, intron retention, variable donor sites, variable acceptor sites, variable promoters, variable terminators, and mutually exclusive exons. Finally, the diffSplice command in SUPPA2 was used to analyze the differential splicing events between different samples;
[0069] Differential expression analysis and functional analysis: Mann-Whitney U test method was used to analyze the different quality embryo culture fluid cfRNA differential abundance genes, and the Benjamini-Hochberg adjusted FDR value was used, and the FDR <0.01 and |Fold Change| > 2 were retained as genes with differential abundance. Subsequently, DAVID (https: / / david.ncifcrf.gov) and Metascape were used to perform KEGG and GO functional annotation on the differential abundance genes, so as to reveal the molecular processes of different quality embryos.
[0070] ⑤ cfRNA feature screening: To construct the prediction model of embryo quality, the Lasso regression algorithm is used to screen the cfRNA molecular markers with the highest prediction potential. The specific process is as follows: First, based on the embryo quality label, the Lasso regression is used to select the features of the whole cfRNA expression data; the optimal regularization parameter λ is determined through ten-fold cross-validation to balance the model complexity and prediction performance; then the R language glmnet package is used to fit the model, and the coefficients of irrelevant or redundant variables are automatically compressed to zero, so as to obtain a set of key cfRNA feature genes (see Table 1 for example). The feature genes screened by this method can be dynamically adjusted according to the actual data, and are not limited to the genes listed in Table 1, and can also be applied to other feature combinations screened by this process.
[0071] Table 1 Example gene set
[0072] 4 Multi-modal fusion model construction
[0073] Based on the residual network, a special cross-modal interaction module Res_CAM_Black is constructed, and a double-branch feature extraction network is constructed in turn to learn the features of multi-modal data. The model structure diagram is as shown in Figure 2 .
[0074] The input data of the embryo image branch is three-channel RGB data, which is extracted by the network weight to obtain the deep feature , is the network mapping function. In order to improve the reliability of image information evaluation, the gene expression characteristics of cfRNA sequencing analysis are introduced, and the deep information is extracted through a specific network branch to obtain the deep feature . Finally, the double-branch features are dynamically fused by the hybrid expert network, and the final result is predicted by the prediction layer.
[0075] In order to obtain high-quality representation information in the above network structure, we construct an efficient attention mechanism, such as Figure 3 , and the specific data processing process is as follows:
[0076]
[0077] where X represents the input feature data, Sigmoid is the activation function, are the max-pooling and mean-pooling respectively.
[0078] Compared with traditional feature fusion methods (such as feature concatenation or direct summation), the MoE (Mixture of Experts) framework can adaptively learn and fuse complex data feature distributions at multiple scales by introducing a multi-expert network structure. The core of this architecture is to dynamically adjust the weights of each expert network through a gating mechanism, so that multiple experts can work together to complete complex tasks. As shown in Figure 4 , the MoE model can adaptively divide the input feature space into multiple subspaces and match the most suitable expert network for each subspace, effectively supporting multi-level and multi-scale feature fusion and representation learning.
[0079] In this study, the feature processing flow of MoE can be represented as:
[0080]
[0081] where is the weight assigned to the expert network by the gating network, satisfying the following constraint condition:
[0082]
[0083] The gating network is the core component of the MoE model, mainly responsible for dynamically assigning weights to different experts based on input features, achieving adaptive task allocation and expert responsibility optimization. In specific implementation, the number of experts N (usually 2-8, preferably 4) can be set, and a 1-3 layer fully connected network can be used to construct the gating structure. The input features are subjected to a trainable linear transformation by the gating network to generate initial scores for each expert, which are then normalized by the Softmax function to obtain the weight distribution of each expert. The weights can be applied to specific layers of the dual-branch structure (such as post-fusion at the kth layer of feature extraction or supporting multi-layer fusion mechanism) to flexibly regulate the integration of expert outputs. The calculation process is as follows:
[0084]
[0085] is the weight of the i-th expert network, and N is the number of expert networks. This mechanism can adjust the weights of each expert in real time according to the input features, ensuring that each sample is processed by the most suitable expert, thereby improving the prediction accuracy and computational efficiency of the model. During training, the gating network continuously optimizes the task allocation strategy and dynamically adjusts the responsibility range of each expert, so that each expert can focus on its specialized data subspace, achieving more efficient feature modeling. This adaptive multi-scale feature fusion mechanism significantly enhances the model's generalization ability among complex data patterns, providing flexible and robust architectural support for multi-level feature representation learning.
[0086] 5 Model performance verification and experimental results
[0087] The pre-processed dataset (sample size 200) was divided into training and test sets in the ratio of 80% and 20% successively, and the model was trained only on the training set. The embryo images and cfRNA features were input into the deep learning model, which learned the image and cfRNA features and integrated and mapped the different levels of features to the corresponding labels (“0” representing an unusable embryo and “1” representing a usable embryo), and the performance results are shown in Table 2.
[0088] Table 2 Performance comparison of multi-modal model and single-modal model
[0089]
[0090] The results show that the multi-modal fusion comprehensively improves the prediction performance. In addition, five-fold stratified cross-validation is used to evaluate the stability of the model. The variance (σ²) of the performance of the multi-modal model is 0.007, which is significantly lower than that of the single-image model (0.012) and the single-cfRNA model (0.019), indicating that the model has stronger robustness and generalization ability.
[0091] 6 Ablation experiment analysis
[0092] To further verify the necessity and technical contribution of the attention alignment mechanism and the MoE gate dynamic fusion in the present application to the overall performance improvement, an ablation experiment is designed for comparison. As shown in Table 3, not all feature fusion can bring performance improvement, and removing any core component will cause a significant decrease in model performance.
[0093] Table 3 Performance comparison of ablation experiment
[0094]
[0095] The results show that the performance of the model is significantly reduced after removing the attention alignment mechanism or the MoE. In summary, the dynamic allocation of the attention alignment and the MoE gate is a synergistic effect and indispensable, and the performance gain brought by their combination cannot be easily achieved by conventional feature splicing by those skilled in the art.
[0096] 7 Clinical verification
[0097] Using a five-fold cross-validation strategy, the model is trained on the training set and the performance of predicting embryo usability is evaluated on the validation fold. The model exhibits stable high prediction ability (average AUC > 0.90), proving its potential for clinical application.
[0098] The above merely describes preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Any change or equivalent replacement within the disclosed technical principles of the present application, which can be easily conceived by those skilled in the art, shall fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be defined by the appended claims.
Claims
1. A method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data, characterized in that, Includes the following steps: (1) Image and sample acquisition: Collect morphological static images of embryos, time-difference imaging video sequences, and corresponding embryo culture medium samples from assisted reproductive patients, and record the patients' clinical information and embryo transfer outcomes; (2) Image preprocessing: Quality screening, standardization and cropping of static images and video frames, and automatic extraction of the region of interest (ROI) of the embryo using an algorithm based on Hough circle detection; (3) Acquisition of cfRNA molecular characteristics: The culture medium samples obtained in step (1) were used to construct an ultra-micro cfRNA sequencing library and perform high-throughput sequencing; then the sequencing data were subjected to sequence alignment, fragment feature analysis, alternative splicing analysis and differential expression analysis, and finally a set of cfRNA characteristic genes that can characterize embryo quality were screened. (4) Multimodal feature construction: The embryo image features obtained in step (2) and the cfRNA features obtained in step (3) are respectively input into the dual-branch deep network, and the feature focusing module containing channel attention and spatial attention is used to improve the expressive power of the features. The image features and cfRNA features are aligned and a fused feature vector is generated based on the dynamic fusion mechanism of hybrid expert MoE. The hybrid expert MoE mechanism includes a gating network for dynamically weighting the contributions of different modalities. (5) Embryo quality prediction: The fusion features are input into the predictor, and the trained deep model outputs the embryo availability or pregnancy probability.
2. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The static image acquisition meets the following requirements: uniform light source, clear boundary between the zona pellucida and trophoblast, single embryo presentation, no obstruction, no fragmentation, and is taken on the 5th–6th day after fertilization.
3. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The Hough circle detection algorithm is used to automatically extract the ROI region of the embryo, and locate the center and radius of the embryo based on gradient distribution, satisfying: (x−a) 2 +(y−b) 2 =r 2 .
4. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The cfRNA sequencing library was reverse transcribed using random primers, ribosomal cDNA was removed, and PCR amplification was performed before sequencing on a sequencing platform. The data was then filtered, aligned, deduplicated, and analyzed for fragment characteristics.
5. A method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The differential abundance analysis of cfRNA was performed using the Mann-Whitney U test and FDR correction.
6. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The screening of cfRNA characteristic genes that can predict embryo quality was performed using Lasso regression, and the optimal regularization parameter λ was determined by 10-fold cross-validation.
7. A method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The image branch and the cfRNA branch respectively construct convolutional feature extraction networks, outputting a depth feature vector ϕ(X). b W b ) and ϕ(X rna W rna The image branch can be selected as a residual convolutional network structure.
8. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The multimodal feature extraction module integrates channel attention and spatial attention mechanisms: channel attention generates feature weights by combining max pooling and average pooling operations, while spatial attention is used to focus on the distribution of key regions in the image. On this basis, the module further introduces the MoE expert hybrid mechanism to dynamically and adaptively weight and fuse features from different modalities.
9. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The model was trained using the Adam optimizer with a learning rate of 0.0001, and the training and test sets were divided at an 80%:20% ratio.
10. The method for screening high-quality embryos based on the fusion of embryo images and cfRNA multimodal data according to claim 1, characterized in that, The model performance was evaluated using accuracy, F1 score, precision, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC), and clinical validation was performed using 5-fold stratified cross-validation.