Embryo optimization method and system with multi-modal temporal fusion and anatomical prior constraints

CN122842974APending Publication Date: 2026-09-29XUZHOU MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611308075.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,这些现有技术方案仍存在显著不足

Benefits of technology

[0027]本发明通过解剖结构约束的时空注意力与深度时序融合显著提升了胚胎发育潜能的预测精度。在内部回顾性验证数据集上进行五折交叉验证,并经外部独立测试集评估,本发明的优质胚胎选择AUC值达到0.953,较传统静态图像模型的0.832提升了12.1个百分点;同时,模型准确率从79.1%提升至90.2%,特异性从81.3%提升至91.8%,灵敏度从76.5%提升至88.4%,模型稳定性(十次运行AUC方差)从1.44e-4降低至0.25e-4,各项指标均得到显著提升,充分佐证了本发明的性能优势。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842974A_ABST
    Figure CN122842974A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for embryo selection based on multimodal temporal fusion and anatomical prior constraints, specifically relating to the interdisciplinary fields of assisted reproduction and artificial intelligence. The method involves collecting temporal images of embryos and maternal clinical data as input to a 3D convolutional network; using U-net++ to segment the images and generate anatomical prior masks, providing prior knowledge of local details and global structure; integrating the masks into the 3D convolutional network, coupled with a spatiotemporal attention module, to dynamically focus on developmental changes in key embryonic regions; then constructing a multi-scale perception network to extract features at different levels and perform fusion through adaptive gating; subsequently aligning and bidirectionally encoding the clinical data; combining a causal discovery algorithm to screen causal features affecting pregnancy outcomes; and relying on a causal enhancement model and a developmental trajectory model to obtain embryo potential scores, feature contributions, and counterfactual results. The system is then deployed to achieve developmental abnormality identification, pregnancy probability prediction, potential ranking, and transplantation recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of assisted reproduction and artificial intelligence, specifically to a method and system for embryo selection based on multimodal temporal fusion and anatomical prior constraints. Background Technology

[0002] Traditional embryo selection methods mainly rely on embryologists to assess embryo quality by scoring morphology under a microscope at specific time points. This method has inherent limitations such as large subjective differences among assessors, inconsistent standards, and inability to quantify dynamic developmental information.

[0003] Chinese invention patent CN120747631A proposes introducing a residual-aware attention module into embryo image classification to enhance the network's attention to important morphological features; invention application CN119090820A discloses a scheme for multi-scale feature extraction in frozen embryo prediction by combining a feature pyramid network. In addition, some studies have utilized time-lapse photography to extract embryonic development features by separating temporal and spatial domain models and then fusing them later.

[0004] However, these existing technical solutions still have significant shortcomings. First, their attention mechanisms are mostly of general design, lacking targeted guidance for key embryonic anatomical structures. The models may be affected by background artifacts or irrelevant objects in the culture dish, and the decision-making logic may deviate from the professional understanding of embryologists, resulting in weak interpretability. Second, in terms of utilizing multimodal information, existing methods mostly use simple early feature splicing or later decision weighting, failing to deeply integrate the profound correlation between embryonic temporal morphological images and dynamic clinical data of patients, and the characterization of the continuous dynamic process of embryonic development is not refined enough. More fundamentally, existing models rely entirely on data-driven correlations for prediction, unable to distinguish between causal effects and spurious associations between features and pregnancy outcomes. The models are prone to learning data biases, have limited generalization ability, and cannot answer the causal questions of clinical concern.

[0005] Therefore, there is an urgent need in this field for a next-generation intelligent embryo selection method and application system that can simulate the evaluation logic of embryologists, deeply integrate multimodal temporal information, and provide explainable decision-making basis from a causal perspective. Summary of the Invention

[0006] To address these issues, the present invention provides a method and system for embryo selection based on multimodal temporal fusion and anatomical prior constraints, thereby resolving the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for embryo selection based on multimodal temporal fusion and anatomical prior constraints, comprising the following steps:

[0008] S1: Collect time-series image sequences of embryonic development and corresponding maternal multidimensional clinical data, construct a time-modal feature cube indexed by samples, obtain the input feature map of a single image, and use it as the input of the subsequent three-dimensional convolutional neural network. Here, "three-dimensional" does not refer to medical three-dimensional volume data, but rather to the combination of time dimension and two-dimensional image space.

[0009] S2: Apply U-net++ to segment embryo images in time-series image sequences into anatomical structures and generate anatomical prior masks, including key anatomical structure masks and foreground masks. This establishes anatomical prior constraints for focusing on depth feature extraction at different scales, that is, it provides local region prior knowledge for cell-level high-resolution feature extraction branches and global structural relationship prior knowledge for embryo-level context feature extraction branches.

[0010] S3: Introduce the anatomical prior mask obtained in step S2 into the three-dimensional convolutional neural network to obtain the improved three-dimensional convolutional neural network; construct a spatiotemporal attention module composed of channel attention submodule and spatial attention submodule based on anatomical structure constraints, and use the anatomical prior mask to guide the channel weight distribution and constrain the spatial attention weight respectively, so that the improved three-dimensional convolutional neural network can dynamically focus on key anatomical regions and their developmental morphological changes at continuous development time points when extracting temporal features;

[0011] S4: Design a multi-scale perception network to extract high-resolution cell-level features and embryo-level contextual features of the embryo through parallel paths, and introduce an adaptive gating fusion mechanism to dynamically weight and fuse the multi-scale features.

[0012] S5: Align dynamic clinical data into a clinical time series according to the collection time, and use a two-way... Temporal encoding was performed to obtain clinical temporal representations; a feature selection algorithm based on causal discovery was used to screen out a subset of causal features with direct causal effects on pregnancy outcomes from the fused multi-scale embryonic image temporal features and maternal clinical temporal features.

[0013] S6: Based on the selected subset of causal features, the decision forest model with causal enhancement is combined with the embryo development trajectory prediction model based on multimodal temporal fusion data to calculate inference and output the embryo development potential score, the causal contribution of each feature and the counterfactual inference results.

[0014] S7: Deploy an intelligent decision support system to receive time-series image sequences and corresponding new samples of maternal multidimensional clinical data, and output key developmental abnormality point identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

[0015] This invention also discloses an embryo selection system based on multimodal temporal fusion and anatomical prior constraints, characterized in that the system performs the above-described embryo selection method based on multimodal temporal fusion and anatomical prior constraints, including:

[0016] The multi-source time-series data acquisition module is used to simultaneously acquire embryo time-series image sequences and corresponding maternal multidimensional clinical data during the embryo culture process;

[0017] The anatomical structure-guided multi-scale feature extraction module includes an embryonic anatomical structure segmentation unit, a spatiotemporal attention module, and a multi-scale perception network, as detailed below:

[0018] The embryonic anatomical structure segmentation unit achieves pixel-level segmentation of key anatomical structures based on the U-Net++ network, obtaining anatomical prior masks, namely embryonic key anatomical structure masks and embryonic foreground masks; anatomical prior constraints are established for focusing on the extraction of depth features at different scales, that is, providing prior knowledge of local region details for cell-level high-resolution feature extraction and prior knowledge of global structural relationships for embryonic-level contextual feature extraction.

[0019] The spatiotemporal attention module, whose spatial attention submodule integrates anatomical prior masks, guides an improved 3D convolutional neural network based on anatomical constraints to dynamically focus on the developmental changes of key anatomical regions when extracting temporal features.

[0020] The multi-scale perception network extracts high-resolution cell-level features and embryo-level contextual features of the embryo through parallel paths, and dynamically weights and fuses the multi-scale features through an adaptive gating mechanism.

[0021] The cross-modal causal feature selection module includes a temporal feature encoding unit and a cross-modal causal feature selection unit, as detailed below:

[0022] The temporal feature encoding unit aligns dynamic clinical data into a clinical time series according to the acquisition time, and uses bidirectional LSTM for temporal encoding to obtain a clinical temporal representation;

[0023] The cross-modal causal feature selection unit uses a feature selection algorithm based on causal discovery to select a subset of core features that have a direct causal effect on pregnancy outcomes from the fused multi-scale temporal features and clinical features.

[0024] The causal-enhanced intelligent decision-making module, based on the selected subset of causal features, uses a gradient-enhanced decision forest model and combines it with an embryo development trajectory prediction model trained based on multimodal time-series fusion data to calculate inference, outputting embryo development potential score, causal contribution of each feature, and counterfactual inference results.

[0025] An interpretable clinical interaction platform provides a visual interface for dynamic heatmaps of embryonic development, causal contribution waterfall plots, time-series abnormality reports, and personalized transplantation recommendations. This enables the platform to receive time-series image sequences and corresponding new samples of multidimensional maternal clinical data, and output key developmental abnormality identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

[0026] The present invention has the following advantages:

[0027] This invention significantly improves the prediction accuracy of embryonic developmental potential through spatiotemporal attention constrained by anatomical structure and deep temporal fusion. Five-fold cross-validation was performed on an internal retrospective validation dataset, and the AUC value for selecting high-quality embryos using this invention reached 0.953, a 12.1 percentage point improvement compared to the 0.832 of the traditional static image model. Simultaneously, the model accuracy increased from 79.1% to 90.2%, specificity from 81.3% to 91.8%, and sensitivity from 76.5% to 88.4%. Furthermore, the model stability (variance of AUC over ten runs) decreased from 1.44e-4 to 0.25e-4. These significant improvements across all metrics fully demonstrate the performance advantages of this invention.

[0028] This invention, based on a causal reinforcement-based intelligent decision-making module, utilizes causal contribution analysis and counterfactual reasoning capabilities to ensure that every recommendation decision is verifiable and the decision-making process is highly transparent and credible. This greatly enhances clinicians' trust in AI and represents a leap from a "black box" to a "white box."

[0029] This invention enables the system to identify abnormal patterns and predict developmental trajectories in early development to assess embryonic developmental potential through multimodal data temporal fusion modeling. Counterfactual reasoning can provide quantitative intervention suggestions for specific patients, transforming the system from a passive assessment tool into an active clinical decision support engine.

[0030] This invention effectively identifies and filters false associations in data through a feature selection algorithm based on causal discovery, improving the model's generalization ability across different populations and application scenarios. Its modular design facilitates system integration, deployment, and continuous iterative upgrades. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of a superior embryo intelligent selection system based on multimodal information fusion.

[0032] Figure 2 This is a flowchart of the multimodal data processing and feature extraction stages.

[0033] Figure 3 This is a flowchart of the multimodal feature adaptive gating fusion mechanism.

[0034] Figure 4 This is a flowchart of the feature selection and causal reinforcement process based on causal discovery.

[0035] Figure 5 This is a flowchart of the intelligent decision-making assistance process. Detailed Implementation

[0036] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] like Figure 1 As shown, this invention proposes a method for embryo selection based on multimodal temporal fusion and anatomical prior constraints, comprising the following steps:

[0038] S1: Collect time-series image sequences of embryonic development and corresponding maternal multidimensional clinical data, construct a time-modal feature cube indexed by samples, obtain the input feature map of a single image, and use it as the input of the subsequent three-dimensional convolutional neural network. Here, "three-dimensional" does not refer to medical three-dimensional volume data, but rather to the combination of time dimension and two-dimensional image space.

[0039] S2: Apply U-net++ to segment embryo images in time-series image sequences into anatomical structures and generate anatomical prior masks, including key anatomical structure masks and foreground masks. This establishes anatomical prior constraints for focusing on depth feature extraction at different scales, that is, it provides local region prior knowledge for cell-level high-resolution feature extraction branches and global structural relationship prior knowledge for embryo-level context feature extraction branches.

[0040] S3: Introduce the anatomical prior mask obtained in step S2 into the three-dimensional convolutional neural network to obtain the improved three-dimensional convolutional neural network; construct a spatiotemporal attention module composed of channel attention submodule and spatial attention submodule based on anatomical structure constraints, and use the anatomical prior mask to guide the channel weight distribution and constrain the spatial attention weight respectively, so that the improved three-dimensional convolutional neural network can dynamically focus on key anatomical regions and their developmental morphological changes at continuous development time points when extracting temporal features;

[0041] S4: Design a multi-scale perception network to extract high-resolution cell-level features and embryo-level contextual features of the embryo through parallel paths, and introduce an adaptive gating fusion mechanism to dynamically weight and fuse the multi-scale features.

[0042] S5: Align dynamic clinical data into a clinical time series according to the collection time, and use a two-way... Temporal encoding was performed to obtain clinical temporal representations; a feature selection algorithm based on causal discovery was used to screen out a subset of causal features with direct causal effects on pregnancy outcomes from the fused multi-scale embryonic image temporal features and maternal clinical temporal features.

[0043] S6: Based on the selected subset of causal features, the inference is calculated by combining the gradient boosting decision forest model (GBDT) with the embryo development trajectory prediction model based on multimodal temporal fusion data, and the embryo development potential score, the causal contribution of each feature and the counterfactual inference results are output.

[0044] S7: Deploy an intelligent decision support system to receive time-series image sequences and corresponding new samples of maternal multidimensional clinical data, and output key developmental abnormality point identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

[0045] In step S3, based on anatomical structural constraints, a spatiotemporal attention module consisting of a channel attention submodule and a spatial attention submodule is constructed, as follows:

[0046] Step S31: The channel attention submodule is a temporal channel attention submodule guided by the dissection of prior masks, specifically the original input feature map. The temporal features output by the intermediate layer of the 3D convolutional neural network are Embryo foreground mask and key anatomical structures of the embryo according to The spatiotemporal dimensions are interpolated and aligned to obtain the interpolated and aligned embryo foreground mask. Interpolated and aligned key anatomical structures of the embryo ;

[0047] Step S32: In the time dimension and spatial dimensions , The above describes the time series feature maps respectively. Performing global average pooling and global max pooling yields two results. 3D channel descriptor , ;

[0048] Step S33: Using the anatomical prior mask as the weighted region for channel statistics, calculate the channel descriptor that fuses the anatomical constraints. ,in This represents element-wise multiplication. Indicates in , , Global average pooling across dimensions;

[0049] Step S34: , and After concatenation, the data is input into a shared multilayer perceptron to generate importance weights for each channel. :

[0050]

[0051] in Indicates vector splicing; , The parameters are learnable for two fully connected layers; It is a non-linear activation function used to enhance the non-linear expressive power of the model; The activation function is sigmoid. Therefore, each channel weight is determined not only by the global channel feature response. , The determination is also influenced by the characteristic responses of the embryonic foreground and key anatomical structures at consecutive time points. Together, they guide and adapt to three-dimensional temporal feature scenarios composed of time dimension and two-dimensional spatial dimension.

[0052] Step S35: The spatial attention submodule fuses the anatomical prior mask with the spatial dimension of the feature map to obtain the spatial attention weights constrained by the anatomical structure. Specifically:

[0053] Step S351: Process the timing feature map Average pooling and max pooling are performed along the channel dimension, and the concatenation of these two methods is followed by 3D convolution to generate the original spatial attention weights based on the feature map itself. :

[0054]

[0055] in, Indicates in Average pooling operation in dimension, Indicates in Max pooling operation in dimension, Indicates vector splicing. It is the sigmoid activation function. Represents a 3D convolution operation;

[0056] Step S352: Then, apply the interpolated and aligned embryo foreground mask. Interpolated and aligned key anatomical structures of the embryo Compared with the original spatial attention weights Fusion, generating spatial attention weights constrained by anatomical structure :

[0057] ;

[0058] in, Represents the embryo foreground mask aligned with the original spatial attention size, including a pixel-level mask of the entire region occupied by the embryo in the image; A pixel-level mask representing the key anatomical structures of the embryo after alignment with the original spatial attention size, including the pronuclear region, blastomeres region, blastocoel region, inner cell mass region, and trophoblast region; It is the sigmoid activation function. , These are learnable anatomical constraint strength parameters, used to control the embryo foreground mask. The constraint strength of spatial attention weights (i.e., the embryonic region in the image) and the control of key anatomical structure masks. The constraint strength of spatial attention weights on the various key biological anatomical structures within the embryo.

[0059] Step S36: As Figure 2 As shown, the final output feature map is calculated. :

[0060] ;

[0061] in, This represents element-wise multiplication; Here, represents the spatial attention weights constrained by the anatomical structure, and represents the channel attention weights guided by anatomical priors. During the calculation, it is necessary to... Broadcast to along the time and space dimensions to Same spacetime size ,Will Broadcast along the channel dimension to Same spacetime size This allows for the simultaneous implementation of channel importance weighting and spatial anatomical region constraints.

[0062] In step S4, the specific implementation of the multi-scale sensing network and the adaptive gating fusion mechanism is as follows:

[0063] Step S41: The process of designing and building a multi-scale sensing network is as follows:

[0064] Step S411: The final output feature map from step S36 serves as the anatomical constraint temporal feature map. The multi-scale sensing network consists of two parallel encoding paths, including a cell-level high-resolution feature extraction branch. and embryo-level contextual feature extraction branch They are respectively applied to the anatomical constraint temporal feature map. To extract features at different scales, such as details and global context, in parallel, where:

[0065] Step S4111: Cell-level high-resolution feature extraction branch By employing 1×1×1 convolution dimensionality reduction, 3×3×3 dilated convolution, and residual connections, blastomere boundaries, fragmentation ratios, and cell morphology textures are extracted without reducing spatial resolution, preserving high-resolution spatial details and obtaining cell-level features. ;

[0066] Step S4112: Embryo-level contextual feature extraction branch An architecture employing 3×3×3 convolutions with a stride of 2, combined with spatiotemporal pyramid pooling and global contextual attention, is used to extract contextual features from the overall embryo outline, blastocoel expansion trend, and the overall spatial relationship between the inner cell mass and trophoblast to capture global contextual and semantic information, resulting in low-resolution contextual features. ;

[0067] Step S4113: Process the obtained low-resolution context features in the spatial dimension Perform trilinear interpolation upsampling operation Thus obtaining cell-level characteristics Embryo-level features of the same spatiotemporal size This enables the cell-level high-resolution feature extraction branch. and embryo-level contextual feature extraction branch Feature fusion, thereby, cellular-level features Embryo-level characteristics All are constrained by the same anatomical time-series feature map The samples were extracted separately through parallel paths and used as input for the subsequent adaptive gating fusion mechanism.

[0068] Step S42: Adaptive Gated Feature Dynamic Fusion:

[0069] The obtained cellular-level features with aligned spatiotemporal dimensions Embryo-level characteristics The layers are stitched together along the channel dimension and then passed through a 1×1×1 three-dimensional convolutional layer. The data is then fused and reduced to a single channel, and activated by the Sigmoid function to generate a gated weight map. :

[0070] ;

[0071] Step S43: As Figure 3As shown, the temporal features of the fused multi-scale embryonic images are obtained.

[0072] The specific implementation process of step S5 includes:

[0073] Step S51: Introduce a bidirectional LSTM network to encode the collected maternal dynamic clinical data; let the maternal dynamic clinical data sequence be... The sequence contains If there are n time steps, then the output of the bidirectional LSTM at time t is: Finally, the concatenated hidden state of the last time step or the pooled result of the outputs of all time steps is taken as the encoded parent clinical time-series feature. ;

[0074] Step S52: Design a feature selection algorithm based on causal discovery to filter out a subset of causal features that have a causal effect on pregnancy outcomes; the feature selection algorithm based on causal discovery is a combination of PC algorithm and causal forest, and the specific combination process is as follows:

[0075] Step S521: Let the set of all variable characteristics be... ,in As a feature dimension, it includes temporal features of multi-scale embryo images. and maternal clinical time sequence characteristics Define the target variable For pregnancy outcome, ,in This indicates a successful pregnancy outcome. This indicates a failed pregnancy outcome; dataset ,in For the sample size, For the first The set of variable features for each sample For the first Pregnancy outcomes of one sample;

[0076] Step S522: The cause-effect graph based on the PC algorithm is constructed as follows:

[0077] A partially directed acyclic graph (PDAG) between features was constructed using the PC algorithm to identify maternal clinical time-series features. Temporal features of multi-scale embryo images The potential causal relationship between them is as follows:

[0078] Step S5221: Initialize the undirected graph; construct the complete undirected graph The set of nodes in the graph The set of edges of a graph Includes all node pairs Undirected edges between them;

[0079] Step S5222: Perform conditional independence testing, focusing on maternal clinical time-series characteristics. Temporal features of multi-scale embryo images The potential causal edges lay the groundwork for subsequent V-structure orientation and PDAG generation:

[0080] For each node pair Test in turn on different subsets That is, from the set of nodes Remove the currently pending node pair from the middle. When the remaining set of nodes is reached, and Whether they are conditionally independent; use the G-squared test, the test statistic is:

[0081] ;

[0082] in, This indicates the number of samples for the corresponding value combination;

[0083] like That is, the degrees of freedom are Chi-square distribution Quantiles, then determine and In the given If the conditions are independent, remove the edge between them;

[0084] Step S5223: Orient the edges of the undirected graph according to the V-structure and orientation rules to generate a partially directed acyclic graph (PDAG): if nodes exist... ,and and If there is no boundary between them, then the orientation is... If a path exists ,and and If there is no boundary between them, then the orientation is... ;in, Indicates a directed edge; Indicates an undirected edge;

[0085] Step S523: Extract the target variable from the generated partially directed acyclic graph (PDAG). The set of direct causal parent nodes That is, all pairs The specific method for identifying candidate feature subsets with direct causal influence is as follows:

[0086] Step S524: Traverse the generated partially directed acyclic graph (PDAG) and the target variable. If there exists a directed edge among all adjacent edges. → Then the characteristic variable of the directed edge join in and judged it as a Candidate causal features with direct causal influence. The determination of direct causal influence requires that the feature variable... With target variable There is a pointer between them Directly directed edges, i.e. → Y exists and this edge is not indirectly connected to any other intermediate node. Meanwhile, in subsequent causal forest estimation, The conditional average treatment effect meets the preset effect size threshold. and statistical significance threshold ;

[0087] Step S525: Simultaneously, for the undirected edges in the PDAG that are still undirected... — That is, the characteristic variables are not explicitly shown. With target variable The causal relationship between them, firstly Record in the set of undetermined parent nodes If the timing of embryonic development, the timing of maternal clinical data collection, and prior medical knowledge are combined, it can be determined that... Occurred in Previously, and not present in PDAG → A directed path will then If a candidate parent node passes the causal forest significance test, and passes the time sequence, medical prior, and subsequent CATE significance tests, it can be included in the subset of parent nodes with confirmed causal relationships. If it exists → If a directed path exists, or if a reverse causal relationship cannot be ruled out through chronological order and medical a priori reasoning, then it will not be considered. Include in the set of direct causal parent nodes These are retained only as features to be verified. Therefore, the final set of candidate causal features can be represented as follows: .

[0088] Step S526: Use the final candidate causal feature set The heterogeneous treatment effect is estimated based on the characteristics in the data. The conditional average treatment effect of the causal forest estimation is calculated as follows:

[0089] For the candidate causal feature set in causal graph G Use causal forest to estimate each feature For pregnancy outcomes Conditional average treatment effect (CATE):

[0090] Step S5261: Transfer the dataset Randomly divided into Bootstrap sample set ;

[0091] Step S5262: For each Bootstrap sample set Train a regression tree Each leaf node Corresponding to a sample subset The splitting criterion for the regression tree is to minimize the mean squared error.

[0092] ;

[0093] in, and These are the sample means of the left and right child nodes, respectively. Indicates the splitting threshold;

[0094] Step S5263: For features In each leaf node In the middle, the conditionally averaged treatment effect (CATE) is estimated:

[0095] ;

[0096] in, and leaf nodes middle and The number of samples;

[0097] Step S5264: Average and fuse the CATE estimates of all regression trees to obtain the features. Final conditional average treatment effect estimation:

[0098] ;in, It is the first Trees in feature value Estimation of the conditional average treatment effect at the location;

[0099] Step S527: Causal Feature Subset Screening: Based on the Conditional Average Treatment Effect (CATE) and the statistical significance threshold, select the top... A subset of causal features consists of features that have significant positive or negative causal effects, as follows:

[0100] Step S5271: Perform a statistical significance test; for all variables in the constructed causal graph G that are related to pregnancy outcome The set of direct causal parent nodes with direct causal influence Each candidate causal feature in ,use Test method test Is it significantly different from 0?

[0101] ;

[0102] in, Indicates the first The significance test statistic for the causal effect of each feature; yes The standard error, estimated using Bootstrap:

[0103] ;

[0104] Step S5272: Perform feature filtering; select those that meet the criteria. and of Each feature constitutes a subset of causal features. :

[0105] ;

[0106] in, This is the causal effect size; The preset CATE threshold, i.e., the effect size threshold; The statistical significance threshold is... for The statistical significance value of the test;

[0107] Step S5273: As Figure 4 As shown, the subset of causal features Sort in descending order to get The final subset of causal features is composed of these features.

[0108] Step S6 includes the following steps:

[0109] Step S61: The causal reinforcement gradient boosting decision forest model is based on the standard gradient boosting model, and a causal regularization term is introduced into the loss function. Causal correlation constraints are imposed to ensure that the feature weights learned by the causal-enhanced gradient boosting decision forest model are consistent with the average treatment effect estimated by the causal forest in both direction and magnitude. This ensures that the model weights are consistent with the causal discovery results, thereby improving the interpretability and stability of the model. The total loss function for training the causal-enhanced gradient boosting decision forest model is calculated. :

[0110] ;

[0111] in, For cross-entropy loss, To enhance the causal reinforcement of gradients in the decision forest model, the predicted clinical pregnancy probability is improved. This is the true label, i.e., the actual probability of pregnancy. For balance parameters; causal regularization term Defined as:

[0112] ;

[0113] in, Features learned by the gradient boosting decision forest model for causal reinforcement Features obtained by causal forest estimation The average treatment effect To normalize the absolute value of the causal effect size CATE to the [0,1] interval; This is the penalty coefficient for non-causal feature terms, which are those not in the causal feature subset. Features in;

[0114] Step S62: Train an embryonic development trajectory prediction model using multimodal temporal fusion data. Employ a conditional generative adversarial network (cGAN) combined with random perturbation modeling to predict the future development trajectory of the embryo. The prediction is then integrated into the embryo scoring system using confidence scores. The specific implementation process is as follows:

[0115] Step S621: Using the generator Generate predicted subsequent developmental trajectories; based on the temporal features of the fused multi-scale embryonic images obtained in step 43, assume the observed pre-... Temporal characteristics of embryo images at various time points The subsequent developmental trajectory to be predicted, i.e., the future developmental stage... The embryonic time sequence characteristics at each time point are as follows: The generator is based on a 3D U-Net architecture, with input before... Embryo temporal characteristics at each time point and a random noise vector Output the predicted future number Embryonic temporal characteristics at each time point:

[0116] ;

[0117] Step S622: Use a discriminator Determine the predicted trajectory The authenticity, that is, the authenticity of something labeled as authentic. The similarity of actual features at each time point is as follows:

[0118] The 3D PatchGAN structure is used, with the input being the temporal features of observed embryo images. With real future images (True sample group) or predicted image The concatenation of (fake sample groups) over time outputs the probability of authenticity:

[0119] ;

[0120] Step S623: Calculate the model loss function: The loss function of the entire conditional generative adversarial network consists of two parts, including the discriminator loss function. and generator loss function These are used to constrain the discriminator and generator to update network parameters, respectively; where the discriminator loss function is... The standard binary cross-entropy loss aims to maximize the discriminator's ability to distinguish true trajectories.

[0121] ;

[0122] in, For the first The complete true sequence of each sample The sample generated by the generator will be the future. Feature sequences at each time point;

[0123] Generator loss function Consistent with the losses from combat and reconstruction ( The algorithm consists of two parts (regularization and regularization), which calculates and minimizes the probability that the trajectory generated by the generator will be identified as a false trajectory by the discriminator.

[0124] ;

[0125] in, For the first Future of individual samples In the real sequence at each time point The characteristic value at time; The sample generated by the generator will be the future. In the time series The characteristic value at time; These are weighting coefficients, used to balance adversarial loss and trajectory fitting loss;

[0126] Step S624: Calculate the confidence score of the predicted trajectory: Combine the probability of the generator's predicted trajectory being true with the trajectory error to obtain the final confidence score of the embryonic development trajectory prediction.

[0127] ;

[0128] Among them, the generator predicts the trajectory. With the actual trajectory The mean squared error; the confidence level of embryonic developmental trajectory will also serve as an important influencing factor in subsequent embryonic developmental potential scoring;

[0129] Step S625: Calculate the results for any input sample using the trained causal boosting gradient boosting decision forest model. Pregnancy outcome prediction probability:

[0130] ;

[0131] in, That is, the model operation function. For core feature subset The corresponding input feature vector, The activation function is responsible for converting the model's output values ​​into probabilities.

[0132] Step S626: Calculate the output pregnancy outcome prediction probability. The embryonic developmental potential score is calculated by weighting and fusing the results with the confidence score of the embryonic developmental trajectory, and then mapping the results to a standardized 10-point scoring interval.

[0133] ;

[0134] in, and The weights for the influence of learnable embryonic developmental potential scores are as follows:

[0135] Step S627: Apply the SHAP model to calculate the causal contribution of each feature, observe the impact of different features on the prediction of the final pregnancy outcome, thereby improving the interpretability of the model; for the input sample Each of its features causal contribution The product of its SHAP value and the sign of the characteristic causal effect:

[0136] ;

[0137] in, Features For the sample Predicted probability The SHAP value; the sign of the product indicates the direction of the causal influence of the feature on the pregnancy outcome, including positive contribution or negative constraint, and the magnitude of its absolute value indicates the order of contribution.

[0138] Step S628: For the input sample If one of the features The value was intervened as The counterfactual reasoning engine answers the question of how the predicted results change; the sample after intervention is denoted as... Substituting this into the causal enhanced gradient boosting decision forest model and then passing it through an activation function, the counterfactual inference result is calculated. :

[0139] ;

[0140] in, To train the causal augmented gradient boosting decision forest model, This refers to the counterfactual prediction result calculated and inferred by the model after intervention or perturbation is applied to the feature terms in the input sample; the change in the counterfactual score is... It is used to quantify the impact of interventions on embryonic developmental potential scores from the perspectives of clinical variables or changes in embryonic developmental characteristics.

[0141] In step S7, the identification of key developmental abnormalities is achieved through a temporal anomaly detection algorithm, and the implementation process is as follows:

[0142] The temporal trajectory of embryonic development is modeled, and the reconstruction error based on autoencoder or the sequence prediction error based on Transformer is used as the anomaly score. When the anomaly score at a certain time point exceeds the threshold, it is marked as a developmental anomaly point and associated with anatomical structure. Specific anomaly descriptions such as "asynchronous blastomeres on day 3" and "delayed cell mass development on day 5" are output. The specific implementation process is as follows:

[0143] Step S71: Based on the fused multi-scale embryonic image temporal features obtained in step 43, the embryonic temporal feature sequence is set as follows: ,in Indicates the first Feature vectors at each time point For feature dimension, Number of time points;

[0144] Self-encoder ,in For the encoder, feature vectors are mapped to latent representations. ; For the decoder, the latent representation is reconstructed as ;

[0145] Transformer prediction model Before input The feature sequence at the i-th time point is used to predict the i-th time point. Features at each time point ;

[0146] Step S72: Anomaly detection of reconstruction error based on autoencoder

[0147] The autoencoder is trained using temporal feature sequences from normal embryos, with the loss function being reconstruction loss.

[0148] ;

[0149] in, This represents the number of normal embryo samples. For the first The first normal embryo in the second Characteristics of each point in time;

[0150] The Adam optimizer is used, with a learning rate set to... Training continues until the loss reconstructed on the validation set converges;

[0151] For the embryonic time sequence to be tested Calculate the reconstruction error at each time point:

[0152] ;

[0153] Calculate the outlier score, using the reconstruction error as the outlier score:

[0154] ;

[0155] The threshold is then determined using the reconstruction error distribution of normal embryos. Take the 95th percentile of the remodeling error of a normal embryo:

[0156] ;

[0157] in, Indicates the sample index; Indicates the time step index;

[0158] Finally, outliers are determined based on the outlier score and threshold; when At that time, mark the first These time points are considered outliers;

[0159] Step S73: Transformer-based sequence prediction error anomaly detection

[0160] A Transformer-based prediction model was trained using temporal feature sequences from normal embryos. The Transformer model employed a 6-layer encoder and a 6-layer decoder, with 256 hidden layers and 8 attention heads. The Adam optimizer was used, and the learning rate was set to [value missing]. The loss function is the prediction loss:

[0161] ;

[0162] For the embryonic time sequence to be tested Calculate the prediction error at each time point: ;

[0163] in, The reconstruction error of the autoencoder is used as the anomaly score.

[0164] Calculate the outlier score, using the prediction error as the outlier score:

[0165] ;

[0166] The threshold is then determined using the prediction error distribution of normal embryos. Take the 95th percentile of the prediction error for normal embryos:

[0167] ;

[0168] Finally, outliers are determined based on the outlier score and threshold; when At that time, mark the first These time points are considered outliers;

[0169] Step S74: Correlation between outliers and anatomical structures: Extraction of anatomical features for each outlier time point. Extract the anatomical structural features at this time point; the anatomical structural features are as follows:

[0170] Blastomer characteristics: The number of blastomeres is The coefficient of variation in blastomer size is ; Indicates time The standard deviation of the size of all blastomeres at that time; Indicates time The average size of all blastomeres at that time;

[0171] Inner cell mass characteristics: The area of ​​the inner cell mass is Inner cell mass cell density is ; Indicates time The number of cells in the cell cluster within a time period;

[0172] Fragmentation rate characteristics: Fragmentation rate is ; Indicates time The total area of ​​fragments within the embryo at that time; Indicates time The total area of ​​the entire embryo at that time;

[0173] Step S75: Determine the abnormality based on the anatomical structure features and the abnormality description rules, as follows:

[0174] like and The diagnosis was asynchronous blastomeres on day 3.

[0175] like and The cell cluster development was determined to be delayed within the first 5 days; among them and Normal embryos at the 1st Mean and standard deviation of cell cluster area within 14 days;

[0176] like It was determined to be the first The fragmentation rate of the sky is abnormally high; among them and Normal embryos at the 1st Mean and standard deviation of the fragmentation rate;

[0177] Finally, an anomaly description is generated based on the anomaly, including the anomaly time point, anomaly score, anomaly condition, and anatomical structural features.

[0178] Step S76: For a patient's current treatment cycle, a new sample set consisting of several candidate embryos. , Based on the number of candidate embryos and the identification of key developmental abnormalities, the system enables prediction of clinical pregnancy probability, ranking of embryos for live birth potential, and individualized transfer recommendations, as detailed below:

[0179] The methods in steps S1-S5 are used to extract features from embryo sequence data and construct causal feature subsets in the new sample set. Then, the causal-enhanced gradient boosting decision forest model trained in step S6 is used to calculate the clinical pregnancy probability of each embryo.

[0180] ;

[0181] in, For the new sample set One embryo causal feature subset Candidate causal features in;

[0182] Based on the calculated clinical pregnancy probability for each embryo Then, combined with the developmental trajectory confidence score output by the embryonic development trajectory prediction model in step S624... The developmental potential score of each embryo is calculated by weighted summation:

[0183] ;

[0184] in, and The weights for the influence of learnable embryonic developmental potential scores are as follows: , Results of embryo live birth potential ranking According to List of embryos arranged in descending order: ;

[0185] Combining the identification results of key developmental abnormalities, clinical pregnancy probability prediction, embryo live birth potential ranking, and maternal clinical characteristics, structured embryo transfer recommendations are generated, such as... Figure 5 As shown;

[0186] Example: If the embryo ranks first in the overall score of No key abnormalities were detected, and the suggested text is: It is recommended to prioritize embryo transfer. The predicted clinical pregnancy probability is ( ×100)%%; the embryo's developmental trajectory is normal; if the embryo ranks first in the overall score... of Therefore, the recommended approach is to choose the best embryo available at the moment. Given the low probability of pregnancy, it is recommended to consider multiple embryo transfer or proceed to the next treatment cycle, taking into account the patient's wishes.

[0187] This invention also discloses an embryo selection system based on multimodal temporal fusion and anatomical prior constraints, characterized in that the system performs the above-described embryo selection method based on multimodal temporal fusion and anatomical prior constraints, including:

[0188] The multi-source time-series data acquisition module is used to simultaneously acquire embryo time-series image sequences and corresponding maternal multidimensional clinical data during the embryo culture process;

[0189] The anatomical structure-guided multi-scale feature extraction module includes an embryonic anatomical structure segmentation unit, a spatiotemporal attention module, and a multi-scale perception network, as detailed below:

[0190] The embryonic anatomical structure segmentation unit achieves pixel-level segmentation of key anatomical structures based on the U-Net++ network, obtaining anatomical prior masks, namely embryonic key anatomical structure masks and embryonic foreground masks; anatomical prior constraints are established for focusing on the extraction of depth features at different scales, that is, providing prior knowledge of local region details for cell-level high-resolution feature extraction and prior knowledge of global structural relationships for embryonic-level contextual feature extraction.

[0191] The spatiotemporal attention module, whose spatial attention submodule integrates anatomical prior masks, guides an improved 3D convolutional neural network based on anatomical constraints to dynamically focus on the developmental changes of key anatomical regions when extracting temporal features.

[0192] The multi-scale perception network extracts high-resolution cell-level features and embryo-level contextual features of the embryo through parallel paths, and dynamically weights and fuses the multi-scale features through an adaptive gating mechanism.

[0193] The cross-modal causal feature selection module includes a temporal feature encoding unit and a cross-modal causal feature selection unit, as detailed below:

[0194] The temporal feature encoding unit aligns dynamic clinical data into a clinical time series according to the acquisition time, and uses bidirectional LSTM for temporal encoding to obtain a clinical temporal representation;

[0195] The cross-modal causal feature selection unit uses a feature selection algorithm based on causal discovery to select a subset of core features that have a direct causal effect on pregnancy outcomes from the fused multi-scale temporal features and clinical features.

[0196] The causal-enhanced intelligent decision-making module, based on the selected subset of causal features, uses a gradient-enhanced decision forest model and combines it with an embryo development trajectory prediction model trained based on multimodal time-series fusion data to calculate inference, outputting embryo development potential score, causal contribution of each feature, and counterfactual inference results.

[0197] An interpretable clinical interaction platform provides a visual interface for dynamic heatmaps of embryonic development, causal contribution waterfall plots, time-series abnormality reports, and personalized transplantation recommendations. This enables the platform to receive time-series image sequences and corresponding new samples of multidimensional maternal clinical data, and output key developmental abnormality identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

[0198] The system is deployed using a federated learning framework, supports multi-center distributed training, can dynamically detect distribution drift, automatically trigger online learning, and adaptively adjust model parameters to maintain prediction accuracy.

[0199] The system integrates a dynamic decision calibrator that continuously monitors the model's predictive performance in real clinical settings. When distribution drift is detected, it automatically triggers online learning and Bayesian updates to adjust model parameters to maintain predictive accuracy, forming a closed-loop intelligent evolutionary system. The specific implementation process is as follows:

[0200] First, a sliding window mechanism is used to dynamically monitor and detect drift; a window of size is maintained. Recent sample queue sliding window ; Using expected calibration error As a drift detection metric, the calibration deviation of samples within a window is calculated periodically:

[0201] ;

[0202] The predicted probability interval [0,1] is divided into equal intervals. The present invention sets a range of intervals. ; To fall into the first A subset of samples from each interval. and These represent the classification accuracy and average prediction confidence of the sample subset, respectively.

[0203] when Exceeding the preset threshold When a distribution drift is detected, the calibration process is triggered.

[0204] Then, the output probability of the causal-enhanced gradient boosting decision forest model is calibrated using a temperature scaling method; in the sliding window queue dataset. The optimal temperature parameters are learned by minimizing the negative log-likelihood loss:

[0205] ;

[0206] in, For sliding window queues The predicted clinical pregnancy outcome for each new sample is calculated, and the calibrated predicted probability is as follows:

[0207] ;

[0208] Finally, online learning and Bayesian updates are performed based on the temperature calibration results; if calibration is completed, the calculation... Still above the threshold This triggers online learning of the gradient boosting decision forest model with causal reinforcement, using the current global parameters. Let the prior mean be a Gaussian prior. ;in, To control the hyperparameters of the balance between new and old knowledge, The identity matrix; in the sliding window queue dataset The updated parameters are obtained by maximizing the posterior probability:

[0209] ;

[0210] in, The log-likelihood of the window data drives the model to learn a new distribution; The L2 regularization term constrains the new parameters to not deviate too far from the original parameters; after the update, it is... replace This leads to the adaptive evolution of the gradient boosting decision forest model with causal reinforcement.

[0211] This embodiment is based on the intelligent embryo selection method proposed in this invention, which uses an attention mechanism based on multimodal temporal fusion and anatomical structural constraints, and proceeds as follows:

[0212] S1: Multimodal temporal data acquisition and construction of a time-modal feature cube indexed by samples. This involves acquiring embryonic time-series image sequences at multiple key time points from fertilized egg to blastocyst stage, and simultaneously acquiring corresponding maternal dynamic clinical data. For a single sample... Its characteristics can be expressed as The dimensional features containing temporal and spatial information. By aligning batch time-series image sequences, clinical variables, and timestamps, a formal representation can be constructed. The input dataset, where For the sample size, For the number of time points, For modal channels, , This represents the spatial dimension of the image. This structured data representation forms the basis for subsequent deep analysis.

[0213] S2: Embryo Anatomy Segmentation and Prior Mask Generation Based on U-Net++. A pre-trained U-Net++ segmentation network is used to perform pixel-level parsing of temporal images. Through an encoder-decoder structure and dense skip connections, high-precision output is achieved for each time point. Multiple binary masks for the embryo ,in Representing different anatomical structures ( =1 represents the inner cell mass. =2 represents the nutrient layer. These masks constitute an indispensable set of spatial prior knowledge with clear biological significance for subsequent steps. .

[0214] S3: Spatiotemporal attention feature extraction incorporating anatomical priors. This invention designs an improved 3D convolutional neural network that embeds a spatiotemporal attention module constrained by anatomical structures. A learnable constraint strength parameter is introduced, enabling the network to adaptively balance the contributions of data-driven attention and anatomical priors, avoiding over-suppression or under-attention caused by a fixed mask. Given the feature map output of the l-th convolutional layer... This module generates enhanced feature maps according to the following formula. :

[0215] Channel attention: for feature maps In the time dimension and spatial dimensions , Global average pooling and max pooling are performed on the above, that is, in addition to the channel dimension. Aggregate across all dimensions outside the scope to generate two The 3D channel descriptors are merged after passing through a shared multilayer perceptron (MLP) to obtain the channel attention weight vector. .

[0216] Spatial attention (introducing anatomical constraints): First, for In channel dimension Perform average pooling and max pooling on the top and concatenate them to obtain a... Feature descriptors are convolved to generate initial spatial attention weights. The key anatomical structure mask obtained in step S2 corresponding to the current time point. and embryo foreground mask Introduced as a learnable constraint:

[0217] ;

[0218] in, It is the Sigmoid activation function. For standard convolutional layers, and These are trainable parameters used to adaptively balance the contribution of data-driven attention and anatomical priors. This allows the network to focus on key anatomical regions of the embryo from the early stages of training, and to be further refined during training to pay more attention to developmental changes in these key anatomical regions.

[0219] Feature enhancement: The final output feature map is ,in This represents element-wise multiplication. This process is repeated at multiple layers of the network, enabling the model to focus on the dynamic changes of key anatomical structures at different levels of abstraction.

[0220] S4: Multi-scale perception and adaptive gating fusion. To simultaneously capture cellular-level details and embryonic-level global morphology, a dual-branch feature pyramid is constructed. High-resolution branches process the original image to extract detailed features. Low-resolution branch processing downsamples the image to extract contextual features. An adaptive gating fusion unit is introduced to integrate low-level contextual features. Upsampling To high-level details For the same space dimensions, Subsequently, the two are concatenated along the channel dimension, fused and reduced to a single channel through a 1×1×1 convolutional layer, and then activated by the Sigmoid function to generate a spatially gated weight map. , .in This indicates channel splicing. Features after fusion. , This is element-wise multiplication. This mechanism allows the network to dynamically determine whether to rely more on detailed features or contextual features at each spatial location based on the content of the input image.

[0221] S5: Cross-modal dynamic feature selection based on causal discovery. This involves using the temporal-multi-scale image feature vectors extracted in S3 and S4. Compared with preprocessed clinical feature vectors Early splicing was To eliminate redundancy and spurious associations, a causal discovery framework based on the PC algorithm is applied. This algorithm progressively constructs the relationship between features and pregnancy outcomes through conditional independence tests and G-squared. The part of the directed acyclic graph between them identifies the feature set of the direct causal parent nodes. and subsets of parent nodes that do not explicitly show a direct causal relationship but are confirmed to have a causal relationship through testing. Subsequently, a causal forest was used to refine the final candidate causal feature set. Heterogeneous treatment effects are estimated based on the characteristics in the data, and the conditionally averaged treatment effect (CATE) is quantified. Finally, conditions with an absolute CATE value greater than a preset threshold are selected. Furthermore, the statistically significant features constitute a concise feature subset with strong causal explanatory power. This step is the cornerstone of the model's interpretability and robustness.

[0222] S6: Training and Counterfactual Reasoning of Causal Regularized Decision Models. Train a gradient boosting decision tree model. To prevent the model from learning biased correlations, a standard binary cross-entropy loss function can be used. An innovative causal regularization term is added to the basics. This allows the model weights to be better aligned with the causal discovery features. Thus, the model's total loss function can be expressed as: Total losses .in, The model assigns features The importance weight can be calculated by ranking the importance of features. It is the CATE value of this feature estimated in S5. It is a balance parameter. This is the penalty coefficient for non-causal feature terms. This regularization term forces the model weights to align with causal effects. After training, the model can not only predict probabilities... ;in, That is, the model operation function. For core feature subset The corresponding input feature vectors are responsible for converting the model's output values ​​into probabilities. Its inherent counterfactual reasoning engine can be intervened upon. The counterfactual reasoning result is calculated based on a certain characteristic value (such as "set the age to 35"). ,observe The changes in these input indicators can be used to answer clinical questions such as "What changes will occur in the output if a certain input indicator is intervened or a specific perturbation is applied?"

[0223] S7: Integrated intelligent decision support and visualization output. The S1-S6 processes are engineered into a low-latency inference pipeline and encapsulated as a RESTful API service. The front-end interaction platform calls this service, uploads data, and returns structured JSON results within seconds. The front-end renders these results as an embryonic development heatmap (based on S3's attention weights) and a causal contribution waterfall plot (based on S6's...). and The system includes a time-series development trajectory curve and a counterfactual simulation results panel, providing clinicians with a panoramic and interactive decision-making dashboard.

[0224] Based on the above method, a hierarchical, modular, and loosely coupled embryo intelligent selection system based on multimodal temporal fusion and anatomically constrained attention is proposed. The modules of this system communicate with each other through standardized data interfaces. The specific structure and connection relationships are as follows:

[0225] Data Access and Governance Layer: This layer includes a multi-source time-series data acquisition module and a data preprocessing engine. The acquisition module, through an adapter pattern, supports automatic data extraction or passive data reception from various heterogeneous data sources, such as hospital LIS, electronic medical records, and microscope workstations with scheduled image capture. The preprocessing engine is responsible for cleaning, aligning, and formatting the raw data into the input dataset defined by S1, serving as the unified starting point for all subsequent analyses.

[0226] Core Intelligent Processing Layer: This layer is the computing hub of the system, adopting a microservice architecture and containing three sequentially executed and functionally decoupled core services:

[0227] (a) Anatomy-Guided Feature Extraction Service: This service encapsulates the complete workflow from S2 to S4. Internally, it first calls the anatomical segmentation sub-service to segment the embryonic anatomical structure and generate prior masks. Then, the spatiotemporal feature extraction sub-service extracts the spatiotemporal attention features that have been incorporated into the anatomical priors. Finally, the multi-scale fusion sub-service performs feature weighting to generate adaptive spatial gating, and then processes the image and mask in parallel. This service receives data cubes through a message queue, and after processing, outputs high-dimensional feature vectors to a shared cache.

[0228] (b) Causal Feature Selection Service: This service subscribes to a feature cache and executes S5's causal discovery and feature selection algorithm. It includes a causal graph learning unit and an effect estimation unit, which work together to ultimately select the best features. The corresponding causal graph model parameters are stored in the model database.

[0229] (c) Intelligent Decision Service: This service is the "brain" of the system, integrating the causal augmented prediction model trained with S6 with the counterfactual reasoning engine. It loads the latest models from the model database and Defined as a tool that performs real-time reasoning on input features and can respond to counterfactual query requests from higher layers.

[0230] Application and Interaction Layer: This layer centers on an interpretable clinical interaction platform, providing a web front-end interface. Upon user interaction, the platform sends a request to the task scheduler. The scheduler coordinates the sequence of calls to underlying services: first, it triggers the data access layer to process new data; then, it sequentially calls the three services of the core processing layer; finally, it integrates all results, including predicted scores, heatmap data, causal graphs, and counterfactual results, before returning them to the front-end for visualization. Authentication and routing between layers and modules are achieved through an API gateway, ensuring high availability and scalability of the system.

[0231] Example 1: Complete System Training and Performance Verification

[0232] In this embodiment, a historical dataset containing 8500 complete IVF cycles was used to train and validate the system described in this invention, divided into a training set, a validation set, and an independent test set in a 6:2:2 ratio. Each cycle included embryo time-series images on days 1, 3, 5, and 6 post-fertilization, along with over 50 corresponding clinical variables. To scientifically evaluate the actual effectiveness and contribution of each innovative module of this invention, four comparative models were trained:

[0233] (1) Model A (Baseline Model): The static single-time-point image classification model was evaluated using images of embryos at the blastocyst stage on day 5 post-fertilization. ResNet50 was used as the backbone feature extractor. The last fully connected classification layer was removed, and the output of the final global average pooling layer was used as a 1024-dimensional feature vector. The processing flow was as follows: the input was a single RGB blastocyst image, which was uniformly scaled to 224×224 pixels. After the image was forward-propagated through ResNet50, the resulting 1024-dimensional feature vector was directly fed into a small classification network consisting of two fully connected layers (the first layer of this network transforms the feature dimension from 1024 to 256 and inputs it into the ReLU activation function; the second layer flattens the 256-dimensional input from the upper layer into a 1-dimensional vector and then inputs it into the Sigmoid activation function), and finally directly outputs the predicted pregnancy probability value of the embryo.

[0234] (2) Model B (enhanced model): a multi-scale feature fusion model that incorporates temporal and high semantic features, that is, based on model A, it introduces temporal information of embryonic development and multi-scale morphological features.

[0235] First, temporal processing is performed. The input is expanded to include image sequences at four key time points: days 1, 3, 5, and 6. A shallow 3D convolutional neural network consisting of two cascaded 3D convolutional-pooling modules is constructed to extract spatiotemporal joint features, outputting a feature tensor that incorporates temporal information. Simultaneously, a feature pyramid network is applied to the day 5 blastocyst image. Utilizing the intermediate layer outputs of the ResNet50 network, multi-scale feature fusion is performed through lateral connections and upsampling, thus constructing a multi-scale feature map that integrates high-resolution details and high semantic information. Finally, the dynamic feature vector obtained from the temporal 3D convolutional neural network is concatenated with the multi-scale feature vector extracted from the feature pyramid network for feature integration and decision-making, forming a comprehensive feature vector. This comprehensive vector is ultimately input into a classification network with the same structure as Model A for prediction, yielding the corresponding results.

[0236] (3) Model C (improved model): a temporal attention model with added anatomical structure constraints, that is, an anatomical attention mechanism that simulates the visual evaluation logic of embryologists is added to the architecture of Model B.

[0237] First, anatomical prior generation is performed. In the preprocessing stage, an independently trained U-Net++-based anatomical segmentation network is introduced. Taking the embryo image as input, it outputs pixel-level segmentation masks that accurately identify key regions such as the "inner cell mass," "trophoblast," and "blastocoel." These masks do not participate in the main network training but serve only as prior knowledge. Then, an attention mechanism is integrated to modify the model's network structure. A custom dual-path attention module is inserted after each basic convolutional block of the 3D convolutional neural network. This module computes two attention maps in parallel: one is standard channel attention, used to calibrate the importance of feature channels; the other is anatomically guided spatial attention, which weights the spatial weight map learned by the network itself and the anatomical structure mask from the segmentation network corresponding to the current time point, achieving element-level semantic superposition, and is then activated by the Sigmoid function. The weights of the anatomical masks are trainable parameters, allowing the network to adaptively adjust its dependence on prior knowledge. In this way, when the network processes image sequences along the time dimension, its spatial focus region is dynamically and softly constrained to biologically meaningful anatomical structures, thereby extracting more discriminative and interpretable joint "structure-temporal" features. The improved model, which incorporates an attention mechanism with anatomical constraints, has the same classification network as Model A and Model B and can predict the corresponding results.

[0238] (4) Model D (Complete Model of the Invention): Further add a causal enhancement decision model, that is, on the basis of Model C, add a causal inference framework, so that the model is upgraded from correlation learning to causal effect learning.

[0239] After model C completes training and is able to extract high-quality features, its parameters are fixed, and forward propagation is performed using all training data to obtain a deep feature set for each sample. This feature set is then concatenated with the original clinical feature table to form a hybrid feature pool. Subsequently, a causal discovery algorithm based on conditional independence testing is applied to analyze this hybrid feature pool and pregnancy outcome labels to construct a causal graph between variables. Based on this, a causal forest algorithm is used to quantify and estimate the conditional average treatment effect of each feature in the graph on pregnancy outcome, retaining only those features whose CATE values ​​are statistically significant and whose absolute values ​​are greater than a preset threshold, forming a concise feature subset with strong causal interpretability. Using the causal feature subset as input, a new, lightweight gradient boosting decision tree model is constructed as the final decision maker. A causal alignment regularization term is added to the standard binary cross-entropy loss as a soft constraint, forcing the prediction model to maintain its internal decision logic as consistent as possible with the discovered causal structure while optimizing classification performance, greatly improving the model's robustness and interpretability.

[0240] Based on the above, the new input samples for inference first pass through the feature extraction pathway of model C to generate initial deep features. These features are then combined with clinical features and filtered according to a predefined subset of causal features. Finally, the filtered features are input into a causal regularized gradient boosting decision-maker to obtain the final predicted probability and the contribution of each feature. Furthermore, by manipulating the variable values ​​in the causal feature subset, the system can also perform counterfactual inference.

[0241] All models were optimized and evaluated using AUC as the primary metric on the same training / validation / test set partitioning. Performance comparisons on the 2000-episode independent test set are shown in Table 1. AUC is recorded as mean ± standard deviation, where the mean is the average of the AUCs from ten independent repeated experiments, and the standard deviation after ± represents the typical fluctuation range of the results across the ten experiments; a smaller standard deviation indicates less performance fluctuation and higher stability between training iterations. Table 1. Performance Comparison of the Complete System (Test Set)

[0242] Data Analysis and Conclusions: Incremental performance improvement: Model A uses only static images as core features and serves as the baseline for all models. It has the lowest performance metrics: mean AUC 0.832, accuracy 79.1%, specificity 81.3%, and sensitivity 76.5%; its stability variance is 1.44e-4, and its standard deviation is 0.012, which is also the highest among the four models, indicating the greatest fluctuation in results. Model B, based on Model A, adds time-series and multi-scale features. All performance metrics are significantly improved: the mean AUC increases to 0.883, accuracy to 83.7%, specificity to 85.2%, and sensitivity to 81.8%; at the same time, the stability variance decreases to 1.00e-4, and the standard deviation decreases to 0.010, indicating that the model's discriminative ability and stability are improved simultaneously. Model C, based on Model B, adds an attention mechanism with anatomical structure constraints. Performance continues to steadily improve: the mean AUC rises to 0.918, accuracy to 87.2%, specificity to 88.9%, and sensitivity to 85.2%; the stability variance further decreases to 0.64e-4, and the standard deviation to 0.008, further enhancing the model's ability to capture key features and the stability of results. Model D, based on Model C, adds a causal selection and regularization module. It is the best model among the four models in all metrics: the mean AUC reaches 0.953, accuracy is 90.2%, specificity is 91.8%, and sensitivity is 88.4%; the stability variance is reduced to 0.25e-4, and the standard deviation is reduced to 0.005. It achieves the best performance in discriminant ability, prediction accuracy, positive and negative sample identification ability, and operational stability.

[0243] In summary, the AUC significantly improved from model A to model D (p<0.01, paired t-test). In particular, the attention mechanism and causal framework that introduced anatomical constraints brought the greatest marginal gain.

[0244] Significantly enhanced stability: Model D has the smallest AUC variance, indicating that causal feature selection and regularization effectively filter out noise and spurious associations in the data, greatly improving the robustness of the model on different subsets of data, which is crucial for clinical deployment.

[0245] Interpretability achieved: The proposed model D can generate a visual attention heatmap and provide a quantitative causal contribution report, achieving a leap from "visible" to "understandable".

[0246] Example 2: Case Study of In-Depth Clinical Decision Analysis and Counterfactual Generation This embodiment demonstrates the system's deep decision support capabilities in complex clinical scenarios through a specific case. As follows: The patient, a 45-year-old female, had an AMH level of 1.2 ng / mL and a previous failed embryo transfer. Three transferable blastocysts were obtained in this cycle. Their sequence data were input into the system of this invention for inference and a comprehensive report was generated, displaying the pregnancy outcome score for each embryo.

[0247] Taking embryo A (9.1 points), which has the highest score, as an example, we analyze the causal decision contribution of this embryo sample score, including core positive contributions and negative conditional constraints. Specifically: Core positive contributions (total +4.1): +2.5 points: Derived from the morphology of cell clusters on day 5 (Grade A). The system attributes this to the high number and density of cells in this structure, which is identified in the causal diagram as having a strong direct causal effect on pregnancy outcome.

[0248] +1.6 points: derived from the blastocyst expansion rate from day 3 to day 5. The system's time-series modeling captured its synchronized and rhythmic expansion, and CATE estimation showed that this dynamic feature contributed significantly.

[0249] Major negative constraints (total -1.8): -1.2 points: Directly attributed to patient age. Causal findings identified it as the root node affecting multiple downstream related features such as oocyte mitochondrial function, and its negative effects were directly learned by the model.

[0250] -0.6 points: attributed to a fragmentation rate of 15% on day 3. Although the absolute value of the system's causal decision judgment is not high, its negative impact is amplified in this age context.

[0251] Based on the above scoring deconstruction, the doctor further used counterfactual reasoning to raise two questions: Question 1: If this patient is 35 years old, what would the embryo score and pregnancy probability be? The system's causal decision output modifies the age node value to 35 in the causal graph and recalculates it along the causal path.

[0252] The results showed that the embryo prediction score improved from 9.1 to 9.7, and the clinical pregnancy probability increased from 65% to 72%, which directly quantified the direct impact of age.

[0253] Question 2: What would happen if the fragmentation rate of the embryo could be reduced to 5% by day 3? The system's causal decision output sets the fragmentation rate node value to 5% in the causal graph and recalculates it along the causal path.

[0254] The results showed that the score improved to 9.5 and the clinical pregnancy probability increased to 69%, which provides a quantitative direction for optimizing the culture of superior embryos.

[0255] Meanwhile, the system issued an abnormal pattern warning for embryos with the lowest score of 5: "The system detected a blastomere size difference coefficient >0.4 on day 2 and early compaction signs on day 3. This abnormal timing pattern is strongly correlated with an implantation rate of less than 5% in historical data. It is recommended to inform the patient that the success rate is extremely low." It is worth noting that in this embodiment, the input is in vitro blastocyst sequence / morphological data. The system performs causal reasoning, embryo pregnancy outcome scoring, causal contribution deconstruction, counterfactual intervention simulation, and implantation success rate warning output. All calculations occur within the computer system, without performing any physical operations on the patient or the embryo. Counterfactual reasoning is a computer-generated virtual node intervention simulation; it does not actually change the patient's age or modify the embryo fragmentation rate. It is merely a simulation of the causal graph within the model and does not constitute any treatment or treatment operation on the human body / embryo.

[0256] Example 3: Deployment and Performance Verification of a Federated Learning System Based on Secure Aggregation This embodiment details the system's simulated three-center deployment strategy within a federated learning framework, designed to address data privacy regulations in multi-center collaboration scenarios. The three centers hold 3000, 4000, and 3500 local periodic data points, respectively. The system employs a client / server federated learning paradigm, but is customized to suit the core functional modules of this invention. The server holds the initial parameters for the global feature extraction network (S2-S4) and the causal decision model (S6); each client locally stores the raw data and installs lightweight client software.

[0257] System deployment and operation process: The server initializes and distributes global model parameters.

[0258] Each client at each center performs the complete forward propagation from S2 to S6 using local data. Cross-modal dynamic feature selection is performed only on their respective data, and the results are not shared initially. The client calculates the update gradients for its local model parameters, adds noise to the gradients using differential privacy techniques, and uploads them to the server after secure multi-party computation and encryption.

[0259] The server uses the FedAvg algorithm to aggregate gradients and update the global model. Specifically, for causal discovery results, the server only aggregates stable causal edges that appear in more than two clients to form a consensus causal graph.

[0260] The server distributes the updated global model. After 50 rounds of federated training, the performance of the federated global model and the independently trained models at each center is evaluated. On an external validation set consisting of 1200 epochs, the AUC of the federated global model is 0.948, while the AUC of the independent models at each center ranges from 0.918 to 0.932, with a mean of 0.925. The federated model significantly outperforms any independent model (p < 0.05).

[0261] This embodiment demonstrates that the multimodal temporal fusion and anatomical prior constraint embryo selection system proposed in this invention can effectively integrate multi-center knowledge through federated learning without sharing any original sensitive clinical data, thereby training a more generalizable and robust global model. The construction of the consensus causal graph further enhances the universality and reliability of the medical knowledge acquired by the model across centers, laying the foundation for establishing industry-wide intelligent evaluation standards.

[0262] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for embryo selection based on multimodal temporal fusion and anatomical prior constraints, characterized in that: Includes the following steps: S1: Collect embryonic time sequence images and maternal dynamic clinical data as input to a three-dimensional convolutional neural network; S2: U-net++ is used to segment embryonic temporal images to obtain anatomical prior masks, including embryonic foreground masks and key anatomical structure masks, which provide local details and global structural prior constraints for cell-level high-resolution features and embryonic-level contextual features extraction, respectively. S3: Introduce anatomical prior masks into the three-dimensional convolutional neural network to construct a spatiotemporal attention module containing channel and spatial attention sub-modules. Use masks to guide channel weights and constrain spatial attention weights, so that the three-dimensional convolutional neural network can dynamically focus on key anatomical regions and their morphological changes along a continuous temporal sequence. S4: Construct a multi-scale perception network, extract cell-level high-resolution features and embryo-level contextual features in parallel, and obtain multi-scale embryo image temporal features by dynamically weighted fusion through an adaptive gating mechanism. S5: Align the maternal dynamic clinical data time to obtain the clinical time sequence, and obtain the clinical characterization through bidirectional time-series coding; use the causal discovery feature selection algorithm to screen the subset of causal features affecting pregnancy outcome from the fused multi-scale embryo image time-series features and clinical time-series features; S6: Based on causal feature subsets, a causal-enhanced gradient boosting decision forest model is adopted, and combined with multimodal temporal fusion embryonic development trajectory prediction model inference, to output embryonic development potential score, feature causal contribution degree and counterfactual inference results; S7: Deploy an intelligent decision support system to receive new sample data and output developmental abnormality identification, pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

2. The embryo selection method based on multimodal temporal fusion and anatomical prior constraints according to claim 1, characterized in that, In step S3, based on anatomical structural constraints, a spatiotemporal attention module consisting of a channel attention submodule and a spatial attention submodule is constructed, as follows: Step S31: The channel attention submodule is a temporal channel attention submodule guided by the dissection prior mask, specifically: the original input feature map The temporal feature map output after the intermediate layer of the 3D convolutional neural network is Embryo foreground mask and key anatomical structures of the embryo according to The spatiotemporal dimensions are interpolated and aligned to obtain the interpolated and aligned embryo foreground mask. Interpolated and aligned key anatomical structures of the embryo ; Step S32: In the time dimension and spatial dimensions , The above describes the time series feature maps respectively. Performing global average pooling and global max pooling yields two results. 3D channel descriptor , ; Step S33: Using the anatomical prior mask as the weighted region for channel statistics, calculate the channel descriptor with fused anatomical constraints. ; Step S34: , and After concatenation, the data is input into a shared multilayer perceptron to generate importance weights for each channel. ; Step S35: The spatial attention submodule fuses the anatomical prior mask with the spatial dimension of the feature map to obtain the spatial attention weights constrained by the anatomical structure. Specifically: Step S351: Process the timing feature map Average pooling and max pooling are performed along the channel dimension, and the concatenation of these two methods is followed by 3D convolution to generate the original spatial attention weights based on the feature map itself. ; Step S352: Then, apply the interpolated and aligned embryo foreground mask. Interpolated and aligned key anatomical structures of the embryo Compared with the original spatial attention weights Fusion, generating spatial attention weights constrained by anatomical structure ; Step S36: Calculate the final output feature map : 。 3. The embryo selection method based on multimodal temporal fusion and anatomical prior constraints according to claim 2, characterized in that, Step S4 is implemented as follows: Step S41: The process of building a multi-scale sensing network is as follows: Step S411: The final output feature map of step S36 serves as the anatomical constraint temporal feature map. The multi-scale sensing network consists of two parallel encoding paths, including a cell-level high-resolution feature extraction branch. and embryo-level contextual feature extraction branch They are respectively applied to the anatomical constraint temporal feature map. To extract features at different scales, such as details and global context, in parallel, where: Step S4111: Cell-level high-resolution feature extraction branch By employing 1×1×1 convolution dimensionality reduction, 3×3×3 dilated convolution, and residual connections, blastomere boundaries, fragmentation ratios, and cell morphology textures are extracted without reducing spatial resolution, preserving high-resolution spatial details and obtaining cell-level features. ; Step S4112: Embryo-level contextual feature extraction branch An architecture employing 3×3×3 convolutions with a stride of 2, combined with spatiotemporal pyramid pooling and global contextual attention, is used to extract contextual features from the overall embryo outline, blastocoel expansion trend, and the overall spatial relationship between the inner cell mass and trophoblast to capture global contextual and semantic information, resulting in low-resolution contextual features. ; Step S4113: Process the obtained low-resolution context features in the spatial dimension Perform trilinear interpolation upsampling operation Thus obtaining cell-level characteristics Embryo-level features of the same spatiotemporal size This enables the cell-level high-resolution feature extraction branch. and embryo-level contextual feature extraction branch The features are fused and used as input for the subsequent adaptive gating fusion mechanism; Step S42: Adaptive Gated Feature Dynamic Fusion: The obtained cellular-level features with aligned spatiotemporal dimensions Embryo-level characteristics The layers are stitched together along the channel dimension and then passed through a 1×1×1 three-dimensional convolutional layer. The data is then fused and reduced to a single channel, and activated by the Sigmoid function to generate a gated weight map. ; Step S43: Obtain the temporal features of the fused multi-scale embryonic images .

4. The embryo selection method based on multimodal temporal fusion and anatomical prior constraints according to claim 3, characterized in that, The specific implementation process of step S5 includes: Step S51: Introduce a bidirectional LSTM network to encode the collected parent dynamic clinical data to obtain the parent dynamic clinical data sequence. Take the concatenated hidden state of the last time step or the pooling result of the output of all time steps as the encoded parent clinical time-series feature. ; Step S52: Design a feature selection algorithm based on causal discovery to filter out a subset of causal features that have a causal effect on pregnancy outcomes; the feature selection algorithm based on causal discovery is a combination of PC algorithm and causal forest, and the specific process is as follows: Step S521: Let the set of all variable characteristics be... ,in As a feature dimension, it includes temporal features of multi-scale embryo images. and maternal clinical time sequence characteristics Define the target variable For pregnancy outcome, ,in This indicates a successful pregnancy outcome. Indicates a failed pregnancy outcome; dataset ,in For the sample size, For the first The set of variable features for each sample For the first Pregnancy outcomes of one sample; Step S522: The cause-effect graph based on the PC algorithm is constructed as follows: A partially directed acyclic graph (PDAG) between features was constructed using the PC algorithm to identify maternal clinical time-series features. Temporal features of multi-scale embryo images The potential causal relationship between them is as follows: Step S5221: Initialize the undirected graph; construct the complete undirected graph The set of nodes in the graph The set of edges of a graph Includes all node pairs Undirected edges between them; Step S5222: Perform conditional independence testing, focusing on maternal clinical time-series characteristics. Temporal features of multi-scale embryo images The potential causal edges lay the groundwork for subsequent V-structure orientation and PDAG generation: For each node pair Test in turn on different subsets That is, from the set of nodes Remove the currently pending node pair from the middle. When the remaining set of nodes is reached, and Whether they are conditionally independent; use the G-squared test, the test statistic is: ; in, This indicates the number of samples for the corresponding value combination; like That is, the degrees of freedom are Chi-square distribution Quantiles, then determine and In the given If the conditions are independent, remove the edge between them; Step S5223: Orient the edges of the undirected graph according to the V-structure and orientation rules to generate a partially directed acyclic graph (PDAG): if nodes exist... ,and and If there is no boundary between them, then the orientation is... If a path exists ,and and If there is no boundary between them, then the orientation is... ;in, Indicates a directed edge; Indicates an undirected edge; Step S523: Extract the target variable from the generated partially directed acyclic graph (PDAG). The set of direct causal parent nodes That is, all pairs The specific method for identifying candidate feature subsets with direct causal influence is as follows: Step S524: Traverse the generated partially directed acyclic graph (PDAG) and the target variable. If there exists a directed edge among all adjacent edges. → Then the characteristic variable of the directed edge join in and judged it as a Candidate causal features with direct causal influence; Step S525: Simultaneously, for the undirected edges in the PDAG that are still undirected... — That is, the characteristic variables are not explicitly shown. With target variable The causal relationship between them, firstly Record in the set of undetermined parent nodes If the timing of embryonic development, the timing of maternal clinical data collection, and prior medical knowledge are combined, it can be determined that... Occurred in Previously, and not present in PDAG → A directed path will then As a candidate parent node, it enters the causal forest significance test. If it passes the time sequence, medical prior, and subsequent CATE significance tests, it is included in the subset of parent nodes with confirmed causal relationships. If it exists → If a directed path exists, or if a reverse causal relationship cannot be ruled out through chronological order and medical a priori reasoning, then it will not be considered. Include in the set of direct causal parent nodes Only retain them as features to be verified; the final set of candidate causal features is represented as follows. ; Step S526: The conditional average treatment effect of the causal forest estimation is calculated as follows: For the candidate causal feature set in the causal graph G Use causal forest to estimate each feature For pregnancy outcomes Conditional average treatment effect (CATE): Step S5261: Transfer the dataset Randomly divided into Bootstrap sample set ; Step S5262: For each Bootstrap sample set Train a regression tree Each leaf node Corresponding to a sample subset The splitting criterion for the regression tree is to minimize the mean squared error. ; in, and These are the sample means of the left and right child nodes, respectively. Indicates the splitting threshold; Step S5263: For features In each leaf node In the middle, the conditionally averaged treatment effect (CATE) is estimated: ; in, and leaf nodes middle and The number of samples; Step S5264: Average and fuse the CATE estimates of all regression trees to obtain the features. Final conditional average treatment effect estimation: ;in, It is the first Trees in feature value Estimation of the conditional average treatment effect at the location; Step S527: Causal Feature Subset Screening: Based on the Conditional Average Treatment Effect (CATE) and the statistical significance threshold, select the top... A subset of causal features consists of features that have significant positive or negative causal effects, as follows: Step S5271: Perform a statistical significance test; for all variables in the constructed causal graph G that are related to pregnancy outcome The set of direct causal parent nodes with direct causal influence Each candidate causal feature in ,use Test method test Is it significantly different from 0? ; in, Indicates the first The significance test statistic for the causal effect of each feature; yes The standard error, estimated using Bootstrap: ; Step S5272: Perform feature filtering; select those that meet the criteria. and of Each feature constitutes a subset of causal features. : ; in, This is the causal effect size; The preset CATE threshold, i.e., the effect size threshold; The statistical significance threshold is... for The statistical significance value of the test; Step S5273: Subset the causal features Sort in descending order to get The final subset of causal features is composed of these features.

5. The embryo selection method based on multimodal temporal fusion and anatomical prior constraints according to claim 4, characterized in that, Step S6 includes the following steps: Step S61: The causal reinforcement gradient boosting decision forest model is based on the standard gradient boosting model, and a causal regularization term is introduced into the loss function. Causal correlation constraints are imposed to ensure that the feature weights learned by the gradient boosting decision forest model with causal reinforcement are consistent with the average treatment effect estimated by the causal forest in terms of direction and magnitude, so as to ensure that the model weights are consistent with the causal discovery results. Calculate the total loss function for training a gradient boosting decision forest model with causal reinforcement. : ; in, For cross-entropy loss, To enhance the causal reinforcement of gradients in the decision forest model, the predicted clinical pregnancy probability is improved. This is the true label, i.e., the actual probability of pregnancy. For balance parameters; causal regularization term Defined as: ; in, Features learned by the gradient boosting decision forest model for causal reinforcement The weight, Features obtained by causal forest estimation The average treatment effect To normalize the absolute value of the causal effect size CATE to the [0,1] interval; This is the penalty coefficient for non-causal feature terms, which are those not in the causal feature subset. Features in; Step S62: Train an embryonic development trajectory prediction model using multimodal temporal fusion data. Employ a conditional generative adversarial network (cGAN) combined with random perturbation modeling to predict the future development trajectory of the embryo. The prediction is then integrated into the embryo scoring system using confidence scores. The specific implementation process is as follows: Step S621: Using the generator Generate predicted subsequent developmental trajectories; based on the temporal features of the fused multi-scale embryonic images obtained in step 43, assume the observed pre-... Temporal characteristics of embryo images at various time points The subsequent developmental trajectory to be predicted, i.e., the future developmental stage... The embryonic time sequence characteristics at each time point are as follows: ; Step S622: Use a discriminator Determine the predicted trajectory The authenticity, that is, the authenticity of something labeled as authentic. The similarity of actual features at each time point is as follows: The 3D PatchGAN structure is used, with the input being the temporal features of observed embryo images. With real future images Or predict image By concatenating data along the time dimension, the probability of authenticity is output. ; Step S623: Calculate the model loss function: The loss function of the entire conditional generative adversarial network consists of two parts, including the discriminator loss function. and generator loss function They are used respectively for the constraint discriminator and generator Update network parameters; where the discriminator loss function is... The standard binary cross-entropy loss aims to maximize the discriminator. The ability to distinguish real trajectories: ; in, For the first The complete true sequence of each sample The sample generated by the generator will be the future. Feature sequences at each time point; Generator loss function Composed of two parts: adversarial loss and reconstruction loss, the probability of the generator-generated trajectory being identified as a false trajectory by the discriminator is calculated and minimized. ; in, For the first Future of individual samples In the real sequence at each time point The characteristic value at time; The sample generated by the generator will be the future. In the time series The characteristic value at time; These are weighting coefficients, used to balance adversarial loss and trajectory fitting loss; Step S624: Calculate the confidence score of the predicted trajectory: Combine the probability of the generator's predicted trajectory being true with the trajectory error to obtain the final confidence score of the embryonic development trajectory prediction. ; in, Predicting trajectories for generator With the actual trajectory The mean square error; Step S625: Calculate the results for any input sample using the trained causal boosting gradient boosting decision forest model. Pregnancy outcome prediction probability: ; in, That is, the model operation function. For core feature subset The corresponding input feature vector, The sigmoid activation function is responsible for converting the model's output values ​​into probabilities. Step S626: Calculate the output pregnancy outcome prediction probability. The embryonic developmental potential score is calculated by weighting and fusing the results with the confidence score of the embryonic developmental trajectory, and then mapping the results to a standardized 10-point scoring interval. ; in, and The weights for the influence of learnable embryonic developmental potential scores are as follows: Step S627: Apply the SHAP model to calculate the causal contribution of each feature, observe the impact of different features on the prediction of the final pregnancy outcome, thereby improving the interpretability of the model; for the input sample Each of its features causal contribution The product of its SHAP value and the sign of the characteristic causal effect: ; in, Features For the sample Predicted probability The SHAP value; the sign of the product indicates the direction of the causal influence of the feature on the pregnancy outcome, including positive contribution or negative constraint, and the magnitude of its absolute value indicates the order of contribution. Step S628: For the input sample If one of the features The value was intervened as The counterfactual reasoning engine answers the question of how the predicted results change; the sample after intervention is denoted as... Substituting this into the causal enhanced gradient boosting decision forest model and then passing it through an activation function, the counterfactual inference result is calculated. : ; in, To train the causal augmented gradient boosting decision forest model, This refers to the counterfactual prediction result calculated and inferred by the model after intervention or perturbation is applied to the feature terms in the input sample; the change in the counterfactual score is... It is used to quantify the impact of interventions on embryonic developmental potential scores from the perspectives of clinical variables or changes in embryonic developmental characteristics.

6. The embryo selection method based on multimodal temporal fusion and anatomical prior constraints according to claim 5, characterized in that, In step S7, the identification of key developmental anomalies is achieved through a temporal anomaly detection algorithm as follows: The temporal characteristic trajectory of embryonic development is modeled, and the reconstruction error based on autoencoder or the sequence prediction error based on Transformer is used as the anomaly score. When the anomaly score at a certain time point exceeds a threshold, it is marked as a developmental anomaly point, associated with anatomical structure, and an anomaly description is output. The specific implementation process is as follows: Step S71: Based on the fused multi-scale embryonic image temporal features obtained in step 43, the embryonic temporal feature sequence is set as follows: ,in Indicates the first Feature vectors at each time point For feature dimension, Number of time points; Self-encoder ,in For the encoder, feature vectors are mapped to latent representations. ; For the decoder, the latent representation is reconstructed as ; Transformer prediction model Before entering The feature sequence at the i-th time point is used to predict the i-th time point. Features at each time point ; Step S72: Anomaly detection of reconstruction error based on autoencoder: The autoencoder is trained using temporal feature sequences from normal embryos, with the loss function being reconstruction loss. ; in, This represents the number of normal embryo samples. For the first The first normal embryo in the second Characteristics of each point in time; The Adam optimizer is used, with a learning rate set to... Training continues until the loss reconstructed on the validation set converges; For the embryonic time sequence to be tested Calculate the reconstruction error at each time point: ; Calculate the outlier score, using the reconstruction error as the outlier score: ; The threshold is then determined using the reconstruction error distribution of normal embryos. Take the 95th percentile of the remodeling error of a normal embryo: ; in, Indicates the sample index; Indicates the time step index; Finally, outliers are determined based on the outlier score and threshold; when At that time, mark the first These time points are considered outliers; Step S73: Transformer-based sequence prediction error anomaly detection: A Transformer-based prediction model was trained using temporal feature sequences from normal embryos. The Transformer model employed a 6-layer encoder and a 6-layer decoder, with 256 hidden layers and 8 attention heads. The Adam optimizer was used, and the learning rate was set to [value missing]. The loss function is the prediction loss: ; For the embryonic time sequence to be tested Calculate the prediction error at each time point: ; in, The reconstruction error of the autoencoder is used as the anomaly score. Calculate the outlier score, using the prediction error as the outlier score: ; The threshold is then determined using the prediction error distribution of normal embryos. Take the 95th percentile of the prediction error for normal embryos: ; Finally, outliers are determined based on the outlier score and threshold; when At that time, mark the first These time points are considered outliers; Step S74: Correlation between anomalies and anatomical structures: Extraction of anatomical features for each anomalous time point. Extract the anatomical structural features at this time point; the anatomical structural features are as follows: Blastomer characteristics: The number of blastomeres is The coefficient of variation in blastomer size is ; Indicates time The standard deviation of the size of all blastomeres at that time; Indicates time The average size of all blastomeres at that time; Inner cell mass characteristics: The area of ​​the inner cell mass is Inner cell mass cell density is ; Indicates time The number of cells in the cell cluster within a time period; Fragmentation rate characteristics: Fragmentation rate is ; Indicates time The total area of ​​fragments within the embryo at that time; Indicates time The total area of ​​the entire embryo at that time; Step S75: Determine the abnormal situation based on the abnormal description rules matching the anatomical structure features; finally, generate an abnormal description containing the abnormal time point, abnormal score, abnormal situation and anatomical structure features based on the abnormal situation. Step S76: For a patient's current treatment cycle, a new sample set consisting of several candidate embryos. , Based on the number of candidate embryos and the identification of key developmental abnormalities, the system enables prediction of clinical pregnancy probability, ranking of embryos for live birth potential, and individualized transfer recommendations, as detailed below: The methods in steps S1-S5 are used to extract features from embryo sequence data and construct causal feature subsets in the new sample set. Then, the causal-enhanced gradient boosting decision forest model trained in step S6 is used to calculate the clinical pregnancy probability of each embryo. ; in, For the new sample set One embryo causal feature subset Candidate causal features in; Based on the calculated clinical pregnancy probability for each embryo Then, combined with the developmental trajectory confidence score output by the embryonic development trajectory prediction model in step S624, The developmental potential score of each embryo is calculated by weighted summation: ; in, and The weights for the influence of learnable embryonic developmental potential scores are as follows: , Results of embryo live birth potential ranking According to List of embryos arranged in descending order: ; By combining the results of key developmental abnormality identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and maternal clinical characteristics, structured embryo transfer recommendations are generated.

7. An embryo selection system based on multimodal temporal fusion and anatomical prior constraints, characterized in that, The system performs the embryo selection method based on multimodal temporal fusion and anatomical prior constraints as described in any one of claims 1-6, comprising: The multi-source time-series data acquisition module is used to simultaneously acquire embryo time-series image sequences and corresponding maternal multidimensional clinical data during the embryo culture process; The anatomical structure-guided multi-scale feature extraction module includes an embryonic anatomical structure segmentation unit, a spatiotemporal attention module, and a multi-scale perception network, as detailed below: The embryonic anatomical structure segmentation unit achieves pixel-level segmentation of key anatomical structures based on the U-Net++ network, obtaining anatomical prior masks, namely embryonic key anatomical structure masks and embryonic foreground masks. This establishes anatomical prior constraints for focusing on depth feature extraction at different scales, providing prior knowledge of local region details for cell-level high-resolution feature extraction and prior knowledge of global structural relationships for embryonic-level contextual feature extraction. The spatiotemporal attention module, whose spatial attention submodule integrates anatomical prior masks, guides an improved 3D convolutional neural network based on anatomical constraints to dynamically focus on the developmental changes of key anatomical regions when extracting temporal features. The multi-scale perception network extracts high-resolution cell-level features and embryo-level contextual features of the embryo through parallel paths, and dynamically weights and fuses the multi-scale features through an adaptive gating mechanism. The cross-modal causal feature selection module includes a temporal feature encoding unit and a cross-modal causal feature selection unit, as detailed below: The temporal feature encoding unit aligns dynamic clinical data into a clinical time series according to the acquisition time, and uses bidirectional LSTM for temporal encoding to obtain a clinical temporal representation; The cross-modal causal feature selection unit uses a feature selection algorithm based on causal discovery to select a subset of core features that have a direct causal effect on pregnancy outcomes from the fused multi-scale temporal features and clinical features. The causal-enhanced intelligent decision-making module, based on the selected subset of causal features, calculates inference through a causal-enhanced gradient boosting decision forest model combined with an embryo development trajectory prediction model based on multimodal temporal fusion data, and outputs embryo development potential score, causal contribution of each feature, and counterfactual inference results. An interpretable clinical interaction platform provides a visual interface for dynamic heatmaps of embryonic development, causal contribution waterfall plots, time-series abnormality reports, and personalized transplantation recommendations. This enables the platform to receive time-series image sequences and corresponding new samples of multidimensional maternal clinical data, and output key developmental abnormality identification, clinical pregnancy probability prediction, embryo live birth potential ranking, and individualized transplantation recommendations.

8. The embryo selection system based on multimodal temporal fusion and anatomical prior constraints according to claim 7, characterized in that, The system integrates a dynamic decision calibrator that continuously monitors the model's predictive performance in real clinical settings. When distribution drift is detected, it automatically triggers online learning and Bayesian updates to adjust model parameters to maintain predictive accuracy, forming a closed-loop intelligent evolutionary system. The implementation process is as follows: First, a sliding window mechanism is used to dynamically monitor and detect drift; a window of size is maintained. Recent sample queue sliding window ; Using expected calibration error As a drift detection metric, the calibration deviation of samples within a window is calculated periodically: ; The predicted probability interval [0,1] is divided into equal intervals. A range; To fall into the first A subset of samples from each interval. and These represent the classification accuracy and average prediction confidence of the sample subset, respectively. when Exceeding the preset threshold When a distribution drift is detected, the calibration process is triggered. Then, the output probability of the causal-enhanced gradient boosting decision forest model is calibrated using a temperature scaling method; in the sliding window queue dataset. The optimal temperature parameters are learned by minimizing the negative log-likelihood loss: ; in, For sliding window queues The predicted clinical pregnancy outcome for each new sample is calculated, and the calibrated predicted probability is as follows: ; Finally, online learning and Bayesian updates are performed based on the temperature calibration results; if calibration is completed, the calculation... Still above the threshold This triggers online learning of the gradient boosting decision forest model with causal reinforcement, using the current global parameters. Let the prior mean be a Gaussian prior. ;in, To control the hyperparameters of the balance between new and old knowledge, The identity matrix; in the sliding window queue dataset The updated parameters are obtained by maximizing the posterior probability: ; in, The log-likelihood of the window data drives the model to learn a new distribution; The L2 regularization term constrains the new parameters to not deviate too far from the original parameters; after the update, it is... replace This leads to the adaptive evolution of the gradient boosting decision forest model with causal reinforcement.

Citation Information

Patent Citations

  • Freezing embryo transplantation outcome prediction method based on multi-modal data

    CN119090820A

  • Embryo development stage classification method and system based on residual perception and wavelet coding

    CN120747631A