A method and system for spontaneous preterm birth prediction based on multimodal data
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-24
- Publication Date
- 2026-08-14
AI Technical Summary
单次评估与序贯信息利用不足:大部分早产预测方法仅分析单张影像,未充分利用孕期多次检查中积累的序贯影像,且缺乏动态变化信息分析能力,难以根据动态信息对早产风险进行有效预测,限制了其在临床决策中的应用价值
本发明通过构建影像序列构建单元、视觉特征提取单元、检查级信息聚合单元、病例级多模态融合单元、早产预测单元,可以从检查序列图像中提取孕检一维特征,得到多个一维图像特征向量;并能够动态捕捉一维图像特征向量之间的语义关联与互补信息,得到若干个检查级表示向量,完成从局部图像特征到全局检查信息的聚合;再通过学习若干个检查级表示向量之间的演变关联,模拟孕期动态规律,并耦合临床文本数据,生成病例级的动态变化信息,实现对整个检查信息的全局聚合,最后输出早产亚型的预测结果,从而可以充分利用孕期多次检查中积累的序贯影像,并且能够准确分析孕检图像的动态变化信息,根据动态变化信息对早产风险进行有效预测,有效提升早产预测方案在临床决策中的应用价值,因而可以有效解决现有早产预测技术中序贯信息利用不足、多模态数据整合有限、临床适应性弱等核心痛点。
Smart Images

Figure CN122575753A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a spontaneous preterm birth prediction method and system based on multimodal data, belonging to the field of preterm birth prediction technology. Background Technology
[0002] Current methods for predicting preterm birth mainly rely on biomarker detection (cfDNA, mRNA, fetal fibronectin) or ultrasound imaging to measure cervical length (CL). These methods require additional testing, have limited coverage, and are costly. Some imaging methods rely on transvaginal ultrasound (TVUS), but due to differences in pregnant women's acceptance, hospital equipment, and geographical location, some areas only have transabdominal ultrasound (TAUS) data available, resulting in a lack of universally applicable and feasible methods for predicting preterm birth.
[0003] Therefore, existing methods for predicting preterm birth have the following main technical problems: Insufficient use of single assessment and sequential information: Most methods for predicting preterm birth only analyze a single image, failing to make full use of sequential images accumulated from multiple examinations during pregnancy, and lack the ability to analyze dynamic changes. This makes it difficult to effectively predict the risk of preterm birth based on dynamic information, thus limiting their application value in clinical decision-making.
[0004] Limited integration of clinical text and multimodal information: Existing methods for predicting preterm birth do not make sufficient use of information such as the pregnant woman's past medical history and clinical text. Text data is mostly used for stratification of high-risk groups rather than for general screening of the entire population. Most ultrasound imaging-based methods only use cervical ultrasound images and cannot effectively use ultrasound images of other tissues from a large number of prenatal examinations.
[0005] The information disclosed in this background section is only for understanding the background of the inventive concept, and therefore may include information that does not constitute prior art. Summary of the Invention
[0006] To address the aforementioned problems or one of them, the present invention aims to provide a spontaneous preterm birth prediction method and system based on multimodal data. This method can fully utilize sequential images accumulated from multiple prenatal examinations and accurately analyze the dynamic changes in these images. Based on these dynamic changes, it can effectively predict the risk of preterm birth, thereby enhancing the application value of preterm birth prediction strategies in clinical decision-making. Therefore, it can effectively solve the core pain points of existing preterm birth prediction technologies, such as insufficient utilization of sequential information, limited integration of multimodal data, and weak clinical adaptability.
[0007] To address the aforementioned problems or one of them, the second objective of this invention is to provide a spontaneous preterm birth prediction method and system based on multimodal data. This method can dynamically capture the semantic associations and complementary information of features from multiple images within the same examination, solving the problem of fragmented information from multiple images in a single examination. Furthermore, by learning the evolutionary associations of representation vectors at different examination levels through case-level multimodal fusion units, it simulates the dynamic patterns of physiological changes during pregnancy. Compared to static single-frame analysis, this effectively improves the sensitivity of the unit model to subtle feature changes in the preterm stage of preterm birth.
[0008] To address the aforementioned problems, or one of them, the third objective of this invention is to provide a spontaneous preterm birth prediction method and system based on multimodal data. This method can perfectly match the actual scenarios of pregnancy follow-up without requiring additional testing items, reducing medical costs and the burden on pregnant women. It can be directly integrated into existing obstetric ultrasound examination procedures, making it highly practical. It provides standardized preterm birth prediction tools for primary hospitals and significantly improves the accuracy, robustness, and clinical adaptability of spontaneous preterm birth prediction. It can effectively assist in the early identification of preterm birth risks and the development of stratified intervention strategies, and has important clinical value in reducing the incidence of preterm birth and improving maternal and infant outcomes.
[0009] To achieve one of the above objectives, the first technical solution of the present invention is as follows: A method for spontaneous preterm birth prediction based on multimodal data includes the following steps: Step 1: Based on the pre-constructed image sequence generation unit, acquire the pregnancy examination images obtained by a pregnant woman through N ultrasound examinations, and according to the pregnant woman's clinical follow-up records, assemble the pregnancy examination images into an examination sequence image according to the examination time order. Step 2: Using a pre-constructed visual feature extraction unit, extract one-dimensional features of the pregnancy test from the examination sequence images to obtain multiple one-dimensional image feature vectors; Step 3: Using the pre-constructed inspection-level information aggregation unit, the semantic association and complementary information between one-dimensional image feature vectors are dynamically captured to obtain several inspection-level representation vectors, thus completing the aggregation from local image features to global inspection information. Step four: Using a pre-constructed case-level multimodal fusion unit, the evolutionary relationships between several examination-level representation vectors are learned to simulate the dynamic patterns of pregnancy and couple clinical text data to generate case-level dynamic change information, thereby achieving global aggregation of the entire examination information. Step 5: Using a pre-built preterm birth prediction unit, the dynamic change information is processed, and the prediction results of the preterm birth subtype are output to achieve spontaneous preterm birth prediction.
[0010] This invention, by constructing an image sequence construction unit, a visual feature extraction unit, an examination-level information aggregation unit, a case-level multimodal fusion unit, and a preterm birth prediction unit, can extract one-dimensional features of pregnancy examination from examination sequence images, obtaining multiple one-dimensional image feature vectors. It can also dynamically capture the semantic relationships and complementary information between these one-dimensional image feature vectors, obtaining several examination-level representation vectors, thus completing the aggregation from local image features to global examination information. Furthermore, by learning the evolutionary relationships between these examination-level representation vectors, it simulates the dynamic patterns of pregnancy and couples clinical text data to generate case-level dynamic change information, achieving global aggregation of the entire examination information. Finally, it outputs the prediction result of the preterm birth subtype. This allows for full utilization of sequential images accumulated from multiple prenatal examinations and accurate analysis of the dynamic changes in pregnancy examination images. Based on these dynamic changes, it effectively predicts the risk of preterm birth, significantly enhancing the application value of preterm birth prediction strategies in clinical decision-making. Therefore, it effectively addresses the core pain points of existing preterm birth prediction technologies, such as insufficient utilization of sequential information, limited multimodal data integration, and weak clinical adaptability.
[0011] Furthermore, this invention can fully integrate multimodal information from clinical text data and pregnancy examination images, thereby effectively utilizing various ultrasound images during pregnancy examinations. Unlike traditional methods that only analyze single ultrasound images, this invention integrates complete image data from N ultrasound examinations of pregnant women through an image sequence generation unit, constructing examination sequence images along a time axis. Therefore, it can achieve full-cycle, continuous feature capture of pregnancy ultrasound images, fully exploring the dynamic changes of key anatomical structures such as cervical morphology and placental position with gestational age, providing core time-dimensional evidence for predicting the risk of preterm birth.
[0012] The visual feature extraction unit of this invention can convert two-dimensional ultrasound images into one-dimensional feature vectors that retain core information. Combined with self-supervised training, it can achieve cross-modal feature alignment for different types of ultrasound images, such as transabdominal and transvaginal ultrasound images, solve the feature deviation problem caused by equipment differences and different imaging methods, and effectively improve the feature extraction accuracy.
[0013] Furthermore, this invention uses the self-attention mechanism of the examination-level information aggregation unit to dynamically capture the semantic association and complementary information of multiple image features within the same examination, thus solving the problem of fragmented information in multiple images from a single examination. Then, through the case-level multimodal fusion unit, it learns the evolutionary association of different examination-level representation vectors to simulate the dynamic laws of physiological changes during pregnancy. Compared with static single-frame analysis, this invention can effectively improve the sensitivity to subtle feature changes in the preterm labor stage (such as progressive shortening of the cervix).
[0014] Furthermore, this invention enables coupled modeling of ultrasound imaging data with clinical textual data (medical history, structured indicators, etc.). Through the global self-attention mechanism of the case-level multimodal fusion unit, semantic-level fusion of imaging anatomical features and clinical risk background is achieved. For example, associating the textual feature of "past preterm birth" with the imaging feature of "cervical internal os morphology" can effectively distinguish between "physiological cervical shortening" and "pathological preterm birth risk," avoiding misjudgment based on single-modal data.
[0015] As a preferred technical measure: Step 1: Based on the pre-built image sequence generation unit, acquire pregnancy examination images obtained from N ultrasound examinations of a pregnant woman, and according to the pregnant woman's clinical follow-up records, assemble the pregnancy examination images into an examination sequence image according to the examination time order as follows: Acquire pregnancy examination images of a pregnant woman obtained through N ultrasound examinations, including transabdominal ultrasound images and / or transvaginal ultrasound images and / or cervical related ultrasound images, and associate them with the pregnant woman's clinical follow-up records, which include the timestamp and examination type of each examination. Median filtering was used to remove noise, and the pregnancy test images were preprocessed and the contrast was equalized to obtain pregnancy test images with the same brightness and small gain differences. Based on the examination timestamps in the clinical follow-up records, all pregnancy examination images corresponding to a single examination are summarized to obtain a pregnancy examination image group; The pregnancy examination images are sorted according to the order in which the examinations were conducted, forming a time-series image sequence.
[0016] As a preferred technical measure: Step two, using a pre-constructed visual feature extraction unit, extracts one-dimensional features from the examination sequence images to obtain multiple one-dimensional image feature vectors. The method is as follows: Each pregnancy test image in the sequence of images is divided into image blocks of fixed size to unify the image input dimensions; The pre-trained visual feature extraction unit is invoked to extract features from each image block to obtain pregnancy test features. The extracted pregnancy test features are then linearly transformed and mapped to a one-dimensional space to generate corresponding one-dimensional pregnancy test feature fragments. The one-dimensional feature segments of all image blocks are spliced together in their original spatial order to form a one-dimensional image feature vector of the pregnancy examination image, thus realizing the conversion from two-dimensional image to one-dimensional feature. The above process is used to process all pregnancy test images in the sequence of images to obtain multiple one-dimensional image feature vectors corresponding to the number of pregnancy test images, ensuring that each one-dimensional image feature vector retains the core feature information of the pregnancy test image.
[0017] As a preferred technical measure: The method for self-supervised training of the visual feature extraction unit is as follows: Collect unlabeled pregnancy examination images, including transabdominal and / or transvaginal ultrasound images related to the placenta and amniotic fluid, and associate them with the pregnant woman's clinical follow-up records, which include timestamps and examination types for each examination. Median filtering was used to remove noise, and the pregnancy test images were preprocessed and the contrast was equalized to obtain pregnancy test images with the same brightness and small gain differences. The pregnancy examination images are then cropped using a multi-scale cropping method to generate large-size global slices and small-size local slices, which serve as training inputs. The large-size global slices are used to capture macroscopic anatomical relationships, while the small-size local slices are used to uncover microscopic textures. Based on a self-supervised pre-training algorithm, a two-branch model with student and teacher branches is constructed, both of which use visual converters. The parameters of the teacher branch are updated by calculating the average exponent of the student branch, and the training parameters are set for the student branch. The preprocessed pregnancy test images are input into the student branch and the teacher branch respectively, and image features are extracted and corresponding feature vectors are generated. The feature vectors of the teacher branch are centered and sharpened to generate pseudo-labels. Based on pseudo-labels, the consistency loss of feature vectors between student branches and teacher branches is calculated using the cross-entropy function; With the goal of minimizing consistency loss, the student branch is trained iteratively while the teacher branch parameters are updated synchronously. After training, the student branch is taken as the final visual feature extraction unit for feature extraction of pregnancy test images.
[0018] As a preferred technical measure: Step 3: Using a pre-constructed inspection-level information aggregation unit, the semantic relationships and complementary information between one-dimensional image feature vectors are dynamically captured to obtain several inspection-level representation vectors. The method is as follows: Multiple one-dimensional image feature vectors are sorted according to the order in which the inspections occur, and the one-dimensional image feature vectors corresponding to a single inspection are summarized to obtain the same image feature group. When there are m one-dimensional image feature vectors in the same image feature group, the m one-dimensional image feature vectors are aggregated through a self-attention mechanism, and the process is as follows: Obtain m one-dimensional image feature vectors and sort them to establish a vector sequence; introduce a check-level information aggregation classification label at the beginning and end of the vector sequence to form the input sequence of the check-level information aggregation unit; calculate the correlation weight between one-dimensional image feature vectors through a self-attention mechanism to dynamically capture the semantic association and complementary information between vectors, obtain interactive feature data, and realize deep feature interaction; use the encoder layer to perform multi-layer processing on the interactive feature data to strengthen key information and weaken redundant information, and finally extract the vector corresponding to the check-level information aggregation classification label in the output sequence as the unique check-level representation vector for this check; When there is only one one-dimensional image feature vector in the same image feature group, the one-dimensional image feature vector is directly converted into an inspection-level representation vector; The generated inspection-level representation vectors are summarized to obtain several inspection-level representation vectors corresponding to each inspection, thus completing the aggregation from local image features to global inspection information.
[0019] As a preferred technical measure: Step four involves using a pre-constructed case-level multimodal fusion unit to learn the evolutionary relationships between several examination-level representation vectors, simulating the dynamic patterns of pregnancy, and coupling this with clinical text data to generate case-level dynamic change information. The method is as follows: Several examination-level representation vectors are collected, sorted by examination timestamp to form a time series; and case-level information is introduced at the beginning of the time series to aggregate classification labels; Then, based on the actual gestational week intervals of each examination in the clinical follow-up records, a continuous time embedding vector is generated, and these vectors are aligned and superimposed onto the corresponding examination-level representation vectors to obtain an input sequence containing time embeddings, enabling the case-level multimodal fusion unit to capture the dynamic correlation of gestational time. Acquire clinical text data, including pregnant women's medical history text information and structured clinical indicators; The textual information of pregnant women's medical history is integrated with structured clinical indicators, and then a text feature vector is generated after nonlinear mapping and dimension alignment by a multilayer perceptron. The text feature vector is then coupled with the input sequence to generate a case-level input sequence. The case-level input sequence is fed into the case-level multimodal fusion unit. The case-level multimodal fusion unit calculates the correlation weight between the examination-level representation vector and the text feature vector through a global self-attention mechanism, learns the feature evolution association across examinations and the complementary relationship of multimodal information, and thus obtains the case sequence. The vector corresponding to the case-level information aggregation and classification label in the case sequence is extracted as the case-level dynamic change information, realizing the global aggregation of all examination information and clinical text data during pregnancy.
[0020] As a preferred technical measure: The method for generating continuous-time embedding vectors based on the actual gestational age intervals of each examination in the clinical follow-up records is as follows: Extract the actual gestational week corresponding to each examination from the pregnant woman's clinical follow-up records, calculate the gestational week interval between two adjacent examinations, and record the starting gestational week of the first examination as a timeline benchmark. A continuous value mapping method is used to convert the extracted gestational week intervals and initial gestational week into continuous time-series feature values; If the starting gestational week of a certain examination is The interval between the previous inspection and the previous inspection is Then the temporal characteristic value of this inspection is ; The temporal characteristic values of all examinations were normalized to eliminate the influence of differences in gestational week range among different pregnant women. The normalized temporal feature values are input into a preset temporal embedding layer and mapped through nonlinear transformation to a high-dimensional vector with the same dimension as the inspection-level representation vector, which is the final continuous temporal embedding vector. This allows the continuous temporal embedding vector to be directly superimposed and fused with the inspection-level representation vector.
[0021] The following method integrates pregnant women's medical history text information with structured clinical indicators, and generates text feature vectors after multilayer perceptron nonlinear mapping and dimension alignment: Collect the pregnant woman's medical history text information, including her history of premature birth, smoking and alcohol consumption, cervical surgery, and assisted reproductive technology. Obtain structured clinical indicators, including age, weight, height, parity, delivery, progesterone usage, and whether it is a twin pregnancy; and associate them by case identifier to form a complete text dataset for each case. The text information of pregnant women's medical history is processed by word segmentation and removal of stop words, and then transformed into a low-dimensional vector of medical history through word embedding; the structured clinical indicators are normalized to eliminate the differences in dimensions and obtain the indicator vector; the low-dimensional vector of medical history and the indicator vector are then concatenated into a unified medical history input vector. The integrated medical history input vector is fed into a multilayer perceptron, and nonlinear mapping is performed through a hidden layer containing an activation function to capture the complex relationship between text and indicators and generate an intermediate feature vector. Based on the dimension of the inspection-level representation vector, the output layer dimension is set to match it. Then, the intermediate feature vector is processed to output a standardized text feature vector, ensuring that the text feature vector and the inspection-level representation vector can be directly fused.
[0022] This invention effectively integrates clinical text data and ultrasound images, using the text data to capture clinical background information that ultrasound images cannot cover, thus avoiding misjudgments caused by relying solely on images. For example, cervical shortening in images may be physiological (such as normal changes in mid-to-late pregnancy) or pathological (a sign of preterm birth risk), while the history of preterm birth and cervical surgery in the text data can serve as key evidence to help the model distinguish between the two. Similarly, twin pregnancies (structured indicators) can explain the cause of excessive uterine dilation in images, avoiding misjudging normal physiological manifestations of multiple pregnancies as a risk of preterm birth.
[0023] Meanwhile, through the global self-attention mechanism, text features and image features can interact dynamically: progesterone usage in the text (intervention measures) can be associated with the stable cervical morphology in the image, and smoking and drinking history (risk factors) can strengthen the risk weight of placental abnormality in the image. Finally, case-level dynamic change information that integrates dynamic changes in the image and clinical background is generated, realizing the global aggregation of all-dimensional information.
[0024] Furthermore, by combining clinical text data, this invention can be fully adapted to actual clinical scenarios, reducing additional testing costs. The clinical text data all come from routine obstetric follow-up records (such as medical history collection and prenatal examination indicator records), eliminating the need for additional testing items (such as biomarker testing). Its integration with ultrasound imaging allows the model to be directly embedded into existing obstetric examination procedures, avoiding increased medical burden on pregnant women and hospital costs, while covering universal screening of the entire population (rather than just high-risk groups), addressing the pain point of weak clinical adaptability of existing methods and improving the feasibility of technology implementation.
[0025] As a preferred technical measure: Step 5, using a pre-built preterm birth prediction unit to process the dynamically changing information and output the prediction results for the preterm birth subtype, is as follows: Acquire case-level dynamic change information, including all prenatal examination images, clinical text information, and time-related dynamic features, to ensure complete global case information is included; Case-level dynamic change information is fed into the fully connected classification layer of the preterm birth prediction unit. The dynamic change information is mapped to a dimensional result that matches the number of preterm birth subtypes through linear transformation. These subtypes include at least extremely preterm birth, early preterm birth, mid-to-late preterm birth, or full-term delivery. The probability normalization Softmax function is used to process the output dimension results of the fully connected layer, and the linearly transformed dimension results are mapped to the probability space in the interval [0,1] to obtain the confidence scores corresponding to each preterm birth subtype, and the sum of the probabilities of all subtypes is 1; Based on the standards of the Obstetrics and Gynecology Association, the final prediction results are determined according to the category corresponding to the maximum probability of each subtype, including the specific predicted probabilities of extremely preterm birth, early preterm birth, mid-to-late preterm birth, and full-term delivery.
[0026] Unlike existing methods that only predict a binary outcome of preterm birth, this invention can output multi-subtype prediction results, including extremely preterm birth, early preterm birth, mid-to-late preterm birth, and full-term delivery. Furthermore, it outputs confidence scores for each subtype using a Softmax function, providing a quantitative reference for clinical risk levels. For example, intervention programs can be prioritized for cases of extremely preterm birth with high confidence levels, improving the targeted nature of clinical interventions.
[0027] In addition, combining clinical text data can effectively improve the accuracy of subtype differentiation. For example, age (advanced maternal age) and parity (primiparous women) in structured indicators can improve the model's predictive sensitivity for extremely preterm birth; assisted reproductive technology (multiple embryo transfer) in medical history texts can be associated with the risk of mid-to-late-term preterm birth, helping the model to refine the risk characteristics of different subtypes and make the prediction results more in line with the needs of clinical stratified intervention.
[0028] To achieve one of the above objectives, the second technical solution of the present invention is as follows: A method for spontaneous preterm birth prediction based on multimodal data includes the following: Acquire one or more pregnancy examination images from each ultrasound examination of a pregnant woman, and arrange the pregnancy examination images into an examination sequence according to the examination time based on the clinical follow-up records of a pregnant woman; One-dimensional features of pregnancy tests are extracted from the examination sequence images to obtain multiple one-dimensional image feature vectors; When multiple one-dimensional image feature vectors exist in the same inspection, the multiple one-dimensional image feature vectors obtained in the same inspection are aggregated to generate a unique inspection-level representation vector for this inspection; when there is only one one-dimensional image feature vector in the same inspection, the one-dimensional image feature vector is converted into an inspection-level representation vector. By learning the evolutionary relationships between several examination-level representation vectors, simulating the dynamic patterns of pregnancy, and coupling text feature vectors, dynamic change information at the case level is generated, achieving global aggregation of information throughout the entire course of the disease. By processing dynamically changing information, the predicted probability of preterm birth subtype is output, realizing a spontaneous preterm birth prediction based on multimodal data.
[0029] This invention constructs a two-level aggregation mechanism at the examination level and the case level. First, it aggregates local features within a single examination into global information (examination level), and then it achieves global aggregation of multiple examinations and text data throughout the pregnancy (case level). This avoids information redundancy and the submersion of key features, and significantly improves the model's ability to distinguish subtypes such as extremely preterm birth (<28 weeks) and early preterm birth (28-34 weeks), thus meeting the clinical need for precise stratification.
[0030] The workflow design of this invention perfectly matches the actual scenario of pregnancy follow-up (modeling according to the examination time sequence and coupling with clinical follow-up records), without the need for additional testing items (such as biomarker testing), reducing medical costs and the burden on pregnant women. It can be directly integrated into the existing obstetric ultrasound examination process, making it highly practical. It provides a standardized tool for predicting preterm birth for primary hospitals, and can significantly improve the accuracy, robustness, and clinical adaptability of spontaneous preterm birth prediction. It can effectively assist in the early identification of preterm birth risks and the development of stratified intervention strategies, and has important clinical value in reducing the incidence of preterm birth and improving maternal and infant outcomes.
[0031] Furthermore, this invention proposes a sequential input mechanism to address the characteristics of multiple follow-ups during pregnancy and irregular time series, fusing textual features, image features, and clinical indicators at different time points in a structured manner. Through feature splicing across time steps, temporal relationship modeling, and a dynamic weighting mechanism, a more accurate characterization of the pregnancy evolution process is achieved, thereby improving the sensitivity and stability of preterm birth subtype prediction and enhancing the robustness, generalization ability, and clinical usability of the model.
[0032] To achieve one of the above objectives, the third technical solution of the present invention is as follows: A spontaneous preterm birth prediction system based on multimodal data, comprising: One or more processing units; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing units, the one or more processing units implement the above-described method for spontaneous preterm birth prediction based on multimodal data.
[0033] Compared with existing technical solutions, the present invention has the following beneficial effects: This invention, by constructing an image sequence construction unit, a visual feature extraction unit, an examination-level information aggregation unit, a case-level multimodal fusion unit, and a preterm birth prediction unit, can extract one-dimensional features of pregnancy examination from examination sequence images, obtaining multiple one-dimensional image feature vectors. It can also dynamically capture the semantic relationships and complementary information between these one-dimensional image feature vectors, obtaining several examination-level representation vectors, thus completing the aggregation from local image features to global examination information. Furthermore, by learning the evolutionary relationships between these examination-level representation vectors, it simulates the dynamic patterns of pregnancy and couples clinical text data to generate case-level dynamic change information, achieving global aggregation of the entire examination information. Finally, it outputs the prediction result of the preterm birth subtype. This allows for full utilization of sequential images accumulated from multiple prenatal examinations and accurate analysis of the dynamic changes in pregnancy examination images. Based on these dynamic changes, it effectively predicts the risk of preterm birth, significantly enhancing the application value of preterm birth prediction strategies in clinical decision-making. Therefore, it effectively addresses the core pain points of existing preterm birth prediction technologies, such as insufficient utilization of sequential information, limited multimodal data integration, and weak clinical adaptability.
[0034] Furthermore, this invention constructs a two-level aggregation mechanism at the examination level and the case level. First, it aggregates local features within a single examination into global information (examination level), and then it achieves global aggregation of multiple examinations and text data throughout the pregnancy (case level). This avoids information redundancy and the submersion of key features, and significantly improves the model's ability to distinguish subtypes such as extremely preterm birth (<28 weeks) and early preterm birth (28-34 weeks), thus meeting the clinical need for precise stratification.
[0035] The design of this invention perfectly matches the actual scenario of pregnancy follow-up (modeling according to the examination time sequence and coupling with clinical follow-up records), without the need for additional testing items (such as biomarker testing), reducing medical costs and the burden on pregnant women. It can be directly integrated into the existing obstetric ultrasound examination process, making it highly practical. It provides standardized preterm birth prediction tools for primary hospitals, and can significantly improve the accuracy, robustness, and clinical adaptability of spontaneous preterm birth prediction. It can effectively assist in the early identification of preterm birth risks and the development of stratified intervention strategies, and has important clinical value in reducing the incidence of preterm birth and improving maternal and infant outcomes. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the first step of the spontaneous preterm birth prediction method based on multimodal data according to the present invention. Figure 2 This is a schematic diagram of a unit model of the present invention; Figure 3 This is a schematic diagram of a process for self-supervised pre-training of the model of the present invention; Figure 4 This is a schematic diagram of the second process of the spontaneous preterm birth prediction method based on multimodal data according to the present invention. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application. This invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the invention as defined by the claims.
[0038] like Figure 1 , Figure 2 As shown, this is the first specific embodiment of the spontaneous preterm birth prediction method based on multimodal data of the present invention: A method for spontaneous preterm birth prediction based on multimodal data includes the following steps: Step 1: Based on the pre-constructed image sequence generation unit, acquire the pregnancy examination images obtained by a pregnant woman through N ultrasound examinations, and according to the pregnant woman's clinical follow-up records, assemble the pregnancy examination images into an examination sequence image according to the examination time order. Step 2: Using a pre-constructed visual feature extraction unit, extract one-dimensional features of the pregnancy test from the examination sequence images to obtain multiple one-dimensional image feature vectors; Step 3: Using the pre-constructed inspection-level information aggregation unit, the semantic association and complementary information between one-dimensional image feature vectors are dynamically captured to obtain several inspection-level representation vectors, thus completing the aggregation from local image features to global inspection information. Step four: Using a pre-constructed case-level multimodal fusion unit, the evolutionary relationships between several examination-level representation vectors are learned to simulate the dynamic patterns of pregnancy and couple clinical text data to generate case-level dynamic change information, thereby achieving global aggregation of the entire examination information. Step 5: Using a pre-built preterm birth prediction unit, the dynamic change information is processed, and the prediction results of the preterm birth subtype are output to achieve spontaneous preterm birth prediction.
[0039] like Figure 3 , Figure 4 As shown, this is the second specific embodiment of the spontaneous preterm birth prediction method based on multimodal data of the present invention: A spontaneous preterm birth prediction method based on multimodal data mainly consists of two parts: a self-supervised pre-training stage and a preterm birth prediction model training stage. The two stages are functionally related and are used to predict preterm birth subtypes from multimodal, missing, and sequential clinical data.
[0040] In the first stage, a self-supervised pre-trained model based on the Vision Transformer is constructed. This stage uses a large number of unlabeled cervical ultrasound images as input, with images sourced from both transabdominal and transvaginal ultrasound methods, to ensure that the model can learn consistent visual representations across modalities. The input images undergo two different data augmentation processes before being sent to the student branch and the teacher branch, respectively. The core of the self-supervised pre-trained model based on the Vision Transformer in the first stage lies in learning consistent features in ultrasound images through a large number of unlabeled ultrasound images.
[0041] The input data first undergoes targeted preprocessing, including automatic extraction of regions of interest (ROI) using edge detection and morphological operations. Then, median filtering is used for noise reduction and contrast-limited adaptive histogram equalization to enhance image contrast. Finally, normalization is performed to eliminate differences in brightness and gain between different brands of ultrasound equipment.
[0042] The preprocessed images are derived into two inputs using a multi-scale cropping strategy: large global cropps to capture macroscopic anatomical relationships and multiple small local cropps to extract microscopic textures. These two inputs are fed into the teacher and student branches respectively, and high-dimensional feature mapping is performed through a multi-head self-attention mechanism within the VisionTransformer. The student branch network predicts the dynamically distributed target generated by the teacher branch network under momentum updates. Finally, the model outputs a centered cross-modal consistency feature vector, effectively removing visual differences caused by imaging methods and retaining only the core representations in the ultrasound images, providing a highly generalizable feature base for subsequent tasks. Specifically, the process includes the following steps: The image is first divided into fixed-size image patches, which are then mapped into a sequence of one-dimensional image feature vectors (tokens) via a decomposable patch embedding structure. This patch embedding structure reduces the differences in low-level statistical features between transabdominal and transvaginal images, enabling alignment between the two in the encoder. Each patch, after a linear transformation, is mapped into a one-dimensional image feature vector in a unified high-dimensional manifold space, i.e., a one-dimensional image feature vector (token). These sequences of one-dimensional image feature vectors (tokens) are then used as input to the Transformer encoder. This structure, through this mapping mechanism, effectively shields the low-level noise interference introduced by the imaging method, reduces the differences in low-level statistical features between transabdominal and transvaginal images, and allows the encoder to ignore modality specificity and focus on common anatomical structural features, achieving alignment between the two at the representation level.
[0043] The student branch uses a trainable Vision Transformer encoder, and the parameters of the teacher branch are updated by the student branch through an exponential moving average.
[0044] The two branches each generate one-dimensional image feature vectors (Tokens). The model unit is trained by calculating the consistency loss between student and teacher features, ensuring that the student branch learns stable representations under different enhancement conditions. Specifically, this loss uses a cross-entropy function to measure the difference in probability distribution between the one-dimensional image feature vectors (Tokens) output by the two branches. The teacher branch output is centered and sharpened to generate high-quality pseudo-labels, which then guide the student branch to learn anatomical structural features that are highly consistent with the teacher branch and have semantic invariance under complex modality transformations.
[0045] Assume the feature vector output by the student branch is The feature vector output by the teacher branch is The method for transforming probability distributions is as follows: First, the features are transformed into a probability distribution using the Softmax function. and Its expression is as follows: in: is the temperature parameter, used to control the smoothness of the distribution; c is the centering vector of the teacher branch (updated by the historical mean), used to prevent the model from outputting a single class; K is the number of feature dimensions after linear layer mapping; and i is the i-th index.
[0046] The cross-entropy formula is typically used to calculate the consistency loss, and its calculation formula is as follows: This loss function forces the output of the student branch. As close as possible to the pseudo-labels generated by the teacher's branch .
[0047] The visual transformation unit and self-supervised pre-trained model obtained after this training stage can be used as visual feature extraction units in the second stage for image feature extraction, providing a robust and consistent feature space.
[0048] In the second phase, a multimodal sequential model for predicting preterm birth subtypes is constructed, which includes the following: Step 1: Based on the pregnant woman's clinical follow-up records, all ultrasound examinations are arranged into an examination sequence according to their actual timeline. Each examination may contain one or more images, and transabdominal and transvaginal modalities are allowed to be mixed. Each image is processed by the visual feature extraction unit obtained in the first stage to extract a one-dimensional image feature vector (Token) related to the image features.
[0049] Step 2: The one-dimensional image feature vectors (Tokens) of multiple images from the same inspection are fed into the inspection-level information aggregation unit (Transformer) for aggregation, generating a unique inspection-level representation vector. This representation vector contains the image information of the inspection. The specific method is as follows: Each set of inspection image sequences is transformed into a one-dimensional image feature vector by a visual feature extraction unit, and a learnable global Check_ClsToken is introduced at the beginning.
[0050] After the one-dimensional image feature vector is fed into the inspection-level information aggregation unit (Transformer), it dynamically captures the semantic relationships and complementary information between images through a self-attention mechanism, achieving deep feature fusion. Finally, the vector corresponding to the CheckToken in the output sequence is extracted as the unique representation vector for that inspection, i.e., the inspection-level representation vector, thus completing the aggregation from local image features to global inspection information. When an image is missing in an inspection, the structure of this invention allows skipping that inspection without affecting the integrity of the overall sequence.
[0051] Step 3: Multiple examination-level representation vectors are input into the case-level multimodal fusion unit (Transformer) in chronological order of examination occurrence. To model the dynamic patterns of cervical changes during pregnancy, this invention introduces a time encoding method within the case-level multimodal fusion unit (Transformer), generating continuous temporal embeddings based on actual clinical examination intervals, enabling the model to learn the correlations between different gestational weeks. Multiple examination-level feature representations are arranged chronologically, with a learnable global Case_ClsToken added at the beginning of the sequence, and are input into the case-level multimodal fusion unit (Transformer). To model the dynamic patterns of pregnancy, the model generates continuous temporal embeddings based on the actual gestational week of examination and superimposes them onto the corresponding examination-level features. Through a self-attention mechanism, the model learns the evolutionary correlations between different examination points. Finally, the output of Case_ClsToken is extracted as the dynamic change information of the case, achieving global aggregation of information throughout the entire disease course.
[0052] In this embodiment, the temporal embedding vector is a feature vector used to quantify different examination time intervals during pregnancy and provide temporal information for the model. Its core function is to allow the case-level multimodal fusion unit to capture the dynamic evolution of the examination-level representation vector with gestational week, simulate the physiological changes during pregnancy, and solve the problem of difficulty in modeling unordered temporal data. The method for generating continuous temporal embedding vectors needs to be combined with gestational week information from clinical follow-up records and is implemented according to the following process: Extract the actual gestational week corresponding to each examination from the clinical follow-up records of pregnant women (e.g., examinations at 12, 16, and 20 weeks), calculate the gestational week interval between two adjacent examinations (e.g., 16-12=4 weeks, 20-16=4 weeks), and record the starting gestational week of the first examination (as a timeline benchmark).
[0053] Continuous-time coding uses a continuous value mapping method to transform the extracted gestational week intervals and initial gestational week into continuous numerical features (as opposed to discrete-time coding), which includes the following: If the starting gestational week of a certain examination is The interval between the previous inspection and the previous inspection is Then the temporal characteristic value of this inspection is (For example, if the first check-up is at 12 weeks, the result is 12; if the second check-up is 4 weeks later, the result is 16). The temporal feature values of all examinations were normalized (e.g., mapped to the [0,1] interval) to eliminate the influence of differences in gestational week range among different pregnant women.
[0054] High-dimensional vector mapping takes normalized continuous temporal feature values, inputs them into a pre-defined temporal embedding layer (usually a fully connected network or a sinusoidal positional encoding module), and maps them through a nonlinear transformation to a high-dimensional vector with the same dimension as the inspection-level representation vector, i.e., the final continuous temporal embedding vector. For example, if the inspection-level representation vector has a dimension of 512, then the temporal embedding vector also needs to be mapped to 512 dimensions to ensure that it can be directly superimposed and fused with the inspection-level representation vector in the subsequent process.
[0055] In this embodiment, the case-level dynamic change information token output by the case-level multimodal fusion unit (Transformer) represents the overall image features of the case. Simultaneously, the case's medical history textual information (such as history of premature birth, smoking and alcohol consumption, cervical surgery, and whether assisted reproduction was performed) and structured clinical indicators (such as age, weight, height, parity, delivery, progesterone use, and whether twins were conceived) are integrated into a feature vector. This vector is then nonlinearly mapped and dimensionally aligned using a multilayer perceptron (MLP) and directly encoded as textual feature data tokens containing prior clinical knowledge. These textual feature data tokens, along with the one-dimensional image feature vectors (Tokens) of the image, participate in attention calculations within the fusion layer of the case-level multimodal fusion unit (Transformer) to jointly constitute the input sequence, interacting through a global self-attention mechanism. Under this mechanism, each image feature vector (Token) absorbs information from other modalities according to the relevance weights: textual feature data (Token) about clinical indicators provides key risk background for image features, while one-dimensional image feature vectors (Token) about images supplement clinical data with intuitive anatomical evidence, thereby realizing deep coupling and joint modeling of multimodal information at the semantic level.
[0056] Step 4: Finally, the fused case-level dynamic change information is input into the preterm birth prediction unit. The preterm birth prediction unit has a fully connected classification layer that can output the predicted probability of preterm birth subtypes, including categories such as extremely preterm, preterm, late preterm, and full-term delivery. The specific process is as follows: The fused case-level dynamic change information is fed into a fully connected classification layer. After linear transformation, the output is mapped to the probability space using the Softmax function to obtain the confidence score for each delivery category. Following the authoritative standards of the American College of Obstetricians and Gynecologists (ACOG), the samples are classified into four categories based on the index corresponding to the maximum value in the predicted probability distribution: extremely preterm labor (gestational age at delivery <28 weeks), early preterm labor (gestational age at delivery >=28 weeks but <34 weeks), mid-to-late preterm labor (gestational age at delivery >=34 weeks but <37 weeks), and term delivery (gestational age at delivery >=37 weeks).
[0057] To improve the interpretability of the model, this invention provides a Grad-CAM-based visualization analysis method during the inference phase.
[0058] Specifically, after the case-level multimodal fusion unit (Transformer) generates prediction results, the activation weights of the specified level of the visual feature extraction unit are calculated by backpropagating the gradients related to the predicted category, and a saliency heatmap is generated. To ensure that the visualization results accurately correspond back to the original ultrasound image, this invention incorporates a one-dimensional image feature vector (Token) to image patch inverse mapping during the heatmap generation process, mapping the activated regions back to their actual spatial locations in the image. The resulting heatmap highlights the key cervical structural regions that the model focuses on during prediction, providing auxiliary analytical support for clinicians.
[0059] To enhance the interpretability of preterm birth prediction models based on the Transformer architecture, this invention constructs a visual Gradient Weighted Class Activation Mapping (Grad-CAM) analysis method during the inference phase. This method goes beyond general visualization; it designs a reverse mapping and spatial location mapping mechanism from a "one-dimensional semantic space" to a "two-dimensional anatomical space" specifically for the serialized features of one-dimensional image feature vectors (Tokens) in cervical ultrasound images. The specific algorithm processing flow is as follows: Gradient backpropagation and activation weight calculation: First, define the model's prediction score for a certain category as... (That is, the Logits value of the model output layer before Softmax). The output of the last layer of the Visual Encoder is selected as the target feature layer, denoted as A. A contains K feature maps, and the feature map of the k-th channel is represented as... After inference is complete, the system performs a gradient backpropagation, but does not update the model parameters; instead, it calculates the predicted score. Relative to feature layer The gradient of each element (a one-dimensional image feature vector (Token) or pixel) .
[0060] Based on this gradient, the importance weight of the k-th feature channel to the target class c is calculated using global average pooling. The specific mathematical model for its processing is as follows: Where Z is the total number of elements in the feature layer, and (i,j) is the index of the feature unit.
[0061] The reverse mapping of one-dimensional image feature vectors (tokens) from sequence to grid is employed in this invention because the feature layer A is represented in data form as a one-dimensional sequence of one-dimensional image feature vectors (tokens), rather than a traditional two-dimensional image grid. To locate key anatomical structures of the cervix, reverse mapping is required. The system first performs a weighted summation of the feature maps from all K channels to obtain a category-specific localization map. Its expression is as follows: The ReLU function is used to filter out negative activations (i.e., remove regions that negatively impact the prediction results). Subsequently, a one-dimensional image feature vector (Token) inverse mapping is performed: The input image is divided into P×P image patches (e.g., patches), and the length of the one-dimensional image feature vector (Token) sequence processed by the model (after removing the CLS classification one-dimensional image feature vector (Token)) is N. The system maps the one-dimensional image feature vector (Token) back to two-dimensional grid coordinates (h', w') based on its index nin[1, N] in the original sequence, thus transforming the one-dimensional vector... Reshape into a two-dimensional low-resolution heatmap Its dimension is sqrt{N} × sqrt{N}.
[0062] Spatial Position Mapping: To ensure that the heatmap accurately covers the original cervical ultrasound image I (size H × W), spatial position mapping is required. The resolution of the image is much lower than that of the original image (e.g., L_{grid} is 14×14, while I is 224×224). This invention employs a **bilinear interpolation** algorithm to... Projected onto the spatial coordinate system of the original image.
[0063] The specific space mapping function is defined as follows: Here, Upsample represents the spatial dimension enhancement function, which aims to enhance the low-resolution feature matrix. The feature matrix can be extended to a higher resolution using a specific mathematical interpolation algorithm. The original image size is represented by a binary tuple (H, W), which is a reference resolution for the target mapping. Its value is strictly equal to the height H and width W of the original input image I.
[0064] Ultimately, the generated After normalization, the data is overlaid on the original ultrasound image in the form of a pseudo-color heatmap. The highlighted areas represent the cervical morphological features (such as the internal cervical os and the length of the cervical canal) that the model focuses on when predicting the risk of preterm birth, thus realizing the visualization transformation from "data variables" to "clinical anatomical semantics".
[0065] This invention employs a self-supervised pre-trained visual feature extraction unit to extract features from cervical ultrasound images, supporting both transabdominal and transvaginal ultrasound image inputs simultaneously. During the encoding stage, the image is divided into fixed-size image patches, and a sequence of one-dimensional image feature vectors (Tokens) is generated through decomposable image patch embedding, thereby achieving cross-modal alignment and low-level feature consistency. Simultaneously, this invention encodes clinical text information, fusing the text-based one-dimensional image feature vectors (Tokens) with the image-based one-dimensional image feature vectors (Tokens) during the sequence modeling stage to achieve multimodal joint representation.
[0066] This invention can extract robust and generalizable features under different acquisition methods (i.e., transvaginal ultrasound and transabdominal ultrasound) and image quality conditions, improving the accuracy and reliability of preterm birth prediction. Multimodal fusion enhances the comprehensive understanding of patient history and imaging information, reduces reliance on large amounts of labeled data, and improves system applicability and clinical scalability.
[0067] Furthermore, in the inspection-level aggregation stage, this invention effectively integrates the features of multiple images from the same inspection to generate an inspection-level representation, which is used to express the overall information of a single inspection. Through aggregation operations, the model can capture complementary information and local change features between images within the same inspection, ensuring high-quality feature representation for each inspection. This improves the efficiency and consistency of multi-image information integration, making the inspection-level features more representative in subsequent sequence modeling, while reducing computational complexity and ensuring the efficiency of the training and inference processes.
[0068] In the case-level sequence modeling stage, this invention inputs examination-level representations arranged chronologically into the sequence model to characterize the dynamic changes in cervical morphology over time during pregnancy. Through temporal coding and attention mechanisms, the model can capture temporal dependencies and trend changes across examinations, while also handling cases with different gestational week intervals or missing examinations. This achieves complete case-level feature integration, thereby enhancing the invention's understanding and predictive capabilities of sequential data, making preterm birth subtype prediction more accurate and reliable. Furthermore, it adapts to practical application scenarios with varying clinical follow-up frequencies and data completeness, improving the system's adaptability and practicality.
[0069] This invention provides a visualization and interpretability analysis method to saliency-visualize key image regions of interest during the model's prediction process. Through spatial mapping of heatmaps with original images, clinicians can intuitively understand the model's decision-making basis and perform auxiliary diagnosis or result verification. This enhances the interpretability and reliability of the model's predictions, provides intuitive reference for clinical practice, helps identify potential model biases or anomalies, supports model optimization and clinical research, and enhances the model's usability in real-world scenarios. Other embodiments of this invention's spontaneous preterm birth prediction method based on multimodal data include: A spontaneous preterm birth prediction method and a preterm birth subtype prediction system based on multimodal data can be implemented through various technical solutions. While the core functions remain unchanged, there are several alternative methods in the specific implementation.
[0070] During the self-supervised pre-training phase, the visual feature encoder can use other self-supervised learning methods to replace DINOv3, such as SimCLR, MoCo, or BYOL, to obtain robust image representations by enhancing and aligning unlabeled ultrasound images.
[0071] In the inspection-level aggregation stage, the feature aggregation structure that originally used the Transformer can be replaced by convolution or pooling operations. By performing convolution or global pooling on the features of multiple images in the same inspection, the inspection-level representation is generated, thereby reducing computational complexity while retaining the ability to integrate information from multiple images.
[0072] In the case-level sequence modeling stage, the original case-level multimodal fusion unit (Transformer) can be replaced with other sequence modelers, such as LSTM, GRU or Temporal Convolution Network (TCN), to achieve the integration of temporal information across examination sequences and capture the pattern of cervical morphological changes during pregnancy.
[0073] In the visualization and interpretability module, the original Grad-CAM method can be replaced with other attention or gradient analysis techniques, such as Guided Backpropagation, Integrated Gradients, or Transformer AttentionRollout, to generate saliency maps that highlight the regions of interest of the model, thereby enabling image visualization analysis.
[0074] The above-mentioned alternative technical solutions can be used individually or in combination. All of them can achieve the prediction of preterm birth subtypes without changing the core function of the present invention, while taking into account multimodal, sequential data processing and interpretability analysis.
[0075] A specific embodiment of the method of the present invention: Using the method of this invention, a preterm birth risk assessment method based on multimodal deep learning is constructed, and a preterm birth subtype risk assessment system based on multimodal self-supervised learning is formed. The specific experimental environment and implementation steps are as follows: S1. Prepare initial data and preprocess the data, including the following: This embodiment used 110,388 ultrasound images (involving observations of organs or tissues such as the placenta and amniotic fluid) from the same hospital for the first stage of self-supervised pre-training. The clinical application validation phase included 3,250 cases of pregnant women, with 2,272 cases in the training set (containing approximately 15,000 cervical images) and 978 cases in the test set.
[0076] All cases included complete multimodal data, including imaging sequences (imaging methods included transabdominal and transvaginal ultrasound) and accompanying text information (including: age, weight, height, parity, history of progesterone use, history of preterm birth, history of smoking and alcohol consumption, history of cervical surgery, whether it was a twin pregnancy and assisted reproductive technology).
[0077] S2. Training the unit model, which includes the following: Phase 1: Self-supervised pre-training based on the visual transformation unit (Transformer); Phase 2: Multimodal sequential modeling and temporal evolution extraction; Phase 3: Classification Decision and Performance Verification - Classification Judgment.
[0078] S3. Determining and predicting the category of preterm birth, including the following: The model outputs a four-class probability distribution through a fully connected layer and a softmax function. The classification criteria strictly follow the ACOG (American College of Obstetricians and Gynecologists) guidelines, determining the class of a sample based on the highest probability value: Extremely preterm births (<28 weeks): 57 cases in the training set and 44 cases in the test set; Preterm births (28-34 weeks): 131 cases in the training set and 74 cases in the test set; Late preterm birth (34-37 weeks): 296 cases in the training set and 160 cases in the test set; Full-term delivery (≥37 weeks): 1789 cases in the training set and 698 cases in the test set.
[0079] S4. Evaluate the experimental results and performance, including the following: On the internal test set (978 cases), the classification performance of this invention is as follows (AUC value and 95% confidence interval): Premature birth prediction: AUC = 0.85 (0.81-0.89); Early preterm birth prediction: AUC = 0.81 (0.75-0.87); Preterm birth prediction in the mid-to-late stages: AUC = 0.78 (0.72-0.84); Prediction of full-term delivery: AUC = 0.84 (0.80-0.88); Overall average performance: AUC = 0.82 (0.77-0.87).
[0080] AUC (Area Under the ROC Curve) is a core evaluation metric for measuring the predictive performance of binary classification models, with a value range of [0,1].
[0081] S5. Perform visualization analysis, which includes the following: In this embodiment, the image regions with the highest model attention are mapped using the attention map of the visual transformation unit (Transformer). The visualization results show that the model can accurately locate changes in the morphology of the internal cervical os and areas of cervical canal shortening, achieving interpretability verification from feature extraction to risk assessment.
[0082] Therefore, this invention can extract robust representations from multimodal clinical data. Through self-supervised pre-training, the model can acquire ultrasound feature extraction capabilities from a large amount of unlabeled ultrasound image data. The multimodal data includes textual information (such as medical records, examination records, etc.) and image information, wherein the image input simultaneously covers transabdominal (TAUS) and transvaginal (TVUS) cervical ultrasound images. This invention has robust handling capabilities for arbitrary modal missing data, and can still output stable predictions even when a single modality is missing or the information is incomplete.
[0083] Furthermore, this invention can model and stitch features from sequential follow-up data, improving its adaptability to real clinical processes. Addressing the characteristics of multiple follow-ups during pregnancy and irregular time series, this invention proposes a sequential input mechanism that fuses textual features, image features, and clinical indicators from different time points in a structured manner. Through cross-time step feature stitching, temporal relationship modeling, and a dynamic weighting mechanism, a more accurate characterization of the pregnancy evolution process is achieved, thereby improving the sensitivity and stability of preterm birth subtype prediction.
[0084] This invention further constructs a model interpretability analysis tool, including visualization of salient regions in images, text contribution analysis, and multimodal collaborative feature display, so that the basis for model prediction is presented intuitively, making it easier for doctors to understand the model's decision-making logic, thereby supporting clinical researchers to explore preterm birth-related mechanisms and risk patterns from different dimensions.
[0085] In summary, this invention aims to construct a preterm birth subtype prediction system that more closely reflects real clinical procedures through multimodal, self-supervised, and interpretable sequential prediction techniques, thereby improving the model's robustness, generalization ability, and clinical usability.
[0086] A server embodiment applying the method of the present invention: A server comprising: One or more processing units; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing units, the one or more processing units implement the above-described method for spontaneous preterm birth prediction based on multimodal data.
[0087] The storage device can be internal memory, external memory, cache memory, or other special memory. The processing unit has signal processing capabilities and can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, an off-the-shelf programmable gate array, or other programmable logic device.
[0088] An embodiment of a device applying the method of the present invention: An electronic device is provided with a computer-readable storage medium on which a computer program is stored. When the program is executed by a processing unit, it implements the above-described method for spontaneous preterm birth prediction based on multimodal data.
[0089] Computer-readable storage media refers to physical carriers capable of storing computer-recognizable data, instructions, or programs. These media must meet the core characteristic of being "readable by a computer" (i.e., the data exists in the form of electrical, magnetic, or optical signals and can be converted into binary information that a computer can process through appropriate devices). The physical carrier can be a magnetic storage medium, optical storage medium, semiconductor storage medium, or other storage media.
[0090] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0091] The unit or model in this application is an object that constitutes an objective description of form and structure through physical or virtual representation. The object is not the same as a physical object, and is not limited to physical or virtual. It can be a data processing function, software program, processing mode, usage method, operation mode, workflow, application process, electronic hardware, circuit module, processing system, system imitation or simulation object.
[0092] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features within the scope of the technology disclosed in the present invention; and these modifications or substitutions will not cause the substance of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any modifications or equivalent substitutions that do not deviate from the spirit and scope of the present invention should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for spontaneous preterm birth prediction based on multimodal data, characterized in that: Includes the following steps: Step 1: Based on the pre-constructed image sequence generation unit, acquire the pregnancy examination images obtained by a pregnant woman through N ultrasound examinations, and according to the pregnant woman's clinical follow-up records, assemble the pregnancy examination images into an examination sequence image according to the examination time order. Step 2: Using a pre-constructed visual feature extraction unit, extract one-dimensional features of the pregnancy test from the examination sequence images to obtain multiple one-dimensional image feature vectors; Step 3: Using the pre-constructed inspection-level information aggregation unit, the semantic association and complementary information between one-dimensional image feature vectors are dynamically captured to obtain several inspection-level representation vectors, thus completing the aggregation from local image features to global inspection information. Step four: Using a pre-constructed case-level multimodal fusion unit, the evolutionary relationships between several examination-level representation vectors are learned to simulate the dynamic patterns of pregnancy and couple clinical text data to generate case-level dynamic change information, thereby achieving global aggregation of the entire examination information. Step 5: Using a pre-built preterm birth prediction unit, the dynamic change information is processed, and the prediction results of the preterm birth subtype are output to achieve spontaneous preterm birth prediction.
2. The spontaneous preterm birth prediction method based on multimodal data as described in claim 1, characterized in that: Step 1: Based on the pre-built image sequence generation unit, acquire pregnancy examination images obtained from N ultrasound examinations of a pregnant woman, and according to the pregnant woman's clinical follow-up records, assemble the pregnancy examination images into an examination sequence image according to the examination time order as follows: Acquire pregnancy examination images of a pregnant woman obtained through N ultrasound examinations, including transabdominal ultrasound images and / or transvaginal ultrasound images and / or cervical related ultrasound images, and associate them with the pregnant woman's clinical follow-up records, which include the timestamp and examination type of each examination. Median filtering was used to remove noise, and the pregnancy test images were preprocessed and the contrast was equalized to obtain pregnancy test images with the same brightness and small gain differences. Based on the examination timestamps in the clinical follow-up records, all pregnancy examination images corresponding to a single examination are summarized to obtain a pregnancy examination image group; The pregnancy examination images are sorted according to the order in which the examinations were conducted, forming a time-series image sequence.
3. The spontaneous preterm birth prediction method based on multimodal data as described in claim 1, characterized in that: Step two, using a pre-constructed visual feature extraction unit, extracts one-dimensional features from the examination sequence images to obtain multiple one-dimensional image feature vectors. The method is as follows: Each pregnancy test image in the sequence of images is divided into image blocks of fixed size to unify the image input dimensions; The pre-trained visual feature extraction unit is invoked to extract features from each image block to obtain pregnancy test features. The extracted pregnancy test features are then linearly transformed and mapped to a one-dimensional space to generate corresponding one-dimensional pregnancy test feature fragments. The one-dimensional feature segments of all image blocks are spliced together in their original spatial order to form a one-dimensional image feature vector of the pregnancy examination image, thus realizing the conversion from two-dimensional image to one-dimensional feature. The above process is used to process all pregnancy test images in the sequence of images to obtain multiple one-dimensional image feature vectors corresponding to the number of pregnancy test images, ensuring that each one-dimensional image feature vector retains the core feature information of the pregnancy test image.
4. The spontaneous preterm birth prediction method based on multimodal data as described in claim 3, characterized in that: The method for self-supervised training of the visual feature extraction unit is as follows: Collect unlabeled pregnancy examination images, including transabdominal and / or transvaginal ultrasound images related to the placenta and amniotic fluid, and associate them with the pregnant woman's clinical follow-up records, which include timestamps and examination types for each examination. Median filtering was used to remove noise, and the pregnancy test images were preprocessed and the contrast was equalized to obtain pregnancy test images with the same brightness and small gain differences. Then, a multi-scale cropping method is used to crop the pregnancy examination images to generate large-size global slices and small-size local slices, which are used as training inputs; Large-size global slices are used to capture macroscopic anatomical relationships, while small-size local slices are used to uncover microscopic textures. Based on a self-supervised pre-training algorithm, a two-branch model with student and teacher branches is constructed, both of which use visual converters. The parameters of the teacher branch are updated by calculating the average exponent of the student branch, and the training parameters of the student branch are set. The preprocessed pregnancy test images are input into the student branch and the teacher branch respectively, and image features are extracted and corresponding feature vectors are generated. The feature vectors of the teacher branch are centered and sharpened to generate pseudo-labels. Based on pseudo-labels, the consistency loss of feature vectors between student branches and teacher branches is calculated using the cross-entropy function; With the goal of minimizing consistency loss, the student branch is trained iteratively while the teacher branch parameters are updated synchronously. After training, the student branch is taken as the final visual feature extraction unit for feature extraction of pregnancy test images.
5. The spontaneous preterm birth prediction method based on multimodal data as described in claim 1, characterized in that: Step 3: Using a pre-constructed inspection-level information aggregation unit, the semantic relationships and complementary information between one-dimensional image feature vectors are dynamically captured to obtain several inspection-level representation vectors. The method is as follows: Multiple one-dimensional image feature vectors are sorted according to the order in which the inspections occur, and the one-dimensional image feature vectors corresponding to a single inspection are summarized to obtain the same image feature group. When there are m one-dimensional image feature vectors in the same image feature group, the m one-dimensional image feature vectors are aggregated through a self-attention mechanism, and the process is as follows: Obtain m one-dimensional image feature vectors and sort them to establish a vector sequence; introduce check-level information aggregation and classification labels at the beginning and end of the vector sequence to form the input sequence of check-level information aggregation units; calculate the correlation weights between one-dimensional image feature vectors through a self-attention mechanism, dynamically capture the semantic association and complementary information between vectors, obtain interactive feature data, and realize deep feature interaction. The encoder layer is used to process the interactive feature data in multiple layers, strengthen the key information and weaken the redundant information. Finally, the vector corresponding to the aggregated classification label of the inspection level information in the output sequence is extracted as the unique inspection level representation vector of this inspection. When there is only one one-dimensional image feature vector in the same image feature group, the one-dimensional image feature vector is directly converted into an inspection-level representation vector; The generated inspection-level representation vectors are summarized to obtain several inspection-level representation vectors corresponding to each inspection, thus completing the aggregation from local image features to global inspection information.
6. The spontaneous preterm birth prediction method based on multimodal data as described in claim 1, characterized in that: Step four involves using a pre-constructed case-level multimodal fusion unit to learn the evolutionary relationships between several examination-level representation vectors, simulating the dynamic patterns of pregnancy, and coupling this with clinical text data to generate case-level dynamic change information. The method is as follows: Several examination-level representation vectors are collected, sorted by examination timestamp to form a time series; and case-level information is introduced at the beginning of the time series to aggregate classification labels; Then, based on the actual gestational week intervals of each examination in the clinical follow-up records, a continuous time embedding vector is generated, and these vectors are aligned and superimposed onto the corresponding examination-level representation vectors to obtain an input sequence containing time embeddings, enabling the case-level multimodal fusion unit to capture the dynamic correlation of gestational time. Acquire clinical text data, including pregnant women's medical history text information and structured clinical indicators; The textual information of pregnant women's medical history is integrated with structured clinical indicators, and then a text feature vector is generated after nonlinear mapping and dimension alignment by a multilayer perceptron. The text feature vector is then coupled with the input sequence to generate a case-level input sequence. The case-level input sequence is fed into the case-level multimodal fusion unit. The case-level multimodal fusion unit calculates the correlation weight between the examination-level representation vector and the text feature vector through a global self-attention mechanism, learns the feature evolution association across examinations and the complementary relationship of multimodal information, and thus obtains the case sequence. The vector corresponding to the case-level information aggregation and classification label in the case sequence is extracted as the case-level dynamic change information, realizing the global aggregation of all examination information and clinical text data during pregnancy.
7. The spontaneous preterm birth prediction method based on multimodal data as described in claim 6, characterized in that: The method for generating continuous-time embedding vectors based on the actual gestational age intervals of each examination in the clinical follow-up records is as follows: Extract the actual gestational week corresponding to each examination from the pregnant woman's clinical follow-up records, calculate the gestational week interval between two adjacent examinations, and record the starting gestational week of the first examination as a timeline benchmark. A continuous value mapping method is used to convert the extracted gestational week intervals and initial gestational week into continuous time-series feature values; If the starting gestational week of a certain examination is The interval between the previous inspection and the previous inspection is Then the temporal characteristic value of this inspection is ; The temporal characteristic values of all examinations were normalized to eliminate the influence of differences in gestational week range among different pregnant women. The normalized temporal feature values are input into a preset temporal embedding layer and mapped through nonlinear transformation to a high-dimensional vector with the same dimension as the inspection-level representation vector, which is the final continuous temporal embedding vector, so that the continuous temporal embedding vector can be directly superimposed and fused with the inspection-level representation vector. The following method integrates pregnant women's medical history text information with structured clinical indicators, and generates text feature vectors after multilayer perceptron nonlinear mapping and dimension alignment: Collect pregnant women's medical history text information, including past history of premature birth, history of smoking and drinking, history of cervical surgery, and assisted reproductive technology. Obtain structured clinical indicators, including age, weight, height, parity, delivery, progesterone usage, and whether it is a twin pregnancy; and associate them by case identifier to form a complete text dataset for each case. The text information of pregnant women's medical history is processed by word segmentation and removal of stop words, and then transformed into a low-dimensional vector of medical history through word embedding; the structured clinical indicators are normalized to eliminate the differences in dimensions and obtain the indicator vector; the low-dimensional vector of medical history and the indicator vector are then concatenated into a unified medical history input vector. The integrated medical history input vector is fed into a multilayer perceptron, and nonlinear mapping is performed through a hidden layer containing an activation function to capture the complex relationship between text and indicators and generate an intermediate feature vector. Based on the dimension of the inspection-level representation vector, the output layer dimension is set to match it. Then, the intermediate feature vector is processed to output a standardized text feature vector, ensuring that the text feature vector and the inspection-level representation vector can be directly fused.
8. The spontaneous preterm birth prediction method based on multimodal data as described in claim 1, characterized in that: Step 5, using a pre-built preterm birth prediction unit to process the dynamically changing information and output the prediction results for the preterm birth subtype, is as follows: Acquire case-level dynamic change information, including all prenatal examination images, clinical text information, and time-related dynamic features, to ensure complete global case information is included; Case-level dynamic change information is fed into the fully connected classification layer of the preterm birth prediction unit. The dynamic change information is mapped to a dimensional result that matches the number of preterm birth subtypes through linear transformation. These subtypes include at least extremely preterm birth, early preterm birth, mid-to-late preterm birth, or full-term delivery. The probability normalization Softmax function is used to process the output dimension results of the fully connected layer, and the linearly transformed dimension results are mapped to the probability space in the interval [0,1] to obtain the confidence scores corresponding to each preterm birth subtype, and the sum of the probabilities of all subtypes is 1; Based on the standards of the Obstetrics and Gynecology Association, the final prediction results are determined according to the category corresponding to the maximum probability of each subtype, including the specific predicted probabilities of extremely preterm birth, early preterm birth, mid-to-late preterm birth, and full-term delivery.
9. A method for spontaneous preterm birth prediction based on multimodal data, characterized in that: Includes the following: Acquire one or more pregnancy examination images from each ultrasound examination of a pregnant woman, and arrange the pregnancy examination images into an examination sequence according to the examination time based on the clinical follow-up records of a pregnant woman; One-dimensional features of pregnancy tests are extracted from the examination sequence images to obtain multiple one-dimensional image feature vectors; When multiple one-dimensional image feature vectors exist in the same inspection, the multiple one-dimensional image feature vectors obtained in the same inspection are aggregated to generate a unique inspection-level representation vector for this inspection; when there is only one one-dimensional image feature vector in the same inspection, the one-dimensional image feature vector is converted into an inspection-level representation vector. By learning the evolutionary relationships between several examination-level representation vectors, simulating the dynamic patterns of pregnancy, and coupling text feature vectors, dynamic change information at the case level is generated, achieving global aggregation of information throughout the entire course of the disease. By processing dynamically changing information, the predicted probability of preterm birth subtype is output, realizing a spontaneous preterm birth prediction based on multimodal data.
10. A spontaneous preterm birth prediction system based on multimodal data, characterized in that: It includes: One or more processing units; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing units, the one or more processing units implement a spontaneous preterm birth prediction method based on multimodal data as described in any one of claims 1-9.