Fetal state dynamic monitoring method and system based on multi-modal deep learning
Through multimodal deep learning method, combined with fetal heart rate signal and uterine contraction signal, feature extraction and timing analysis are carried out, which solves the problem of accuracy and low efficiency of fetal status monitoring in the prior art, and realizes accurate monitoring of fetal status and timely discovery of abnormal status.
Patent Information
- Application Number
- CN202510409338.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The existing fetal status monitoring methods have problems such as uneven manual judgment experience, low efficiency, difficulty in real-time continuous monitoring, and multimodal deep learning methods fail to fully explore multimodal data information.
The fetal state monitoring method based on multimodal deep learning is adopted, and fetal heart rate signal and uterine contraction signal are collected, sliding window segmentation, denoising processing, time-frequency analysis, cross attention feature fusion and timing feature extraction of Transformer encoder, and finally fetal state classification is realized through a fully connected classification layer.
It realizes accurate and dynamic monitoring of fetal status, reduces manual monitoring errors, improves monitoring efficiency and accuracy, and can detect abnormal status in a timely manner to ensure fetal safety.
Smart Images

Figure CN120323948A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fetal electronic monitoring, and particularly to a method and system for dynamically monitoring fetal status based on multimodal deep learning. Background Art
[0002] As an important fetal monitoring means in clinical practice, cardiotocography (CTG) evaluates fetal health by analyzing the fetal heart rate (FHR) and the uterine contraction signal (UC) of the mother. However, there are many significant drawbacks in manually judging fetal status. The experience of medical staff in interpreting CTG signals varies. In the face of complex CTG signals, medical staff with insufficient experience are difficult to accurately identify the fetal status information contained therein. Different medical staff may make different interpretations of the same CTG signal, and medical staff with insufficient experience may not be able to accurately judge whether the fetus is in potential danger, thus delaying the intervention time.
[0003] At the same time, manual judgment is inefficient and it is difficult to achieve real-time and continuous monitoring of fetal status. In recent years, although computer-aided analysis systems have been applied, there are still deficiencies. For example, existing systems are mostly based on limited fetal monitoring guidelines and only consider some morphological parameters such as the fetal heart rate baseline, accelerations and decelerations, and baseline variability. Traditional machine learning algorithms still require manual feature extraction, which is a cumbersome process and prone to information loss. Most deep learning-based methods only use a single signal such as the fetal heart rate for analysis and cannot comprehensively reflect fetal status. In recent years, some multimodal deep learning strategies have also been explored and studied, but most of the fusion methods are simple and cannot fully exploit the information in multimodal data.
[0004] Therefore, there is an urgent need for a more effective method and system for dynamically monitoring fetal status to improve the current situation. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method and system for dynamically monitoring fetal status based on multimodal deep learning, which uses the fetal heart rate signal and the uterine contraction signal to monitor fetal status.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] The method for dynamically monitoring fetal status based on multimodal deep learning provided by the present invention includes the following steps:
[0008] S1 Data acquisition: Acquire the fetal heart rate signal FHR and the uterine contraction signal UC;
[0009] S2 Data preprocessing: Segment the acquired fetal heart rate signal FHR and uterine contraction signal UC into multiple segmented time signals through a sliding window, and denoise each segmented signal.
[0010] S3 Time-frequency analysis: Perform wavelet transform on the fetal heart rate signal FHR and the uterine contraction signal UC respectively to generate time-frequency spectrograms;
[0011] S4 Feature extraction: Use a convolutional neural network to extract the time-frequency domain features of the fetal heart rate signal FHR and the uterine contraction signal UC respectively;
[0012] S5 Cross-attention feature fusion: Obtain the fusion features of the fetal heart rate signal FHR and the uterine contraction signal UC through a cross-attention module and a feature connection module;
[0013] S6 Time series analysis: Perform positional encoding on the fusion features and then input them into a Transformer encoder for temporal feature extraction;
[0014] S7 Recognition: Complete the classification prediction of the fetal state through a fully connected classification layer.
[0015] Furthermore, in the S2 data preprocessing step, cubic spline interpolation and adaptive filtering techniques are used to remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and the uterine contraction signal UC; or
[0016] The convolutional neural network in the step S4 includes a number of cascaded 3D convolutional layers and max pooling layers; a batch normalization + ReLU activation function layer is set after the 3D convolutional layer.
[0017] Furthermore, the cross-attention module in the step S5 is carried out according to the following steps:
[0018] S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix;
[0019] S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0020] S53 Perform weighted averaging on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature;
[0021] S54 Calculate the similarity between the uterine contraction signal UC feature as the query Q matrix and the fetal heart rate signal FHR feature as the key K matrix;
[0022] S55 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0023] S56 Perform weighted averaging on the attention weights and the fetal heart rate signal FHR feature as the value V matrix to obtain the uterine contraction signal UC - fetal heart rate signal FHR cross-attention feature;
[0024] S5 calculates the cross-attention fusion feature according to the following formula:
[0025]
[0026]
[0027] where d k is the dimension of the key matrix;
[0028] Attention FHR-UC is the cross-attention feature of the fetal heart rate signal FHR - uterine contraction signal UC;
[0029] Attention UC-FHR is the cross-attention feature of the uterine contraction signal UC - fetal heart rate signal FHR;
[0030] CrossAttention is the cross-attention fusion feature.
[0031] Furthermore, the feature connection module in step S5 splices the three feature tensors of the fetal heart rate signal FHR feature, uterine contraction signal UC feature, and cross-attention fusion feature according to the following steps:
[0032] First, three two-dimensional tensors are obtained after max pooling;
[0033] Then the three two-dimensional tensors are spliced;
[0034] Finally, a feature fusion tensor is obtained through a linear layer.
[0035] Furthermore, the position encoding in step S6 is performed according to the following formula:
[0036]
[0037] In the formula, PE represents the position encoding;
[0038] Generated by sine and cosine functions of different frequencies;
[0039] O is the position where it is located;
[0040] i is the corresponding dimension;
[0041] d m is the length of the feature vector after 3D convolution of each frame of the feature map.
[0042] Furthermore, the Transformer encoder includes a multi-head attention mechanism, and the specific steps of the multi-head attention mechanism are as follows:
[0043] The input feature vectors are first linearly transformed to obtain a query matrix Q, a key matrix K, and a value matrix V respectively. Let the comprehensive feature vector be where d model is the dimension of the feature vector. For each head i (i = 1, 2, …, h), its corresponding query matrix Q i , key matrix K i , and value matrix V i are:
[0044]
[0045] wherein, have dimensions of d model ×d q , d model ×d k , d model ×d v ;
[0046] S62 Calculate the attention score score ij :
[0047]
[0048] S63 Normalize the attention score to obtain the attention weight attn ij :
[0049]
[0050] S64 Calculate the output MHA of the multi - head attention layer according to the attention weight:
[0051] MHA = Concat(o1, o2, …, o h )W O ;
[0052] wherein, W O is a linear transformation matrix used to map the concatenated head outputs back to d model dimensions.
[0053] Furthermore, the Transformer encoder includes a feed - forward network layer. The feed - forward network includes two linear layers and a ReLU activation function. Let the input dimension of the first linear layer be d model , and the output dimension be d ff , then its calculation process is:
[0054] FFN1 = MHA×W1 + b1
[0055] where W1 is the weight matrix of the first linear layer, with dimensions of d model ×dff , b1 is a bias vector with a dimension of d ff ; then, a non-linear transformation is performed through the ReLU activation function:
[0056] FFN relu = ReLU(FFN1)
[0057] The second linear layer maps the dimension from d ff back to d model , and the calculation process is as follows:
[0058] FFN2 = FFN relu ×W2 + b2
[0059] where W2 is the weight matrix of the second linear layer with a dimension of d ff ×d model , and b2 is a bias vector with a dimension of d model .
[0060] The fetal state dynamic monitoring system based on multi-modal deep learning provided by the present invention includes:
[0061] A fetal heart rate signal acquisition module for acquiring the fetal heart rate signal FHR;
[0062] A uterine contraction signal acquisition module for acquiring the uterine contraction signal UC;
[0063] A data preprocessing module: used to perform segmented processing on the acquired fetal heart rate signal FHR or uterine contraction signal UC using a sliding window respectively, and remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and uterine contraction signal UC;
[0064] A time-frequency analysis module for performing wavelet transform on the fetal heart rate signal FHR or uterine contraction signal UC respectively to generate a time-frequency spectrogram; wavelet analysis is used for the fetal heart rate signal FHR or uterine contraction signal UC;
[0065] A feature extraction module that uses a 3D convolutional neural network to extract the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC respectively;
[0066] A cross-attention feature fusion module for inputting the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC into a cross-attention module and a feature connection module for processing to obtain the fusion feature signals of the fetal heart rate signal FHR and uterine contraction signal UC;
[0067] A time series analysis module for performing position encoding on the fusion features and then inputting them into a Transformer encoder to extract time series features;
[0068] An identification module for completing fetal state classification prediction through a fully connected classification layer.
[0069] Furthermore, the cross-attention feature fusion module is used to input two feature signals into the cross-attention module and the feature connection module for processing, so as to obtain the fusion features of the fetal heart rate signal FHR and the uterine contraction signal UC. The specific steps are as follows:
[0070] S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix;
[0071] S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0072] S53 Perform weighted averaging on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR-uterine contraction signal UC cross-attention feature;
[0073] S54 Calculate the similarity between the uterine contraction signal UC feature as the query Q matrix and the fetal heart rate signal FHR feature as the key K matrix;
[0074] S55 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0075] S56 Perform weighted averaging on the attention weights and the fetal heart rate signal FHR feature as the value V matrix to obtain the uterine contraction signal UC-fetal heart rate signal FHR cross-attention feature;
[0076] S57 Calculate the cross-attention fusion feature;
[0077] The calculation formula for cross-attention is:
[0078]
[0079]
[0080] where d k is the dimension of the key matrix;
[0081] Attention FHR-UC is the fetal heart rate signal FHR-uterine contraction signal UC cross-attention feature;
[0082] Attention UC-FHR is the uterine contraction signal UC-fetal heart rate signal FHR cross-attention feature;
[0083] CrossAttention is the cross-attention fusion feature.
[0084] The fetal state dynamic monitoring system based on multimodal deep learning provided by the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.
[0085] The beneficial effects of the present invention are as follows:
[0086] The fetal state dynamic monitoring method and system based on multimodal deep learning provided by the present invention uses a sliding window to segment the fetal heart rate signal and uterine contraction signal into multiple segmented time signals, and performs denoising preprocessing on each segmented signal; performs wavelet transform on the signal to obtain a time-frequency spectrogram, and inputs it into a 3D convolutional neural network to extract features; introduces a cross-attention module to fuse the features of the two modalities, inputs it into a transformer encoder to complete the extraction of temporal features, and finally realizes fetal state monitoring through a fully connected layer. This method utilizes multimodal fusion and deep learning technologies to accurately and dynamically monitor the fetal state. It can effectively reduce the manual monitoring error, improve the monitoring efficiency and accuracy, provide strong support for medical staff to understand the fetal situation in real time, help to detect fetal abnormal states in time, and ensure the safety of the fetus.
[0087] This method comprehensively utilizes the fetal heart rate signal and uterine contraction signal. Through a series of operations such as the acquisition, preprocessing, time-frequency spectrogram generation, feature extraction and fusion of the two signals, it fully integrates the key information in the multimodal data. Compared with the single-modal monitoring method, multimodal fusion can more comprehensively depict the state of the fetus in the uterus, avoid misjudgment caused by the limitations of a single signal, and significantly improve the accuracy and reliability of fetal state monitoring. The fetal heart rate signal (FHR signal) can reflect the activity of the fetal heart, while the uterine contraction signal (UC) is related to the contraction state of the uterus. The combination of the two can more accurately judge whether the fetus is in distress or other abnormal states, providing a more powerful basis for clinical decision-making.
[0088] In the process of feature extraction and time series analysis, this method adopts advanced deep learning algorithms such as 3D convolutional neural network and Transformer encoding layer. 3D CNN can automatically learn the deep features of the signal from the time-frequency domain, effectively avoiding the cumbersome process of manual feature extraction and information loss problems in traditional machine learning methods. The Transformer encoder can better capture the long-range dependence relationship in the sequence data, extract the key attention information, accurately model the temporal changes of the fetal state, and further improve the accuracy of classification prediction. This end-to-end deep learning architecture can automatically learn complex patterns and features from a large amount of data, greatly improving the intelligent level and efficiency of the monitoring system.
[0089] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0090] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention provides the following drawings for description.
[0091] Figure 1 is the overall flowchart of the present invention;
[0092] Figure 2 is a schematic diagram of the 3DCNN network structure;
[0093] Figure 3 is a schematic diagram of the cross-attention module;
[0094] Figure 4 is a schematic diagram of the multi-head attention structure. Detailed Embodiments
[0095] The following further describes the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0096] Embodiment 1
[0097] Please refer to Figure 1 , the fetal state dynamic monitoring method based on multi-modal fusion and deep learning provided by the embodiment of the present invention includes the following steps:
[0098] S1 Data Acquisition: Collect the fetal heart rate signal FHR and uterine contraction signal UC. The fetal heart rate signal is collected by a Doppler ultrasound probe placed on the pregnant woman's abdomen, and the uterine contraction pressure signal is measured by a pressure sensor placed on the pregnant woman's abdomen.
[0099] S2 Data Preprocessing: The collected fetal heart rate signal FHR and uterine contraction signal UC are respectively divided into 8 segments using a sliding window, with each segment being 10 minutes long. Cubic spline interpolation and adaptive filtering techniques are used to remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and uterine contraction signal UC.
[0100] S3 Time-frequency analysis: Perform wavelet transforms on the fetal heart rate signal FHR and the uterine contraction signal UC respectively to generate time-frequency spectrograms. For the fetal heart rate signal FHR, use the Daubechies wavelet with a scale parameter α = 2 and a translation step size that is 2 times the sampling period; for the uterine contraction signal UC, use the Morlet wavelet with a scale parameter α = 3 and a translation step size that is 2 times the sampling period.
[0101] S4 Feature extraction: Use two parallel 3D convolutional neural networks (3DCNNs) to extract the time-frequency domain features of the fetal heart rate signal FHR and the uterine contraction signal UC respectively; as Figure 2 shown, each 3DCNN includes several cascaded 3D convolutional layers and max pooling layers; a batch normalization + ReLU activation function layer is set after the 3D convolutional layer.
[0102] The 3D convolutional neural network in this embodiment is composed of 3 three-dimensional convolutional layers and 3 pooling layers cross-stacked with each other. The convolutional layer is provided with multiple convolutional kernels, the size of the convolutional kernel is 3, the stride is 1, and the padding method is "same". The pooling layer is a max pooling layer, the size of the pooling kernel is 2, and the stride is 2. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function layer.
[0103] S5 Cross-attention feature fusion: Input the two feature maps into the cross-attention module and the feature connection module to obtain the fusion features of the fetal heart rate signal FHR and the uterine contraction signal UC;
[0104] The cross-attention module, as Figure 3 shown, establishes the connection between the two modalities and fuses the features of different modalities. The specific steps are as follows:
[0105] S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix;
[0106] S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0107] S53 Perform weighted averaging on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature;
[0108] S54 Calculate the similarity between the uterine contraction signal UC feature as the query Q matrix and the fetal heart rate signal FHR feature as the key K matrix;
[0109] S55 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0110] S56 weights the attention weights and the fetal heart rate signal FHR features as the value V matrix for weighted averaging to obtain the uterine contraction signal UC - fetal heart rate signal FHR cross-attention features;
[0111] S57 calculates the cross-attention fusion features;
[0112] The calculation formula for cross-attention is:
[0113]
[0114] where d k is the dimension of the key matrix;
[0115] Attention FHR-UC is the fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature;
[0116] Attention UC-FHR is the uterine contraction signal UC - fetal heart rate signal FHR cross-attention feature;
[0117] CrossAttention is the cross-attention fusion feature;
[0118] Figure 3 mainly includes a fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature calculation module and a uterine contraction signal UC - fetal heart rate signal FHR cross-attention feature calculation module, where
[0119] Q FHR 、K FHR 、W FHR respectively represent the query matrix, key matrix, and value matrix of the fetal heart rate signal FHR; W q1 、W k1 、W v1 respectively represent the parameter matrices for generating Q FHR 、K FHR and V FHR and are adjusted as model parameters during training;
[0120] Q UC 、K UC 、V UC respectively represent the query matrix, key matrix, and value matrix of the fetal heart rate signal UC; W q2 、W k2 、W v2 respectively represent the parameter matrices for generating Q UC 、K UC and V UC and are adjusted as model parameters during training;
[0121] AFHR―uc and A UC―FHR respectively represent the cross - attention weights of the fetal heart rate signal FHR - uterine contraction signal UC and the cross - attention weights of the uterine contraction signal UC - fetal heart rate signal FHR;
[0122] Softmax is a normalization function;
[0123] The feature connection module in this embodiment splices and fuses the three feature tensors of the fetal heart rate signal FHR feature, uterine contraction signal UC feature and cross - attention fusion feature, and the specific steps are as follows:
[0124] First, three two - dimensional tensors are obtained after max - pooling;
[0125] The three two - dimensional tensors are spliced;
[0126] A tensor of the fusion feature is obtained through a linear layer.
[0127] S6 Time - series analysis: In this embodiment, the fusion feature is position - encoded and then input into a Transformer encoder for time - series feature extraction. The specific steps are as follows:
[0128]
[0129] In the formula, PE represents position encoding, which is generated by sine and cosine functions of different frequencies. O is the position, i is the corresponding dimension, and dm is the length of the feature vector after 3D convolution of each frame of the feature map;
[0130] The encoded feature vector is input into the multi - head attention mechanism to focus on the data from multiple perspectives, capture complex dependencies, enhance the model's expression ability and adaptability, so as to process the fetal state monitoring data more comprehensively and accurately.
[0131] As Figure 4 shown, the specific steps of the multi - head attention mechanism are as follows:
[0132] S61 The input feature vector first undergoes a linear transformation to obtain a query matrix Q, a key matrix K, and a value matrix V respectively; Let the comprehensive feature vector be where d model is the dimension of the feature vector. For each head i (i = 1, 2, …, h), its corresponding query matrix Q i , key matrix K i and value matrix V i :
[0133]
[0134] where have dimensions of d model ×dq , d model ×d k , d model ×d v ;
[0135] S62 calculates the attention score score ij :
[0136]
[0137] S63 normalizes the attention score to obtain the attention weight attn ij :
[0138]
[0139] S64 calculates the output MHA of the multi-head attention layer according to the attention weight:
[0140] MHA = Concat(o1, o2, …, o h )W O ;
[0141] Among them, W O is a linear transformation matrix used to map the concatenated head outputs back to d model dimensions.
[0142] S65 feeds the output MHA of the multi-head attention layer into the feed-forward network layer. The feed-forward network consists of two linear layers and a ReLU activation function. Let the input dimension of the first linear layer be d model , and the output dimension be d ff , then its calculation process is:
[0143] FFN1 = MHA × W1 + b1;
[0144] Among them, W1 is the weight matrix of the first linear layer, with dimensions d model ×d ff , and b1 is the bias vector, with dimensions d ff ; then it undergoes a non-linear transformation through the ReLU activation function:
[0145] FFN relu = ReLU(FFN1);
[0146] The second linear layer maps the dimension from d ff back to d model , and the calculation process is:
[0147] FFN2 = FFN relu × W2 + b2;
[0148] Among them, W2 is the weight matrix of the second linear layer, with a dimension of d ff ×d model , b2 is the bias vector, with a dimension of d model , and finally, the feature vector after Transformer encoding is obtained.
[0149] S7 recognition: Through the fully connected classification layer, the abstract features obtained after a series of previous processes are integrated and mapped, and they are converted into a probability distribution form corresponding to the fetal state classification categories (normal, suspicious, and abnormal). By setting a threshold, the classification sensitivity and specificity are adjusted to balance the misjudgment rate, thereby realizing the final classification prediction of the fetal state and providing a basis for clinical decision-making.
[0150] The number of neurons in the fully connected classification layer is 3. The fetal state is divided into three categories: normal, suspicious, and abnormal. The softmax activation function is used to convert the output into a probability distribution form. At the same time, a threshold is set to adjust the classification sensitivity and specificity to balance the misjudgment rate.
[0151] The fetal state dynamic monitoring system based on multi-modal deep learning provided in this embodiment includes
[0152] A fetal heart rate signal acquisition module for acquiring the fetal heart rate signal FHR;
[0153] A uterine contraction signal acquisition module for acquiring the uterine contraction signal UC;
[0154] A data preprocessing module: used to respectively perform sliding window segmentation processing on the acquired fetal heart rate signal FHR or uterine contraction signal UC, and remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and uterine contraction signal UC;
[0155] A time-frequency analysis module for respectively performing wavelet transform on the fetal heart rate signal FHR or uterine contraction signal UC to generate a time-frequency spectrogram; wavelet analysis processing is adopted for the fetal heart rate signal FHR signal or uterine contraction signal UC;
[0156] A feature extraction module that uses a 3D convolutional neural network (3DCNN) to respectively extract the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC;
[0157] A cross-attention feature fusion module for inputting the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC into the cross-attention module and the feature connection module for processing to obtain the fusion feature signals of the fetal heart rate signal FHR and uterine contraction signal UC;
[0158] A time series analysis module for performing positional encoding on the fusion features, and then inputting them into the Transformer encoder to extract temporal features;
[0159] An identification module, configured to complete the classification prediction of the fetal state through a fully connected classification layer;
[0160] In this embodiment, the fully connected classification layer integrates and maps the abstract features obtained after a series of previous processes, converts them into a probability distribution form corresponding to the fetal state classification categories (normal, suspicious, and abnormal), and adjusts the classification sensitivity and specificity by setting a threshold to balance the misjudgment rate, thereby realizing the final classification prediction of the fetal state and providing a basis for clinical decision-making.
[0161] The number of neurons in the fully connected classification layer of this embodiment is 3. The fetal state is divided into three categories: normal, suspicious, and abnormal. The softmax activation function is used to convert the output into a probability distribution form, and at the same time, a threshold is set to adjust the classification sensitivity and specificity to balance the misjudgment rate.
[0162] The cross-attention feature fusion in this embodiment: Input two feature maps into the cross-attention module and the feature connection module to obtain the fusion feature of the fetal heart rate signal FHR and the uterine contraction signal UC;
[0163] The cross-attention module in this embodiment, as Figure 3 shown, establishes the connection between two modalities and fuses the features of different modalities. The specific steps are as follows:
[0164] S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix;
[0165] S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0166] S53 Perform weighted averaging on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR-uterine contraction signal UC cross-attention feature;
[0167] S54 Calculate the similarity between the uterine contraction signal UC feature as the query Q matrix and the fetal heart rate signal FHR feature as the key K matrix;
[0168] S55 Normalize the similarity matrix using the Softmax function to obtain the attention weights;
[0169] S56 Perform weighted averaging on the attention weights and the fetal heart rate signal FHR feature as the value V matrix to obtain the uterine contraction signal UC-fetal heart rate signal FHR cross-attention feature;
[0170] S57 Calculate the cross-attention fusion feature;
[0171] The calculation formula of cross-attention is as follows:
[0172]
[0173] where d k is the dimension of the key matrix;
[0174] Attention FHR-UC is the cross-attention feature of fetal heart rate signal FHR - uterine contraction signal UC;
[0175] Attention UC-FHR is the cross-attention feature of uterine contraction signal UC - fetal heart rate signal FHR;
[0176] CrossAttention is the cross-attention fusion feature.
[0177] The above-described embodiments are merely preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.
Claims
1. A method for dynamically monitoring fetal status based on multimodal deep learning, characterized in that: It includes the following steps: S1 Data acquisition: Acquire the fetal heart rate signal FHR and the uterine contraction signal UC; S2 Data preprocessing: Segment the acquired fetal heart rate signal FHR and uterine contraction signal UC into multiple segmented time signals through a sliding window, and denoise each segmented signal; S3 Time-frequency analysis: Perform wavelet transform on the fetal heart rate signal FHR and the uterine contraction signal UC respectively to generate time-frequency spectrograms; S4 Feature extraction: Use a convolutional neural network to extract the time-frequency domain features of the fetal heart rate signal FHR signal and the uterine contraction signal UC respectively; S5 Cross-attention feature fusion: Obtain the fusion features of the fetal heart rate signal FHR and the uterine contraction signal UC through a cross-attention module and a feature connection module; S6 Time series analysis: Perform positional encoding on the fusion features, and then input them into a Transformer encoder for temporal feature extraction; S7 Recognition: Complete the classification prediction of the fetal state through a fully connected classification layer.
2. The method for dynamically monitoring fetal state based on multi-modal deep learning according to claim 1, characterized in that, In the step S2 of data preprocessing, cubic spline interpolation and adaptive filtering technology are used to remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and the uterine contraction signal UC; or The convolutional neural network in the step S4 includes a number of 3D convolutional layers and max pooling layers connected in series; a batch normalization + ReLU activation function layer is arranged after the 3D convolutional layer.
3. The method for dynamically monitoring fetal state based on multi-modal deep learning according to claim 1, characterized in that, The cross-attention module in the step S5 is carried out according to the following steps: S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix; S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights; S53 Perform weighted average on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature; S54 Calculate the similarity between the uterine contraction signal UC feature as the query Q matrix and the fetal heart rate signal FHR feature as the key K matrix; S55 Normalize the similarity matrix using the Softmax function to obtain the attention weights; S56 Perform weighted average on the attention weights and the fetal heart rate signal FHR feature as the value V matrix to obtain the uterine contraction signal UC - fetal heart rate signal FHR cross-attention feature; S57 Calculate the cross-attention fusion feature according to the following formula: where d k is the dimension of the key matrix; Attention FHR-UC It is the cross-attention feature of the fetal heart rate signal FHR - uterine contraction signal UC; Attention UC-FHR is the cross-attention feature of uterine contraction signal UC - fetal heart rate signal FHR; CrossAttention is the cross-attention fusion feature.
4. The method for dynamically monitoring fetal state based on multi-modal deep learning according to claim 1, characterized in that, The feature connection module in the step S5 splices the three feature tensors of the fetal heart rate signal FHR feature, the uterine contraction signal UC feature, and the cross-attention fusion feature according to the following steps: First, after max pooling, three two-dimensional tensors are obtained; Then, the three two-dimensional tensors are concatenated; Finally, a feature fusion tensor is obtained through a linear layer.
5. The dynamic fetal state monitoring method based on multi-modal deep learning according to claim 1, wherein the position encoding in step S6 is performed according to the following formula: In the formula, PE represents the position encoding; It is generated by sine and cosine functions of different frequencies; O is the position where it is located; i is the corresponding dimension; d m It is the length of the feature vector after 3D convolution for each frame of feature map.
6. The dynamic fetal state monitoring method based on multi-modal deep learning according to claim 1, wherein the Transformer encoder includes a multi-head attention mechanism, and the specific steps of the multi-head attention mechanism are as follows: The feature vectors input at S61 are first linearly transformed to obtain a query matrix Q, a key matrix K, and a value matrix V respectively; let the comprehensive feature vector be F = [f1, f2, …, f dmodel , where d model is the dimension of the feature vector. For each head i (i = 1, 2, …, h), its corresponding query matrix Q i , key matrix K i , and value matrix V i are as follows: Q i = F × W Qi ; Among them, The dimensions of model are d q × d model × d k × d model × d v ; S62 Calculate the attention score score ij : Normalize the attention score to obtain the attention weight attn ij : S64 Calculate the output MHA of the multi-head attention layer according to the attention weights: MHA = Concat(o1, o2, …, o h )W O ; Among them, W O is a linear transformation matrix for mapping the concatenated head output back to d model dimensions.
7. The dynamic fetal state monitoring method based on multi-modal deep learning according to claim 1, wherein The Transformer encoder includes a feed-forward network layer, and the feed-forward network includes two linear layers and a ReLU activation function. Suppose the input dimension of the first linear layer is d model , and the output dimension is d ff , then its calculation process is as follows: FFN1 = MHA × W1 + b1 Among them, W1 is the weight matrix of the first linear layer, with dimension d model ×d ff , and b1 is the bias vector, whose dimension is d ff ; then, a non-linear transformation is performed through the ReLU activation function: FFN relu = ReLU(FFN1) The second linear layer maps the dimension from d ff back to d model , and the calculation process is as follows: FFN2 = FFN relu × W2 + b2 Among them, W2 is the weight matrix of the second linear layer, with dimension d ff ×d model , b2 is the bias vector, with dimension d model .
8. A fetal state dynamic monitoring system based on multimodal deep learning, characterized in that It includes: A fetal heart rate signal acquisition module for acquiring the fetal heart rate signal FHR; A uterine contraction signal acquisition module for acquiring the uterine contraction signal UC; A data preprocessing module: used to respectively perform sliding window segmentation processing on the acquired fetal heart rate signal FHR or uterine contraction signal UC, and remove the high-frequency noise and baseline drift phenomena of the fetal heart rate signal FHR and uterine contraction signal UC; A time-frequency analysis module for respectively performing wavelet transform on the fetal heart rate signal FHR or uterine contraction signal UC to generate a time-frequency spectrogram; wavelet analysis is used for the fetal heart rate signal FHR signal or uterine contraction signal UC; A feature extraction module that uses a 3D convolutional neural network to respectively extract the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC; A cross-attention feature fusion module for inputting the time-frequency domain feature signals of the fetal heart rate signal FHR or uterine contraction signal UC into a cross-attention module and a feature connection module for processing to obtain the fusion feature signals of the fetal heart rate signal FHR and uterine contraction signal UC; A time series analysis module for performing position encoding on the fusion features, and then inputting them into a Transformer encoder to extract time series features; An identification module for completing fetal state classification prediction through a fully connected classification layer.
9. The dynamic fetal state monitoring system based on multi-modal deep learning according to claim 8, wherein the cross-attention feature fusion module is used to input two feature signals into a cross-attention module and a feature connection module for processing, so as to obtain the fusion features of the fetal heart rate signal FHR and uterine contraction signal UC. The specific steps are as follows: S51 Calculate the similarity between the fetal heart rate signal FHR feature as the query Q matrix and the uterine contraction signal UC feature as the key K matrix; S52 Normalize the similarity matrix using the Softmax function to obtain the attention weights; S53 Perform weighted averaging on the attention weights and the uterine contraction signal UC feature as the value V matrix to obtain the fetal heart rate signal FHR - uterine contraction signal UC cross-attention feature; S54 calculates the similarity between the uterine contraction signal UC features as the query Q matrix and the fetal heart rate signal FHR features as the key K matrix; S55 normalizes the similarity matrix using the Softmax function to obtain the attention weights; S56 performs a weighted average of the attention weights and the fetal heart rate signal FHR features as the value V matrix to obtain the uterine contraction signal UC - fetal heart rate signal FHR cross-attention features; S57 calculates the cross-attention fusion features; The calculation formula for cross-attention is: where d k is the dimension of the key matrix; Attention FHR-UC It is the cross-attention feature of the fetal heart rate signal FHR - uterine contraction signal UC; Attention UC-FHR is the cross-attention feature of uterine contraction signal UC - fetal heart rate signal FHR; CrossAttention is the cross-attention fusion feature.
10. A fetal state dynamic monitoring system based on multimodal deep learning, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7 above.
Citation Information
Cited By
Method and system for constructing fetal heart monitoring auxiliary interpretation model based on similarity learning
CN120727295A
Method and system for constructing fetal heart monitoring auxiliary interpretation model based on one-dimensional time sequence signal and double neural network architecture
CN120748765A
Blood sample detection method and system based on multi-modal biological sample
CN122361781A