Method and system for labor pain assessment based on adaptive fusion of multi-modal signals and storage medium
Patent Information
- Application Number
- CN202611242837.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的在于提供一种基于多模态信号自适应融合的分娩疼痛评估方法、系统及存储介质,以解决现有疼痛评估技术中存在的以下技术问题:现有基于单一生理信号的疼痛评估方法易受个体差异与环境噪声干扰,评估准确性不足;现有多模态疼痛评估方案未针对分娩场景的周期性宫缩特征进行时间锚定,缺乏以宫缩峰值时刻为基准提取目标时间段内生物行为信号并建立刺激—反应同步关联的技术手段;现有方案缺少根据面部表情、呼吸和声音等模态在不同宫缩阶段的可靠性自适应生成模态权重的机制,也缺少从融合时序特征中提取宫缩周期级疼痛峰值表征的查询聚合机制,导致疼痛评估结果与产程进展缺乏精准关联,难以区分宫缩相关疼痛、异常疼痛以及第二产程用力行为引起的非疼痛干扰
[0036]1、通过将宫缩压力信号作为周期性刺激信号,并以宫缩峰值时刻作为时间锚点提取目标时间段内的生物行为信号,本发明能够将疼痛反应与宫缩周期建立明确对应关系,解决现有疼痛评估结果与产程进展不同步的问题,使输出结果具有更清晰的临床时序背景和可解释性。
Smart Images

Figure CN122805213A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence in the field of medical and health technology, specifically to the application of deep learning in pain assessment, and in particular to a method, system and storage medium for assessing labor pain based on multimodal signal adaptive fusion. Background Technology
[0002] Labor pain assessment is a core component of obstetric analgesia management. Traditional assessment methods primarily rely on subjective reports from the mother, such as the Visual Analogue Scale (VAS). However, mothers often struggle to accurately express their pain during contractions, and assessments are conducted at long intervals, hindering continuous monitoring. While existing objective assessment methods based on single physiological signals (such as heart rate or facial expression) can reduce subjectivity, these single-modal signals are susceptible to individual differences and environmental noise, resulting in insufficient accuracy.
[0003] A Chinese patent with publication number CN120899167A discloses a pain assessment system based on multimodal physiological signals, which assesses pain by collecting data such as facial expressions, vocal features, and physiological signs and combining them with deep learning. However, this solution does not time-anchor the periodic stimulation characteristics of uterine contraction pressure signals in the labor process, and lacks the technical means to extract the pain response time window based on the peak time of contractions and establish a synchronous stimulus-response correlation. At the same time, the solution does not adaptively allocate weights based on the temporal changes of biological behavioral signals such as facial expressions, breathing, and voice within the contraction cycle, nor does it set up a query aggregation mechanism for extracting pain peak representations. Therefore, it is difficult to achieve labor pain assessment that is synchronous with the progress of labor and has clinical interpretability.
[0004] Therefore, while existing multimodal pain assessment schemes can reduce the uncertainty of a single modality to some extent by utilizing multi-source signals, they still have the following shortcomings: First, they do not establish time anchors by using uterine contraction pressure signals as periodic stimulus signals, and there is a lack of precise correspondence between pain response segments and uterine contraction cycles; Second, they do not adaptively generate modal weights based on the reliability of different modalities at different stages of uterine contractions, and are easily affected by individual differences in parturients, environmental noise, and pushing behavior during the second stage of labor; Third, they lack a query aggregation mechanism for peak uterine contraction pain status, making it difficult to extract the most relevant representations to peak pain from continuous temporal features. Summary of the Invention
[0005] The purpose of this invention is to provide a method, system, and storage medium for assessing labor pain based on adaptive fusion of multimodal signals, in order to solve the following technical problems existing in current pain assessment technologies: existing pain assessment methods based on single physiological signals are easily affected by individual differences and environmental noise, resulting in insufficient assessment accuracy; existing multimodal pain assessment schemes do not time-anchor the periodic uterine contraction characteristics of the labor scenario, and lack technical means to extract biological behavioral signals within the target time period based on the peak time of uterine contractions and establish a synchronous stimulus-response relationship; existing schemes lack a mechanism for adaptively generating modal weights based on the reliability of modalities such as facial expressions, breathing, and voice at different stages of uterine contractions, and also lack a query aggregation mechanism for extracting the cyclical pain peak representation of uterine contractions from the fused temporal features, resulting in a lack of accurate correlation between pain assessment results and labor progress, and difficulty in distinguishing between contraction-related pain, abnormal pain, and non-pain interference caused by pushing behavior in the second stage of labor.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] This invention provides a method for assessing labor pain based on adaptive fusion of multimodal signals, comprising the following steps: collecting at least two biological behavioral signals and one periodic stimulation signal, wherein the biological behavioral signals include at least two of facial expression signals, respiratory signals, and sound signals, and the periodic stimulation signal includes uterine contraction pressure signals; determining the peak time of uterine contractions based on the periodic stimulation signal, using it as a time anchor point, and extracting the maternal biological behavioral signals within a preset time window before and after the peak time as the biological behavioral signals for a target time period; extracting temporal features from the at least two biological behavioral signals within the target time period to obtain corresponding modal temporal features; inputting the modal temporal features into an adaptive modal gating module to generate modal weights corresponding to each modality, and performing cross-modal fusion of the modal temporal features based on the modal weights to obtain fused temporal features; extracting pain peak representations from the fused temporal features through a pain peak query aggregation module; and outputting pain assessment results based on the pain peak representations. This method, through uterine contraction peak time anchoring, adaptive modal gating, and pain peak query aggregation, enables the pain assessment results to remain synchronized with the uterine contraction cycle and labor progress.
[0008] Furthermore, the periodic stimulation signal is a uterine contraction pressure signal, and the method is used for labor pain assessment; the target time period extraction includes extracting at least two of the facial expression signal, breathing signal, and sound signal within a preset time before and after the peak of the uterine contraction pressure signal, based on the peak time of the uterine contraction pressure signal, in order to establish a synchronous correlation between the intensity of uterine contraction stimulation and pain behavior response.
[0009] Furthermore, the biological behavioral signals include facial expression signals. Feature extraction of the facial expression signals includes: face detection, alignment, and cropping of facial videos within a target time period, and extraction of temporal features of facial expressions using a pre-trained facial video encoder. These temporal features of facial expressions are used to characterize pain-related dynamic facial expression changes such as frowning, eye closure, deepening of nasolabial folds, and facial muscle tension.
[0010] Preferably, the pre-trained face video encoder employs a face video Transformer encoder pre-trained based on a mask reconstruction task, such as the MARLIN architecture. This pre-trained encoder can provide stable facial dynamic representations in small-sample medical scenarios and can choose to freeze all parameters, partially unfreeze high-level parameters, or use efficient parameter fine-tuning methods for transfer learning, depending on the scale of the training data.
[0011] Furthermore, the biological behavioral signal includes a respiratory signal, and feature extraction of the respiratory signal includes: filtering and normalizing the respiratory waveform within the target time period, and extracting respiratory temporal features through at least one of a one-dimensional convolutional network, a recurrent neural network, a temporal convolutional network, or a Transformer encoder. The respiratory temporal features are used to characterize dynamic changes related to pain response, such as respiratory rate, amplitude, rhythm stability, and breath-holding, shallow rapid breathing, or hyperventilation.
[0012] Furthermore, the biological behavioral signal includes a sound signal, and feature extraction of the sound signal includes: denoising, framing, and windowing the sound signal within the target time period; extracting at least one of Mel-spectral features, Mel-frequency cepstral coefficients, energy features, or fundamental frequency features; and extracting temporal features of the sound through a sound coding network. These temporal features are used to characterize pain-related acoustic changes such as groans, shouts, changes in sound intensity, changes in fundamental frequency, and spectral distribution.
[0013] Furthermore, the adaptive modal gating module receives at least two modal temporal features and outputs corresponding modal weights. These modal weights are used to adjust the contribution of each modal temporal feature in the cross-modal fusion process. The adaptive modal gating module includes a multilayer perceptron and a normalization function. The multilayer perceptron generates weight scores based on each modal temporal feature, and the normalization function converts these weight scores into modal weights that satisfy normalization constraints.
[0014] Preferably, the cross-modal fusion step includes:
[0015] S1: Perform time alignment and unified dimensional projection on the temporal features of each modality to obtain the temporal features of the modality in the same feature space;
[0016] S2: Input the modal temporal features into the adaptive modal gating module to generate at least two of the following: facial modal weights, breathing modal weights, and voice modal weights;
[0017] S3: Based on the modal weights, the corresponding modal temporal features are weighted, and the weighted modal temporal features are input into the cross-modal fusion module for temporal interaction modeling to obtain fused temporal features.
[0018] This fusion process aligns the scale and distribution of features across different modalities through unified dimensional projection, suppresses interference from low-reliability modalities through adaptive modal gating, and preserves complementary information between different modalities through cross-modal temporal interaction, thereby improving the robustness of labor pain assessment under individual differences, delivery room noise, and the interference of pushing during the second stage of labor.
[0019] Furthermore, the method also includes a pain peak query aggregation step: setting a learnable pain peak query vector, generating time attention weights based on the correlation between the pain peak query vector and each time position feature in the fused temporal features, and weighting and aggregating the fused temporal features based on the time attention weights to obtain a pain peak representation; the output pain assessment results include outputting pain level classification results through a classification prediction branch, and outputting a continuous pain intensity score through a regression prediction branch.
[0020] Preferably, the pain level classification results include three levels: mild, moderate, and severe; the continuous pain intensity score is a value from 0 to 10. This classification and scoring standard corresponds to the commonly used VAS visual analog scale in clinical practice, making it easy for clinical medical staff to understand and use.
[0021] Furthermore, the method also includes a training step: based on a dataset containing facial expression signals, breathing signals, sound signals, periodic stimulus signals, and pain annotation results, the feature extraction module, adaptive modal gating module, cross-modal fusion module, pain peak query aggregation module, and output module are jointly trained; the pain annotation results include at least one of pain level annotation and continuous pain intensity score.
[0022] This invention also provides a labor pain assessment system based on adaptive fusion of multimodal signals, the structure of which can be found in [reference needed]. Figure 1The system includes a multimodal signal acquisition module, a target time period extraction module, a feature extraction module, an adaptive gating module, a feature fusion module, a pain peak query and aggregation module, and an output module. The multimodal signal acquisition module acquires at least two of the following: facial expression signals, respiratory signals, and sound signals, as well as uterine contraction pressure signals. The target time period extraction module determines the peak time of uterine contractions based on the uterine contraction pressure signals and extracts biological behavioral signals within a preset time window before and after the peak. The feature extraction module outputs the temporal features of each modality. The adaptive gating module generates modal weights. The feature fusion module obtains the fused temporal features. The pain peak query and aggregation module extracts the pain peak representation. The output module outputs the pain level classification results and / or a continuous pain intensity score.
[0023] A multimodal signal acquisition module is used to acquire at least two biological behavioral signals and periodic stimulation signals. The biological behavioral signals include at least two of facial expression signals, respiratory signals, and sound signals. The periodic stimulation signals include uterine contraction pressure signals.
[0024] The target time period extraction module is used to determine the uterine contraction cycle or the peak time of uterine contraction based on the periodic stimulation signal, and to extract the biological behavior signal within the corresponding target time period from the biological behavior signal based on the uterine contraction cycle or the peak time of uterine contraction.
[0025] The feature extraction module is used to extract temporal features from at least two biological behavioral signals within the target time period to obtain the corresponding modal temporal features.
[0026] An adaptive modal gating module is used to generate modal weights corresponding to each modality based on the modal temporal features.
[0027] A cross-modal fusion module is used to perform cross-modal fusion of the modal temporal features based on the modal weights to obtain fused temporal features;
[0028] The pain peak query and aggregation module is used to extract pain peak representations from the fused temporal features;
[0029] The output module is used to output pain assessment results based on the pain peak characterization.
[0030] Furthermore, the feature extraction module includes at least two of the following: a facial feature extraction unit, a respiratory feature extraction unit, and a voice feature extraction unit. The facial feature extraction unit employs a pre-trained face video encoder, the respiratory feature extraction unit extracts respiratory rhythm temporal features, and the voice feature extraction unit extracts voice temporal features. The adaptive modal gating module includes a multilayer perceptron and a normalization function, used to output at least two of the following: facial modal weights, respiratory modal weights, and voice modal weights. The pain peak query aggregation module includes a learnable pain peak query vector, used to perform attention aggregation on the fused temporal features to obtain a representation of uterine contraction cycle-level pain peaks.
[0031] Furthermore, the cross-modal fusion module includes a unified dimension projection unit, a weighted fusion unit, and a cross-modal temporal interaction unit; the unified dimension projection unit is used to map temporal features of different modalities to a unified feature space, the weighted fusion unit is used to weight each temporal feature based on the modal weights generated by the adaptive modal gating module, and the cross-modal temporal interaction unit is used to fuse the weighted temporal features through a Transformer encoder, a cross-attention network, or a temporal convolutional network to obtain fused temporal features.
[0032] Furthermore, the output module includes a classification prediction branch and a regression prediction branch. The classification prediction branch includes a fully connected layer and a Softmax activation function to output the pain level classification result; the regression prediction branch includes a fully connected layer and a sigmoid activation function to output a continuous pain intensity score. This dual-branch output design enables the system to simultaneously provide pain level classification and continuous scoring, meeting the needs of different clinical application scenarios.
[0033] Furthermore, the periodic stimulation signal is a uterine contraction pressure signal; the system is used for real-time assessment and monitoring of labor pain. This system is specifically designed for the labor process, enabling continuous monitoring throughout labor and providing real-time, objective pain assessment data for labor pain management.
[0034] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for assessing labor pain based on adaptive fusion of multimodal signals, including steps such as data acquisition, target time period extraction, preprocessing, modal temporal feature extraction, adaptive gating, cross-modal fusion, pain peak query aggregation, and pain assessment output.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. By using uterine contraction pressure signals as periodic stimulation signals and extracting biological behavioral signals within the target time period using the peak time of uterine contractions as the time anchor point, this invention can establish a clear correspondence between pain response and uterine contraction cycle, solving the problem of existing pain assessment results not being synchronized with the progress of labor, and making the output results have a clearer clinical temporal background and interpretability.
[0037] Furthermore, the time anchor point established based on the peak contraction time can also be used to output auxiliary prompts. When there is a significant and persistent time deviation between the peak pain level and the peak contraction pressure, or when a high pain assessment result still occurs during the contraction interval, the system can output a "Pain-Contraction Correlation Abnormality" prompt to the medical staff terminal and display the corresponding pain assessment curve and contraction pressure curve. This prompt is only used to assist medical staff in making further judgments and does not directly replace clinical diagnosis or automated medication decisions.
[0038] 2. By working collaboratively with the adaptive modal gating module and the cross-modal fusion module, the system can dynamically adjust modal weights based on the reliability of facial expression signals, respiratory signals, and sound signals within the target uterine contraction time window, and perform cross-modal temporal interactions within a unified feature space. Compared to simple feature splicing or fixed-weight fusion, this mechanism can reduce the adverse effects of a modality when it is occluded, subjected to environmental noise, or interfered with by pushing behavior during the second stage of labor, thereby improving the robustness of labor pain assessment.
[0039] 3. By employing a pre-trained face video encoder for temporal feature extraction of facial expressions, the system can obtain relatively stable facial dynamic representations in medical scenarios with limited labeled samples. The pre-trained encoder can be MARLIN or other face video Transformer encoders, and its output serves as temporal features of facial modalities in subsequent adaptive modal gating and cross-modal fusion, rather than being used alone as the final basis for pain judgment.
[0040] 4. By extracting temporal features from respiratory and auditory signals respectively, the system can capture changes in respiratory rhythm and evolution of vocal characteristics caused by pain. Respiratory temporal features may include information such as respiratory frequency, amplitude changes, and rhythm stability; auditory temporal features may include information such as energy, fundamental frequency, spectral distribution, and vocal rhythm, thus providing multi-source biological behavior evidence for adaptive modal gating and cross-modal fusion.
[0041] 5. By employing a multi-task loss function to jointly optimize pain level classification and pain intensity regression, and combining class imbalance handling and robust regression loss, the system can simultaneously meet the needs of clinical hierarchical management and continuous pain intensity quantification. During training, the feature extraction module, adaptive modal gating module, cross-modal fusion module, pain peak query aggregation module, and output module can be jointly trained, or trained under small sample conditions using a pre-trained encoder frozen or partially fine-tuned method.
[0042] 6. By continuously and synchronously collecting at least two biological behavioral signals and uterine contraction pressure signals, the system can output pain assessment results corresponding to the uterine contraction cycle, providing medical staff with objective and continuous auxiliary decision-making information, and can be used as one of the inputs of a closed-loop intelligent pain system or clinical decision support system. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the architecture of a labor pain assessment system based on adaptive fusion of multimodal signals, showing the data flow between the multimodal signal acquisition module, target time period extraction module, feature extraction module, adaptive gating module, feature fusion module, pain peak query aggregation module, and output module.
[0044] Figure 2 The diagram illustrates the process of extracting facial image sequences, respiratory waveform segments, and sound segments within a preset time window before and after the uterine contraction peak time determined by the uterine contraction pressure signal as the time anchor point. Detailed Implementation
[0045] This invention provides a method, system, and computer-readable storage medium for assessing labor pain based on adaptive fusion of multimodal signals. For example... Figure 1 As shown, the overall logic of this method is as follows: First, at least two of the following signals are collected: facial expression signal, respiratory signal, and sound signal, as well as uterine contraction pressure signal; then, the peak time of uterine contraction determined by the uterine contraction pressure signal is used as the time anchor point to extract biological behavior signals within a preset time window before and after the peak; then, the temporal features of each modality are extracted separately, and modal weights are generated through an adaptive gating module; subsequently, cross-modal temporal fusion is performed based on the modal weights to obtain fused temporal features; finally, the pain peak characterization is extracted through the pain peak query aggregation module, and the pain level classification results and / or continuous pain intensity scores are output through the output module.
[0046] In the above process, the uterine contraction pressure signal is used to determine the target time period and is not simply spliced as ordinary features; facial expression, breathing and sound signals are used to characterize pain behavior responses; the adaptive modal gating module is used to solve the problem that the reliability of different modalities varies with the individual mother, the stage of labor and the collection environment; the pain peak query aggregation module is used to extract the most relevant features to the uterine contraction peak pain from the continuously fused temporal features.
[0047] The following is combined Figure 1 , Figure 2 The following embodiments further illustrate the present invention. The embodiments are merely illustrative of optional implementations of the present invention, wherein the network structure, feature dimensions, sampling rate, and time window length can be adjusted according to actual device conditions and clinical data distribution, and should not be construed as limiting the scope of the claims.
[0048] Example 1
[0049] This embodiment provides a method for assessing labor pain based on adaptive fusion of multimodal signals. The overall module relationships can be found in [reference needed]. Figure 1 This method uses uterine contraction pressure signals as periodic stimulation signals and the peak time of uterine contractions as time anchors. It extracts biological behavioral signals within the target time period from facial expression, respiratory, and auditory signals, and performs temporal feature extraction on each signal to obtain corresponding modal temporal features. Subsequently, an adaptive gating module generates modal weights, performs cross-modal fusion based on these weights, and obtains pain peak representations through a pain peak query aggregation module. Finally, it outputs pain level classification results and / or continuous pain intensity scores.
[0050] The processing flow of this embodiment includes, in sequence: data acquisition, optional individualized baseline calibration, extraction of uterine contraction peak time window, preprocessing, modal temporal feature extraction, adaptive modal gating, cross-modal fusion, pain peak query aggregation, and pain assessment output.
[0051] Specifically, the pain assessment method in this embodiment includes the following steps.
[0052] Step S1 involves data acquisition. An RGB camera, a non-contact respiratory monitoring device, and a microphone array are deployed next to the mother's bedside to simultaneously acquire facial expression videos, respiratory waveforms, and sound signals. The facial expression video is acquired at a frame rate of 30fps with a resolution of 1920×1080.
[0053] Step S1.1 is an optional individualized baseline calibration. Without increasing the burden on the mother and in accordance with clinical procedures, facial expressions, respiratory rhythms, and voice baselines of the mother at rest can be collected during the intercontraction period in the early stages of labor for subsequent differential quantification of biological behavioral signal characteristics within the target time period.
[0054] Individualized baselines may include the following:
[0055] (1) Facial expression baseline: Facial videos of the mother in a relatively relaxed or intermittent state of uterine contractions were collected. Facial expression baseline features were extracted by a pre-trained face video encoder to measure the differences in facial expression changes such as frowning, closing eyes, deepening of nasolabial folds, and facial muscle tension relative to the individual baseline within the target time period.
[0056] (2) Baseline respiratory rhythm: Record the respiratory rate, amplitude and rhythm stability of the parturient in a relatively stable state, and use it to measure changes such as shallow and rapid breathing, breath-holding, hyperventilation or rhythm disorder within the target time period;
[0057] (3) Sound baseline: Collect the ambient noise of the delivery room and the sound characteristics of the mother in a non-pain state to reduce the impact of ambient noise, normal conversation or equipment noise on the recognition of pain-related vocalizations;
[0058] (4) Baseline of facial expression during exertion: Instruct the mother to take a deep breath and hold it, then grip the hand gripper with maximum force for 30-60 seconds, during which she can breathe, to simulate the exertion process during the second stage of labor. Collect facial videos of the mother gripping the hand gripper with maximum force for 30-60 seconds, and extract 512-dimensional facial expression feature vectors during exertion using the MARLIN encoder, which will serve as a reference baseline for pain recognition during exertion in the second stage of labor;
[0059] (5) Baseline respiratory rhythm: Instruct the parturient to take a deep breath and hold it, then grip the hand gripper with maximum force for 30-60 seconds, during which breathing is allowed, to simulate the pushing process of the second stage of labor. Record the mean respiratory rate, mean amplitude, and rhythm stability index when the parturient grips the hand gripper with maximum force for 30-60 seconds, to identify changes in respiratory rate, amplitude, and rhythm caused by pushing and pain during the second stage of labor;
[0060] (6) Sound baseline: Instruct the parturient to take a deep breath and hold it, then grip the hand gripper with maximum force for 30-60 seconds, during which she can breathe, to simulate the pushing process of the second stage of labor. Collect the acoustic characteristics of the parturient when gripping the hand gripper with maximum force for 30-60 seconds, which will be used to judge the subsequent pushing and pain-related vocalizations in the second stage of labor.
[0061] (7) In cases where it is not possible to safely and compliantly collect the relevant baseline for exertion, the system can use only the resting baseline and the inter-contraction baseline for individualized calibration, without requiring the mother to perform additional exertion actions;
[0062] (8) The above baseline features can be quantified to differentiate from the modal time series features within the target time period and used as auxiliary inputs for the adaptive modal gating module or output module.
[0063] Step S1.2 is the calculation of individualized calibration parameters. The system can calculate the difference quantification features of each modality based on the above baseline features, such as the difference, ratio, distance or normalized residual between the target time period features and the corresponding baseline features, to reduce individual differences among different mothers in terms of facial morphology, breathing habits and vocalization habits.
[0064] This mechanism helps reduce the impact of differences in facial morphology, breathing habits, and vocalization habits among different mothers in resting and exertion states on the model output, thereby improving the generalization ability of cross-maternal pain assessment.
[0065] Breathing waveforms were acquired using a non-contact breathing monitoring device at a sampling rate of 100 Hz, which detects minute movements of the mother's chest and abdomen based on video, radar, or infrared technology. Audio signals were acquired using a microphone array at a sampling rate of 16 kHz, containing 2 to 4 omnidirectional microphones to reduce ambient noise interference.
[0066] Simultaneously, uterine contraction pressure curves are acquired from fetal heart rate monitors or uterine contraction pressure sensors. The uterine contraction pressure signals are used as periodic stimulation signals to determine the peak contraction time. All signals are synchronously recorded using a unified timestamp to ensure the correspondence between different modal signals in the time dimension. Step S2 involves peak contraction detection and target time period extraction. Figure 2 As shown, the peak moment of the uterine contraction pressure signal is used as the time anchor point to extract biological behavior signals within a preset time window before and after the peak. In one example, a 10-second time window from 5 seconds before the peak of the uterine contraction to 5 seconds after the peak can be extracted.
[0067] Within the target time period, facial image sequences, respiratory waveform segments, and sound segments are simultaneously captured to form a synchronous multimodal data segment. The peak value of uterine contractions can be identified using a sliding window peak detection algorithm, and the peak detection threshold can be dynamically adjusted based on the individual mother's baseline pressure or the trend of uterine contraction pressure changes. This time anchoring mechanism is used to establish a clear correspondence between pain assessment results and the uterine contraction stimulation process.
[0068] Step S2 is an optional alert for abnormal pain-contraction correlation. During the time-series alignment process, the system can monitor the synchronization relationship between pain assessment results and uterine contraction pressure signals. When the pain assessment result corresponding to the pain peak deviates from the uterine contraction peak time for a long period, or when a high pain assessment result persists during the contraction interval, the system can output an auxiliary alert to the medical staff terminal and display the correspondence between the pain assessment result and the uterine contraction pressure curve. This alert is used to assist medical staff in further judgment and does not directly output a specific disease diagnosis conclusion.
[0069] Through the aforementioned time anchoring and anomaly alerting mechanism, the system can perform periodic assessments of contraction-related pain and also alert to pain assessment patterns inconsistent with the contraction cycle. Step S3 is preprocessing. Face detection, alignment, and cropping are performed on facial expression videos within the target time period; respiratory signals are filtered and normalized; and audio signals are denoised, framed, windowed, and have acoustic features extracted.
[0070] The face image sequence can be cropped and scaled to a preset resolution to meet the input requirements of the pre-trained face video encoder. The breathing waveform is filtered and normalized, for example, using a bandpass filter of 0.1 to 1 Hz to remove baseline drift and high-frequency noise, retaining the main frequency components of the breathing rhythm, and then performing zero-mean normalization on the filtered breathing waveform.
[0071] The system performs denoising, pre-emphasis, framing, and windowing on the audio signal, and can extract at least one of Mel-spectrum features, Mel-frequency cepstral coefficients, energy features, or fundamental frequency features. These acoustic features are used as input to the audio coding network to extract temporal features of the audio signal.
[0072] Step S3 is modal temporal feature extraction. For example... Figure 1 As shown, facial expression video is input to a pre-trained face video encoder, and the output retains the temporal features of facial expression. Breathing waveform is input to at least one of a one-dimensional convolutional network, a recurrent neural network, a temporal convolutional network, or a Transformer encoder, and the output is breathing temporal features. The sound signal is input to a sound coding network after denoising, framing, windowing, and acoustic feature extraction, and the output is sound temporal features.
[0073] Step S4 is adaptive gating. For example... Figure 1 As shown, at least two modal temporal features are input into an adaptive gating module to generate modal weights corresponding to each modality. These modal weights can dynamically change with the target time period, production stage, signal quality, and individualized baseline differences, and are used to adjust the contributions of different modalities in subsequent cross-modal fusion.
[0074] Steps S5 and S6 involve cross-modal fusion and pain peak query aggregation. For example... Figure 1 As shown, the system weights the temporal features of each modality based on the modal weights, and performs cross-modal temporal interaction modeling through the feature fusion module to obtain fused temporal features; subsequently, the pain peak query aggregation module uses the learnable pain peak query vector to perform attention aggregation on the fused temporal features to obtain the pain peak representation.
[0075] Step S7 involves pain assessment output and continuous monitoring. For example... Figure 1As shown, the output module outputs pain level classification results and / or continuous pain intensity scores based on pain peak characterization. Repeating the above steps for consecutive contraction cycles during labor generates a pain assessment sequence corresponding to each contraction cycle, which can be used to assist in labor pain management.
[0076] This embodiment achieves labor pain assessment based on adaptive fusion of multimodal signals through the above steps. The main technical approach of this embodiment is not to simply concatenate and classify the features of each modality, but rather to determine the target time period using the peak value of uterine contractions as the time anchor point, adjust the contributions of different modalities using adaptive modal gating, and obtain a representation of the pain peak at the uterine contraction cycle level through pain peak query aggregation. This ensures that the assessment results have a clear correspondence with the progress of labor.
[0077] In one embodiment, the above structure enables multimodal complementary verification and maintains high assessment stability even when any of the facial, respiratory, or vocal modalities are disturbed. This method provides objective, continuous, and real-time auxiliary decision-making information for labor pain management.
[0078] Example 2
[0079] This embodiment, based on Embodiment 1, further clarifies an optional configuration for facial feature extraction using the MARLIN architecture. MARLIN is a face video encoder pre-trained on a mask reconstruction task. Through self-supervised learning on large-scale unlabeled face videos, it can provide stable facial dynamic representations in small-sample medical scenarios.
[0080] In one implementation, the input face video clip can be scaled to a preset resolution, such as 224×224; the spatial size and temporal span of the spatiotemporal patch can be determined according to the encoder configuration. Each spatiotemporal patch is linearly projected to form an embedding vector, which is then input into a multi-layer Transformer encoder for spatiotemporal dependency modeling.
[0081] The masking ratio is 75%. During the pre-training phase, 75% of the spatiotemporal tokens are randomly masked, and only the remaining 25% of visible tokens are input into the encoder. Self-supervised learning is performed by reconstructing the masked video blocks. The high masking ratio forces the encoder to learn robust facial structure and dynamic expression representations. The encoder has a depth of 12 layers. Each Transformer encoder layer includes a multi-head self-attention mechanism and a feedforward network.
[0082] In one implementation, the encoder can employ a 12-layer Transformer structure, with each layer containing a multi-head self-attention mechanism and a feedforward network. The embedding dimension can be 512-dimensional, 768-dimensional, or other preset dimensions. Different embedding dimensions only represent different model sizes and do not affect the technical solution of this invention for extracting temporal features of facial expressions using a pre-trained face video encoder.
[0083] When performing transfer learning on obstetric pain assessment tasks, one can choose to freeze all parameters of the pre-trained encoder, partially unfreeze high-level parameters, or use efficient fine-tuning methods such as adapters and low-rank parameters, depending on the data size and risk of overfitting, so that the pre-trained facial dynamic representation can adapt to the data distribution of the target task.
[0084] Example 3
[0085] This embodiment, based on Embodiment 1, further clarifies the detailed configuration of the adaptive modal gating and cross-modal fusion module. This module is used to model the complementary relationship between facial expression, breathing, and voice modalities, and to generate modal weights based on the reliability of different modalities within the target uterine contraction time window, thereby obtaining fused temporal features.
[0086] The adaptive modal gating and cross-modal fusion module is configured as follows. First, the temporal features of facial expression, breathing, and voice are mapped to a unified feature dimension through a linear projection layer. Then, the temporal features of each modality are input into a multilayer perceptron to generate corresponding gating scores, and the facial modality weights, breathing modality weights, and voice modality weights are obtained by Softmax normalization.
[0087] The temporal features of each modality, after being weighted by modal weights, are input into a cross-modal Transformer encoder for temporal interaction modeling to obtain fused temporal features. The cross-modal Transformer encoder may include two layers, each containing a multi-head self-attention mechanism, residual connections, layer normalization, and a feedforward network to learn the correlation and complementarity relationships between different modalities within a unified feature space.
[0088] Each layer can contain multiple attention heads to learn the correlations between temporal features across different subspaces. The self-attention mechanism calculates attention weights based on linear transformations of the query, key, and value, and outputs the temporal features after cross-modal interaction. This attention calculation is used to model the correlations between temporal features at different time points and across different modalities.
[0089] To improve model training stability, the cross-modal fusion module can be configured with structures such as residual connections, layer normalization, feedforward networks, and Dropout. These structures are optional implementations of the cross-modal fusion module and do not alter its function of fusing temporal features based on modal weights.
[0090] In the multi-head self-attention module and feedforward network module, the Dropout operation can be used to reduce the risk of overfitting. The fused temporal features output by the cross-modal Transformer encoder are further input into the pain peak query aggregation module. The pain peak query aggregation module calculates the attention weights at each time position through the learnable query vector and performs weighted aggregation on the fused temporal features to obtain the pain peak representation.
[0091] Example 4
[0092] This embodiment, based on Embodiment 1, further clarifies the design of the loss function during training. To simultaneously achieve joint learning of pain level classification and pain intensity regression, the model can employ a multi-task loss function for end-to-end optimization. The multi-task loss function can be expressed as: L = λcLc + λrLr, where Lc is the pain level classification loss, Lr is the pain intensity regression loss, and λc and λr are weight coefficients. The classification loss can employ cross-entropy loss, Focal Loss, or class balance loss, while the regression loss can employ Huber Loss, smoothing L1 loss, or mean squared error loss. During training, the weight coefficients can be set according to the sample class distribution, the level of rating noise, and the needs of the clinical task.
[0093] Example 5
[0094] This embodiment provides an optional method for acquiring respiratory signals. Unlike Embodiment 1, this embodiment acquires video of the chest and abdominal region using an RGB camera and obtains the respiratory waveform based on a video respiratory extraction algorithm. This approach reduces reliance on additional respiratory monitoring equipment, but it requires a high degree of maternal positioning and visibility of the chest and abdomen, making it suitable for early labor or other scenarios where the patient's position is relatively stable.
[0095] Specifically, the video breathing extraction algorithm based on an RGB camera is as follows: While capturing facial expression video, video of the mother's chest and abdomen region is also captured. The chest and abdomen region is automatically identified using a human pose estimation algorithm or manually calibrated by clinical personnel. The video resolution for the chest and abdomen region is the same as that for the facial expression video, with a frame rate of 30fps. Temporal pixel intensity statistics are performed on the chest and abdomen region, and the Region of Interest (ROI) within this region is selected.
[0096] The mean pixel intensity of each frame within the ROI is calculated. Due to the rise and fall of the chest and abdomen caused by respiratory movements, the mean pixel intensity exhibits periodic changes over time. The temporal mean pixel intensity sequence is bandpass filtered. A bandpass filter ranging from 0.1 to 1 Hz is used to remove baseline drift and high-frequency noise, preserving the main frequency components of the respiratory rhythm. The filtered signal is the extracted respiratory waveform. The extracted respiratory waveform is then normalized.
[0097] Zero-mean normalization was employed to ensure comparability of respiratory signals from different postpartum women. The normalized respiratory waveforms can be input into at least one of a one-dimensional convolutional network, a recurrent neural network, a temporal convolutional network, or a Transformer encoder for feature extraction, outputting respiratory temporal features. These respiratory temporal features are used to characterize dynamic changes related to pain response, such as respiratory rate, amplitude, rhythm stability, breath-holding, shallow and rapid breathing, or hyperventilation.
[0098] Example 6
[0099] This embodiment provides an optional sound feature extraction method. Unlike schemes that use only a single acoustic feature, this embodiment can use log-Mel spectrum features as one of the inputs to the sound coding network. Log-Mel spectrum retains more spectral detail information and can improve the robustness of sound temporal features when there is high noise in the delivery room environment, but the computational cost is relatively high. Specifically, the sound signal is subjected to denoising, pre-emphasis, framing, and windowing processing; the pre-emphasis coefficient can be set to 0.97, the framing parameters can be a frame length of 25 milliseconds and a frame shift of 10 milliseconds, and a Hamming window can be used. Subsequently, a short-time Fourier transform is performed on each frame of signal and the power spectrum is calculated.
[0100] A Mel filter bank is applied to the power spectrum. The Mel filter bank can contain a predetermined number of triangular filters, such as 128 filters, distributed across the Mel scale. Each filter performs a weighted summation of the power spectrum to obtain the Mel spectrum; then, the logarithm of the Mel spectrum is taken to obtain the logarithmic Mel spectral characteristics.
[0101] Log-Melogram spectral feature sequences can be input into a sound coding network for feature extraction. This sound coding network can employ a one-dimensional convolutional network, a recurrent neural network, a temporal convolutional network, a Transformer encoder, or a combination thereof, outputting temporal sound features. These temporal sound features are used to capture pain-related changes in sound intensity, fundamental frequency, spectral distribution, and vocal rhythm.
[0102] Example 7
[0103] This embodiment provides an optional experimental design for comparative verification with existing technologies to illustrate the synergistic effect of the various technical modules of the present invention. Experimental data may include facial expression videos, respiratory waveforms, sound signals, uterine contraction pressure signals, and corresponding pain annotation results. The dataset may be divided into training, validation, and test sets.
[0104] Comparison schemes may include: Scheme A is a multimodal pain assessment scheme that does not use contraction time anchors and adopts a fixed fusion method; Scheme B is a scheme that adopts other time-series fusion methods but does not set contraction peak time windows and pain peak query aggregation; Scheme C is the scheme proposed in this invention, that is, a scheme that adopts contraction peak time anchoring, adaptive modal gating, cross-modal time-series fusion and pain peak query aggregation.
[0105] In an example experiment, the accuracy of pain level classification, continuous error in pain intensity scoring, consistency with clinical VAS scores, correspondence to uterine contraction cycles, and inference speed can be evaluated. The evaluation results can be used to demonstrate that the present invention's approach, compared to approaches that do not use contraction time anchors or do not employ adaptive modal gating, offers improvements in pain assessment accuracy, labor synchronicity, and robustness.
[0106] The above comparative logic shows that introducing a peak contraction time anchor point can synchronize pain assessment results with the progress of labor; introducing adaptive modal gating can dynamically adjust the contribution of each modality when the modality is disturbed; and introducing pain peak query aggregation can enhance the representation ability of the pain peak state at the contraction cycle level.
[0107] Regarding inference speed, Scheme C can meet the requirements of real-time or near-real-time assessment while ensuring assessment accuracy. Further analysis of classification performance for different pain levels shows that a training strategy employing class imbalance helps improve the recall rate of moderate and severe pain samples.
[0108] Regarding the robustness of VAS score prediction, the use of robust regression loss can reduce the impact of subjective rating noise on model training. This embodiment, through comparison with existing technologies, illustrates that the advantages of this invention stem from the synergistic effects of contraction time anchoring, adaptive modal gating, pain peak query aggregation, and multi-task training.
[0109] The above experimental design can be used to verify the technical advantages of this invention. This invention establishes a synchronous correlation model between stimulus and response by introducing uterine contraction signals as time anchors; dynamically adjusts modal weights based on the reliability of different modalities within the target time period using an adaptive modal gating module; extracts uterine contraction cycle-level pain peak representations from fused temporal features using a pain peak query aggregation module; and simultaneously supports pain level classification and continuous pain intensity scoring through multi-task training.
[0110] The aforementioned technical features work together to enable the present invention to improve existing solutions in terms of pain assessment accuracy, labor synchronicity, and assessment stability, providing objective and continuous auxiliary decision-making information for labor pain management.
[0111] In addition to the above embodiments, the present invention also includes the following variants. Variant 1: The uterine contraction pressure signal can be acquired by a fetal heart rate monitor, a uterine contraction pressure sensor, or other devices capable of reflecting changes in the uterine contraction cycle. Variant 2: Facial feature extraction can be performed using MARLIN, VideoMAE, TimeSformer, SlowFast, or other pre-trained face video encoders based on video Transformer or spatiotemporal modeling networks. Variant 3: Temporal feature extraction of respiration and sound can be performed using one-dimensional convolutional networks, recurrent neural networks, temporal convolutional networks, Transformer encoders, or combinations thereof. Variant 4: The adaptive modal gating module can be implemented using a multilayer perceptron, attention network, expert hybrid network, or a weight generation network based on signal quality estimation, and modal weights can be generated through a normalization function. Variant 5: The pain peak query aggregation module can be implemented using a learnable query vector combined with temporal attention, attention pooling, or a weighted aggregation method based on the distance to the uterine contraction peak. Variant 6: The output module can simultaneously output pain level classification results and continuous pain intensity scores, or output only one of the results according to clinical needs.
[0112] This invention also provides a labor pain assessment system based on adaptive fusion of multimodal signals. The system includes a multimodal signal acquisition module, a target time period extraction module, a feature extraction module, an adaptive modal gating module, a cross-modal fusion module, a pain peak query aggregation module, and an output module. The multimodal signal acquisition module is used to acquire at least two of the following: facial expression signals, respiratory signals, and sound signals, as well as uterine contraction pressure signals; the target time period extraction module is used to determine the peak time of uterine contractions based on the uterine contraction pressure signals and extract biological behavioral signals within a preset time window before and after the peak; the feature extraction module is used to output the temporal features of each modality; the adaptive modal gating module is used to generate modal weights; the cross-modal fusion module is used to obtain the fused temporal features; the pain peak query aggregation module is used to extract the pain peak representation; and the output module is used to output the pain level classification results and / or continuous pain intensity scores.
[0113] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for assessing labor pain based on adaptive fusion of multimodal signals, including steps such as data acquisition, target time period extraction, preprocessing, modal temporal feature extraction, adaptive modal gating, cross-modal fusion, pain peak query aggregation, and pain assessment output.
[0114] The computer-readable storage medium can be a read-only memory, random access memory, disk, optical disk, or other non-volatile storage medium. The program can be deployed on medical devices, mobile terminals, edge computing devices, or cloud computing platforms, and can provide pain assessment results to medical terminals, clinical decision support systems, or closed-loop intelligent analgesia systems as auxiliary decision input.
[0115] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for assessing labor pain based on adaptive fusion of multimodal signals, characterized in that, Includes the following steps: Collect at least two biological behavioral signals and one periodic stimulation signal, wherein the biological behavioral signals include at least two of facial expression signals, respiratory signals and sound signals, and the periodic stimulation signal includes uterine contraction pressure signals; The peak time of uterine contractions is determined based on the periodic stimulation signal, and it is used as a time anchor point. The biological behavior signals of the parturient within a preset time window before and after the peak are extracted as the biological behavior signals within the target time period. Temporal features are extracted from at least two biological behavioral signals within the target time period to obtain the corresponding modal temporal features; The modal temporal features are input into the adaptive modal gating module to generate modal weights corresponding to each modality. Based on the modal weights, the modal temporal features are fused across modalities to obtain fused temporal features. The pain peak representation is extracted from the fused temporal features through the pain peak query aggregation module; Pain assessment results are output based on the pain peak characterization.
2. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The biological behavioral signals include facial expression signals. Feature extraction of the facial expression signals includes: face detection, alignment and cropping of facial videos within the target time period, and extraction of temporal features of facial expressions through a pre-trained face video encoder.
3. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The biological behavioral signals include respiratory signals. Feature extraction of the respiratory signals includes: filtering and normalizing the respiratory waveform within the target time period, and extracting respiratory temporal features through at least one of a one-dimensional convolutional network, a recurrent neural network, a temporal convolutional network, or a Transformer encoder.
4. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The biological behavior signal includes a sound signal. Feature extraction of the sound signal includes: denoising, framing and windowing the sound signal within the target time period, extracting at least one of Mel spectrum features, Mel frequency cepstral coefficients, energy features or fundamental frequency features, and extracting sound temporal features through a sound coding network.
5. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The adaptive modal gating module receives at least two modal temporal features and outputs corresponding modal weights. The modal weights are used to adjust the contribution of each modal temporal feature in the cross-modal fusion process.
6. The method according to claim 5, characterized in that, The adaptive modal gating module includes a multilayer perceptron and a normalization function. The multilayer perceptron generates modal weights based on at least two modal temporal features, and the normalization function is used to ensure that each modal weight satisfies the normalization constraint.
7. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The pain peak query aggregation module includes a learnable pain peak query vector, which is used to perform cross-time attention aggregation on the fused temporal features to obtain the pain peak representation. The pain peak query aggregation module generates temporal attention weights based on the correlation between the pain peak query vector and each time position feature in the fused temporal features, and performs weighted aggregation on the fused temporal features based on the temporal attention weights.
8. The method for assessing labor pain based on adaptive fusion of multimodal signals according to claim 1, characterized in that, The method also includes an individualized baseline calibration step: collecting facial expression baseline, respiratory rhythm baseline, and voice baseline of the mother at rest during the early stages of labor; In pain assessment, the biological behavioral signal characteristics within the target time period are quantified to differentiate from the corresponding individualized baseline characteristics, and pain assessment is performed based on the results of the differential quantification.
9. A labor pain assessment system based on multimodal signal adaptive fusion, characterized in that, include: A multimodal signal acquisition module is used to acquire at least two biological behavioral signals and periodic stimulation signals. The biological behavioral signals include at least two of facial expression signals, respiratory signals, and sound signals. The periodic stimulation signals include uterine contraction pressure signals. The target time period extraction module is used to determine the uterine contraction cycle or the peak time of uterine contraction based on the periodic stimulation signal, and to extract the biological behavior signal within the corresponding target time period from the biological behavior signal based on the uterine contraction cycle or the peak time of uterine contraction. The feature extraction module is used to extract temporal features from at least two biological behavioral signals within the target time period to obtain the corresponding modal temporal features. An adaptive modal gating module is used to generate modal weights corresponding to each modality based on the modal temporal features. A cross-modal fusion module is used to perform cross-modal fusion of the modal temporal features based on the modal weights to obtain fused temporal features; The pain peak query and aggregation module is used to extract pain peak representations from the fused temporal features; The output module is used to output pain assessment results based on the pain peak characterization.
10. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the steps of the labor pain assessment method based on multimodal signal adaptive fusion as described in claims 1-8.
Citation Information
Patent Citations
Pain assessment system and method based on multi-modal physiological signals
CN120899167A