A deep learning-based multi-modal real-time intelligent pain assessment auxiliary system

By using a multimodal data acquisition and consistent inference device, the simultaneous acquisition and feature extraction of facial video and near-field speech were achieved. An anchor point set was constructed, and soft alignment and interference recognition were performed. This solved the real-time and accuracy problems of single-modal pain assessment and significantly improved the stability and reliability of pain assessment.

CN121237442BActive Publication Date: 2026-06-26SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
Filing Date
2025-09-30
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing single-modal pain assessment methods have problems in clinical applications, such as insufficient real-time performance, susceptibility to subjective influence of observers, difficulty in comprehensively capturing diverse pain manifestations, and insufficient analysis of dynamic features across time slices, resulting in a high misjudgment rate.

Method used

A deep learning-based multimodal real-time intelligent pain assessment assistance system is adopted. By simultaneously acquiring facial video and near-field speech, multidimensional dynamic features are extracted, an anchor point set is constructed, and combined with short-window rapid scoring and expanded-window verification mechanisms, key event-driven soft alignment and interference identification are achieved, non-pain interference is eliminated, and pain segments are accurately identified.

Benefits of technology

It improves the stability and reliability of pain assessment, reduces the risk of misjudgment, meets the needs of real-time pain identification, facilitates structured recording, and facilitates subsequent medical intervention and data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237442B_ABST
    Figure CN121237442B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of pain assessment, and particularly relates to a multi-modal real-time intelligent pain assessment auxiliary system based on deep learning. The system comprises a multi-modal data acquisition device and a two-section consistency reasoning device. The multi-modal data acquisition device synchronously acquires facial video and near-field voice under a unified time reference, cuts the time axis to form a unified time slice sequence, and outputs the aligned unified time slice sequence and anchor point set. The two-section consistency reasoning device, under the constraint of the aligned unified time slice sequence and anchor point set, first performs short window scoring on each time slice to obtain a pain score and uncertainty, confirms the candidate segment on the score timeline, and outputs the real-time pain score, confidence, and start and end time of the confirmed pain segment. The present application improves the assessment stability and reliability while ensuring real-time response, effectively eliminates non-pain interference, and accurately confirms the pain segment by combining the double threshold rule and anchor point neighborhood consistency verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pain management technology, specifically relating to a multimodal real-time intelligent pain assessment and assistance system based on deep learning. Background Technology

[0002] Pain is one of the most common symptoms in clinical medicine that directly impacts patients' quality of life, and its assessment results are directly related to the formulation and adjustment of analgesia regimens. However, pain is highly subjective, and different patients have significant differences in tolerance, expressive ability, and physiological responses, making it difficult to obtain objective and quantifiable indicators by relying solely on patient self-reporting. In current clinical practice, medical staff often use a combination of facial expression observation, verbal description, and physiological indicator monitoring to assist in assessing the degree of pain in patients. For example, by observing changes in the patient's eyebrows and eyes, mouth shape, muscle tension, and combining this with physiological signals such as heart rate, blood pressure, and respiratory rate, the degree of pain can be inferred. These methods improve the objectivity of the assessment to some extent, but they still have problems such as excessive human intervention, insufficient real-time monitoring, and susceptibility to the influence of the observer's subjective experience.

[0003] In recent years, with the rapid development of computer vision and speech processing technologies, some studies have attempted to introduce automated pain recognition methods based on a single modality. For example, in the video modality, face detection, keypoint tracking, and expression classification models based on convolutional neural networks are used to identify facial feature changes related to pain; in the speech modality, parameters such as short-time energy, fundamental frequency, formants, and spectral features are extracted, and a deep neural network classifier is used to determine whether the patient's speech contains pain-related speech patterns. These single-modal methods have achieved some success in laboratory settings, but still face significant limitations in practical clinical applications. The main problems include: single-modality methods struggle to comprehensively capture the diverse manifestations of pain; video modalities may be affected by facial occlusion and lighting changes; and speech modalities may fail due to patient silence or background noise. Most methods only independently judge single-frame images or short speech segments, lacking dynamic feature analysis across time slices, making it difficult to distinguish between transient interference and genuine pain responses. Non-pain-related transient behaviors such as blinking and plosive sounds often cause changes similar to pain responses at the feature level, leading to misjudgments. Existing single-mode methods generally lack the ability to identify and correct interference. Summary of the Invention

[0004] The main objective of this invention is to provide a deep learning-based multimodal real-time intelligent pain assessment assistance system. By synchronously acquiring and aligning facial video and near-field speech under a unified time reference, multidimensional dynamic features are extracted and anchor point sets are constructed by combining blink initiation and plosive initiation points, achieving key event-driven soft alignment and interference recognition. A two-stage consistency inference mechanism of short-window rapid scoring and expanded-window verification is employed, along with quantified uncertainty, improving assessment stability and reliability while ensuring real-time response. Combining dual-threshold rules and anchor point neighborhood consistency verification effectively eliminates non-pain interference and accurately identifies pain segments, thus significantly improving multimodal fusion accuracy, interference suppression capability, real-time output, and clinical applicability compared to existing technologies.

[0005] To address the aforementioned technical problems, this invention provides a deep learning-based multimodal real-time intelligent pain assessment assistance system, comprising: a multimodal data acquisition device and a two-stage consistency inference device; wherein, the multimodal data acquisition device synchronously acquires facial video and near-field speech under a unified time reference, segments the time axis to form a unified time slice sequence; generates video-side feature vectors and speech-side feature vectors respectively according to a fixed feature list; generates blink start anchors and plosive start anchors according to fixed judgment rules to form an anchor set; performs soft alignment in the neighborhood of anchors based on the anchor set and performs linear interpolation to fill in missing time slices, so as to output the aligned unified time slice sequence and anchor set; the two-stage consistency inference device, under the constraints of the aligned unified time slice sequence and anchor set, first performs short-window scoring on each time slice to obtain a pain score and uncertainty; when the uncertainty is higher than a threshold, performs expanded-window verification to update the pain score and uncertainty; then performs consistency verification based on the anchor set to obtain candidate segments; completes the confirmation of candidate segments on the score timeline and outputs the real-time pain score, confidence level, and start and end times of the confirmed pain segments.

[0006] Furthermore, facial video is continuously acquired at a fixed frame rate of 30 frames per second, and near-field speech is continuously acquired at a fixed sampling rate of 16000 Hz. A timestamp in milliseconds is written for each frame of facial video and each frame of near-field speech, and the timestamps come from the same time base. The time axis is divided into equal-length time slices, with a time slice length of 200 milliseconds. The start of the time slice is advanced in multiples of 200 milliseconds from the start of the acquisition. Within each time slice, all facial video and near-field speech within its time range are collected, and a time slice index, a start timestamp, and an end timestamp are added to the time slice, thereby forming a unified time slice sequence.

[0007] Furthermore, the process by which the multimodal data acquisition device generates video-side feature vectors and speech-side feature vectors based on a fixed feature list includes: performing face localization and key point tracking in each frame of facial video, using the distance between the centers of the two eyes as the scale reference; calculating the segment mean and median rate of change of the following quantities in each time slice: eyelid opening and closing, eyebrow-eye distance, lip gap, corner of mouth upward angle, nasal wing expansion amplitude, head translation amplitude, and facial region optical flow amplitude, to form the video-side feature vector; dividing the near-field speech into frames according to a fixed window length and step size; calculating the segment mean and median rate of change of the following quantities in each time slice: short-time energy, zero-crossing rate, spectral centroid, bandpass energy ratio, spectral flatness, fundamental frequency presence index, and short-time jitter index, to form the speech-side feature vector.

[0008] Furthermore, the process by which the multimodal data acquisition device generates blink initiation anchor points and plosive sound initiation anchor points according to fixed judgment rules and forms an anchor point set includes: on a unified time-slice sequence, generating only two types of anchor points and writing them into the anchor point set according to the following rules: Blink initiation: When, within a certain time slice, the eyelid opening and closing degree is detected to decrease by 0.2 relative to the preset baseline of the metric for a duration of no more than 3 frames, and subsequently recovers by 0.2 relative to the preset baseline of the metric for a duration of no more than 3 frames, while the eyebrow-eye distance in that time slice... When the median rate of change of a segment is positive, the time slice is marked as the blink start anchor point; plosive start point: when the average value of the short-time energy segment is detected to increase by a factor of 2 or more relative to the preset baseline of the metric within a certain time slice, and the average value of the bandpass energy ratio segment increases by an amount of 0.3 or more relative to its preset baseline, and the average value of the spectral flatness segment increases by an amount of 0.2 or more relative to its preset baseline, the time slice is marked as the plosive start anchor point; each record in the anchor point set contains the anchor point type, time slice index, start timestamp, and end timestamp.

[0009] Furthermore, the two-stage consistency inference device obtains the pain score through the following process: Under the constraints of the aligned unified time-slice sequence and the anchor point set, the specific steps for implementing short-window scoring are as follows: For each target time-slice, a short window of length 5 time-slices is taken on the unified time-slice sequence centered on that time-slice; for each feature quantity listed in the fixed feature list, the historical median is calculated within 30 seconds before the target time-slice as the baseline, and the median absolute deviation of the historical interval relative to the baseline is calculated, and the larger of the median and 0.001 is taken as the scale; the absolute value of the difference between the current value of the feature quantity in each time-slice within the short window and the baseline is divided by the scale to obtain the deviation ratio sequence, and the median of the sequence is taken as the short-window representative of the feature quantity. The deviation ratio indicates whether a deviation ratio greater than or equal to 2 is considered abnormal, otherwise it is considered normal. Conversely, if there is a blink start point anchor point within the short window and the feature quantity associated with blinking shows abnormality in no less than 3 time slices, while the feature quantity associated with plosive sounds shows abnormality in only 1 time slice, then the abnormality in that single time slice is considered transient interference and treated as normal. If there is a plosive sound start point anchor point within the short window and the feature quantity associated with plosive sounds shows abnormality in no less than 3 time slices, while the feature quantity associated with blinking shows abnormality in only 1 time slice, then a symmetrical reclassification is performed. After the reclassification is completed, the number of abnormalities in all feature quantities in the target time slice is counted to the total number of feature quantities, and the pain score is set to 10 multiplied by the number of abnormalities divided by the total number of feature quantities.

[0010] Furthermore, the two-stage consistency inference device obtains the uncertainty through the following process: the uncertainty is set to 1 and then subtracted from the ratio of the number of features that are consistent with the target time slice and appear at least 3 times consecutively within the short window to the total number of features; the obtained pain score and uncertainty are written into the score timeline in the order of the timestamps.

[0011] Furthermore, the two-stage consistency inference device generates candidate segments by using a dual-threshold rule on the score timeline. The dual-threshold rule includes an entry threshold and an exit threshold. The entry threshold is used to determine the start of a candidate segment when the pain score rises, and the exit threshold is used to determine the end of a candidate segment when the pain score falls back. The entry threshold is higher than the exit threshold, and the minimum duration of a candidate segment is not less than two time slices.

[0012] Furthermore, the two-stage consistency inference device searches for anchor points in the anchor point neighborhood near the starting point of each candidate segment. The anchor point neighborhood is defined as a fixed symmetrical range in time slices, and the retrieved anchor points are arranged in ascending order of time slice index. If only the blink starting point anchor point is retrieved in the anchor point neighborhood and the short-time energy and spectral flatness of the speech-side feature vector in the candidate segment do not show a common rising pattern consistent with the plosive sound, then the candidate segment is marked as a blink interference candidate. If only the plosive sound starting point anchor point is retrieved in the anchor point neighborhood and the eyelid opening and closing degree and facial optical flow amplitude of the video-side feature vector in the candidate segment do not show a common change pattern consistent with the painful expression, then the candidate segment is marked as a plosive sound interference candidate. If both the blink starting point anchor point and the plosive sound starting point anchor point are retrieved in the anchor point neighborhood, then the candidate segment is marked as a mixed interference candidate. If no anchor point is retrieved in the anchor point neighborhood, then the candidate segment is marked as an anchorless consistency candidate.

[0013] Furthermore, the system receives a set of candidate segments generated and labeled according to a dual-threshold rule; merges adjacent candidate segments with an interval of no more than one time slice; performs rapid verification: segments labeled as blink interference candidates or plosive interference candidates are directly eliminated; segments labeled as mixed interference candidates are retained only when both video and speech anomalies occur simultaneously within the segment; segments labeled as anchorless consistency candidates are retained; boundary determination is performed on the retained candidate segments: based on the score timeline, the time slice that first reaches the entry threshold is taken as the confirmation start point, and the last time slice before the first two consecutive time slices below the exit threshold are taken as the confirmation end point; if the duration between the confirmation start and end is less than two time slices, the segment is canceled; the system outputs the real-time pain score and real-time confidence level for each time slice, with the real-time confidence level being 1 minus the uncertainty; for each confirmed pain segment, the system outputs the confirmation start and end times, and uses the arithmetic mean of the segment's built-in confidence level as the segment confidence level; when the interval between two confirmed pain segments is no more than one time slice, they are merged and the start and end times and segment confidence levels are updated.

[0014] This invention provides a deep learning-based multimodal real-time intelligent pain assessment assistance system with the following advantages: It synchronously acquires facial video and near-field speech under a unified time reference and segments them into equal-length time slice sequences, achieving precise alignment of cross-modal data and effectively avoiding the fusion failure problem caused by temporal misalignment of multimodal data in existing technologies. By extracting multidimensional dynamic features from the video and speech sides under a fixed feature list and constructing an anchor point set based on the blink initiation point and plosive sound initiation point, this invention achieves soft alignment and interference identification of key events at the time slice level, significantly improving the ability to suppress interference from non-painful behaviors. A two-stage consistency inference device, combined with short-window rapid scoring and expanded-window verification mechanisms, enhances the stability of pain assessment results while ensuring real-time performance. It also introduces an uncertainty quantification method, enabling the system to clearly reflect the judgment confidence level even when signal fluctuations or feature changes are not obvious, reducing the risk of clinical misjudgment. By generating candidate segments through dual-threshold rules and combining them with anchor point neighborhood consistency verification, it achieves rapid elimination of transient interference and accurate confirmation of real pain responses, effectively reducing the false alarm rate. The system can output real-time pain scores, real-time confidence levels, and the start and end times of confirmed pain segments in each time slice. This meets the needs of high-frequency monitoring and supports structured recording of pain events, facilitating subsequent medical intervention and data analysis. In summary, this invention has significant advantages over existing technologies in terms of multimodal synchronous acquisition, interference suppression, real-time assessment, and uncertainty quantification. It can be widely applied in medical scenarios requiring real-time pain identification, such as surgical anesthesia monitoring, critical care, and rehabilitation training. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the system structure of a deep learning-based multimodal real-time intelligent pain assessment and assistance system provided in an embodiment of the present invention.

[0017] Figure 2 A schematic diagram of the experimental curve for detecting the blink origin anchor point provided in an embodiment of the present invention;

[0018] Figure 3 This is a schematic diagram of the experimental curve for detecting the starting point anchor point of the blast sound, provided in an embodiment of the present invention. Detailed Implementation

[0019] The method of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0020] refer to Figure 1 A deep learning-based multimodal real-time intelligent pain assessment and assistance system includes: a multimodal data acquisition device and a two-stage consistency inference device;

[0021] The multimodal data acquisition device starts with a unified time base and maintains strict temporal consistency and repeatable feature representation for facial video and near-field speech in a pipeline consisting of a hardware acquisition layer, a timestamp annotation layer, a time slice organization layer, a fixed feature list calculation layer, and an anchor set generation layer. At the acquisition end, a deep learning model is introduced for key point tracking stabilization, image quality enhancement, speech activity detection, and noise robustness enhancement. However, it does not directly output pain scores. Instead, it ensures that each record written into the unified time slice sequence contains reliably reusable video-side feature vectors and speech-side feature vectors, a time slice index, a start timestamp and an end timestamp, and regular blink start anchors and plosive start anchors, thereby providing a stable data source for the two-stage consistent inference device.

[0022] In practical implementation, after startup, the multimodal data acquisition device simultaneously distributes a unified time base to both image and voice acquisition using the same master clock. Facial video is continuously acquired at a fixed frame rate of 30 frames per second, and a timestamp in milliseconds is recorded the instant the data is written to the buffer. Near-field voice is continuously acquired at a fixed sampling rate of 16000 Hz, and a timestamp in milliseconds is recorded the instant the sample block is written to the buffer. The timeline management module generates time slices in multiples of 200 milliseconds from the start of acquisition and assigns them as time slice indices. Each time slice in the unified time slice sequence stores all facial video and near-field voice data within its time range; the formation of feature vectors on the video side... The device first performs face localization and key point tracking on each frame of facial video. Key point tracking is completed by an embedded deep learning key point regressor and traditional optical flow tracking. The former is used to maintain stable detection of reference points for eyelids, eyebrows, corners of the mouth, nose, and head under conditions of rapid pose changes, partial occlusion, or lighting jitter. The latter is used to update key point positions at low cost between adjacent frames to reduce latency. Scale normalization is completed directly at the frame level based on the distance between the centers of the two eyes. This makes the frame-level measurement of eyelid opening and closing, eyebrow-eye distance, lip gap, corner of mouth upward angle, nose expansion amplitude, head translation amplitude, and facial region optical flow amplitude insensitive to shooting distance and slight zoom.

[0023] In the formation of speech-side feature vectors, the multimodal data acquisition device uses a deep learning speech activity detector and a denoising front-end based on spectral masking to jointly suppress air conditioning noise and indoor echo. Then, it frames the data with a fixed window length and step size and calculates short-time energy, zero-crossing rate, spectral centroid, bandpass energy ratio, spectral flatness, fundamental frequency presence index, and short-time jitter index. All frame-level metrics are aggregated into segment mean and segment change rate median within a unified time-slice sequence according to fixed rules, ensuring that subsequent calculations rely only on the time-slice level data structure. In anchor point set generation, the device generates only blink start anchor points and plosive start anchor points on the unified time-slice sequence according to fixed judgment rules. The blink start pattern recognition on the video side is a combination of the eyelid opening / closing trajectory and eyebrow-eye distance change provided by a deep learning keypoint regressor. The speech-side plosive start-point pattern recognition is supported by synchronous jumps in short-time energy, bandpass energy ratio, and spectral flatness. The device adds anchor point type, time slice index, start timestamp, and end timestamp to the anchor point set for time slices that meet the rules. For soft alignment and linear interpolation completion, the device defines a fixed symmetrical range in the neighborhood of the anchor point with the anchor point set as a constraint. It checks the peak and abrupt change times of the video-side feature vector and the speech-side feature vector within this range. If the peaks on both sides cross the time slice boundary and the offset is not greater than a single frame time, the corresponding frame-level sample is reassigned to the nearest time slice to reduce the cross-modal phase difference. If there is feature loss due to instantaneous occlusion or silence in the time slice, the device performs linear interpolation completion on the same feature between adjacent valid time slices and writes the completion mark for the two-stage consistency inference device to perceive.

[0024] The two-stage consistency inference device takes the aligned unified time-slice sequence and anchor point set as input. First, it uses short-window scoring to stably identify common anomalies in video-side and speech-side feature vectors within a small time context, thereby obtaining the pain score and uncertainty for each time slice. Then, when the uncertainty exceeds a threshold, it performs a window expansion check to introduce a longer range of context to reduce the impact of occasional noise. Subsequently, it performs consistency verification based on the anchor point set, so that instantaneous interference caused by blink start anchor point and plosive start anchor point can be quickly distinguished at the rule level. Finally, it generates candidate segments on the score timeline through dual threshold rules and completes the confirmation of candidate segments, thereby continuously outputting real-time pain scores, confidence levels, and the start and end times of confirmed pain segments. The specific implementation process is as follows: the device reads the unified time-slice sequence slice by slice. For each target time slice, it takes a short window of 5 time slices in length centered on that time slice. For each feature quantity listed in the fixed feature list, the historical median is calculated 30 seconds before the target time slice as the baseline, and the larger of the median absolute deviation of the historical interval from the baseline and 0.001 is taken as the scale.

[0025] The absolute value of the difference between the current value of the feature quantity in each time slice and the baseline within the short window is divided by the scale to obtain the deviation ratio sequence. The median of this sequence is taken as the short window representative deviation ratio of the feature quantity. A representative deviation ratio greater than or equal to 2 is judged as abnormal; otherwise, it is normal. The device simultaneously searches the set of anchor points within the short window. When there is a blink start anchor point within the short window and the feature quantity associated with blinking shows abnormality in no less than 3 time slices, while the feature quantity associated with plosive sound shows abnormality in only 1 time slice, the abnormality of that single time slice is regarded as transient interference and treated as normal. When there is a plosive sound start anchor point within the short window... If the anchor point and the feature associated with the plosive sound are abnormal in at least 3 time slices, while the feature associated with blinking is abnormal in only 1 time slice, a symmetrical reclassification is performed. After the reclassification is completed, the device counts the number of abnormalities in all feature quantities in the target time slice and the total number of feature quantities. The pain score is set to 10 multiplied by the number of abnormalities divided by the total number of feature quantities. At the same time, the uncertainty is set to 1 minus the ratio of the number of feature quantities that are consistent with the judgment of the target time slice and appear at least 3 times consecutively within the short window to the total number of feature quantities. The pain score and uncertainty are written into the score timeline according to the timestamp.

[0026] When the uncertainty exceeds the threshold, the device triggers a windowed review for the target time slice. The windowed review repeats the calculation and reassessment steps of the short window score over a longer time slice sequence, but only updates the pain score and uncertainty of the target time slice. It does not change the short window's writing to other time slices. The windowed review also reads the quality marker and missing completion marker and checks the stability of the representative deviation ratio. If the uncertainty is still higher than the threshold after the windowed review, only the pending status is recorded for the time slice to avoid false alarms. As the score timeline continues to be written, the device generates candidate segments on the score timeline using a dual-threshold rule. The entry threshold is used to determine the start of a candidate segment when the pain score rises, and the exit threshold is used to determine the end of a candidate segment when the pain score falls back. The entry threshold is higher than the exit threshold, and the minimum duration of the candidate segment is not less than two time slices.

[0027] For each candidate segment, the device searches for anchor points in the anchor point neighborhood near its starting point. The anchor point neighborhood is defined as a fixed symmetrical range in time slices, and the retrieved anchor points are sorted in ascending order by time slice index. If only the blink start anchor point is found in the anchor point neighborhood and the short-time energy and spectral flatness of the speech-side feature vector in the candidate segment do not show a common rising pattern consistent with the plosive sound, then the candidate segment is marked as a blink interference candidate. If only the plosive sound start anchor point is found in the anchor point neighborhood and the video-side feature vector shows a common rising pattern consistent with the plosive sound in the candidate segment, then the candidate segment is marked as a blink interference candidate. If the degree of fit and the optical flow amplitude in the facial region do not show a common change pattern consistent with the painful expression, the candidate segment is marked as a plosive interference candidate; if both the blink start anchor and the plosive start anchor are found in the anchor point neighborhood, the candidate segment is marked as a mixed interference candidate; if no anchor point is found in the anchor point neighborhood, the candidate segment is marked as an anchorless consistency candidate; the device then receives the set of candidate segments generated and marked according to the dual threshold rule, merges adjacent candidate segments with an interval of no more than 1 time slice, and performs fast verification.

[0028] Segments marked as blink interference candidates or plosive interference candidates are directly eliminated. Segments marked as mixed interference candidates are retained only if both video and speech abnormalities are present simultaneously within the segment. Segments marked as anchorless consistency candidates are retained. Boundary determination is performed on the retained candidate segments: based on the score timeline, the time slice where the entry threshold is first reached is taken as the confirmation start point, and the last time slice before two consecutive time slices below the exit threshold are taken as the confirmation end point. If the duration between the confirmation start and end is less than two time slices, the segment is canceled. The system outputs the real-time pain score and real-time confidence level for each time slice, with the real-time confidence level being 1 minus the uncertainty. For each confirmed pain segment, the system outputs the confirmation start and end times, and uses the arithmetic mean of the segment's built-in confidence level as the segment confidence level. When the interval between two confirmed pain segments does not exceed one time slice, they are merged and the start and end times and segment confidence are updated. In the engineering implementation of the device, deep learning is used to ensure the executability and stability of the above rules in complex environments: before writing the score timeline, the device uses a deep learning quality evaluator to assign quality labels to the video-side feature vectors and speech-side feature vectors corresponding to the time slices. The quality labels affect the windowing review triggering and pending state maintenance strategy but do not change the judgment criteria for short window scoring and reversal. The device uses a lightweight deep learning anomaly prior encoder to perform contextual modeling on the historical score timeline in the background, providing contextual hints for segments with continuously increasing uncertainty to prioritize triggering windowing review and anchor point neighborhood encryption retrieval, thereby improving the ability to separate weak pain sensation and short-term interference superposition scenarios without modifying the rules.

[0029] Furthermore, continuous acquisition of facial video at a fixed frame rate of 30 frames per second means that the interval between adjacent frames is approximately 33.3 milliseconds. With a time slice length of 200 milliseconds, each time slice can stably cover 6 frames, which can be used to calculate quantities that require cross-frame difference and robust median aggregation, such as eyelid opening and closing, eyebrow-eye distance, and facial region optical flow amplitude. Continuous acquisition of near-field speech at a fixed sampling rate of 16000 Hz means that the interval between individual samples is approximately 0.0625 milliseconds. 3200 samples can be covered in the same time slice, which is sufficient to accommodate multiple short-time speech analysis frames and support stable estimation of segment mean and segment change median for metrics such as short-time energy, spectral centroid, bandpass energy ratio, and spectral flatness. Each frame of facial video and each frame of near-field speech is timestamped in milliseconds, and the timestamps come from the same time base. This allows the jitter and buffer queuing delay at the acquisition end to be converted into a definite time stamp, thus achieving deterministic attribution by timestamp when writing a unified time slice sequence. Once the time axis is divided into equal-length slices and the start of the time slice advances in multiples of 200 milliseconds from the start of acquisition, continuous time is quantized into a discrete index with fixed granularity. The time slice index, start timestamp, and end timestamp together constitute a traceable temporal coordinate, maximizing the probability that the same physical event falls on the same index in different modalities during any run.

[0030] Choosing 200 milliseconds as the granularity is a trade-off between real-time performance and statistical stability: this length is sufficient to encompass 6 frames of video, providing a robust median trend even with slight occlusion or motion blur, while avoiding an interaction delay exceeding one time slice; for speech, the 200-millisecond window covers both stable vowel segments and the instantaneous energy and spectral morphology changes of consonant plosives, facilitating subsequent verification of the plosive's origin in the anchor neighborhood. Structurally, the unified time-slice sequence binds the video-side feature vector and speech-side feature vector, time-slice index, start timestamp, and end timestamp to the same record. Any clock drift from the device, once corrected to the unified time base, will not cause cross-modal misalignment accumulation within this structure. In practice, if a single frame is lost in the video or a very short silence occurs in the speech, the timestamp still ensures that the samples before and after it are assigned to the correct time slice. Combined with subsequent soft alignment and linear interpolation, continuity can be restored without disrupting the index monotonicity.

[0031] Furthermore, face localization employs an embedded deep learning detector to output face boundaries and initial values ​​of several key points in each frame of facial video. Key point tracking is accomplished collaboratively by a lightweight key point regressor and dense optical flow. In the current frame, the regressor corrects the initial key point values, and in the preceding and following frames, optical flow propagates the key point positions in the local neighborhood and performs consistency checks. Relocalization is triggered when regression confidence decreases or occlusion causes drift. The interocular center distance is calculated from the central key points of both eyes in each frame and used as a scale benchmark to ensure that distance and displacement remain comparable when the subject moves forward or backward or changes in perspective.

[0032] On the video side, eyelid opening and closing is defined as the normalized result of the instantaneous gap between key points of the upper and lower eyelid margins along the line connecting the pupils on the same side, relative to the distance between the centers of the two eyes, reflecting the instantaneous opening and closing state; eyebrow-eye distance is defined as the vertical distance between the highest point of the eyebrow and the center of the corresponding eye, which, after normalization relative to the distance between the centers of the two eyes, represents the lifting or lowering trend of the forehead-eye area; lip gap is defined as the minimum vertical distance between the inner edges of the upper and lower lips, which, after normalization, reflects the degree of mouth opening and closing; the upward angle of the corners of the mouth is defined as the instantaneous tilt angle of the line connecting the left and right corners of the mouth relative to the horizontal reference, with upward tilt being positive and downward tilt being negative. The α is used to characterize facial expression-driven mouth corner movements; the α is defined as the offset of the lateral distance between the outermost key points of the left and right α outer edges relative to their stable level, and after normalization, it reflects the respiratory-related α abduction and contraction; the α is defined as the normalized result of the inter-frame two-dimensional displacement amplitude with the nose tip or the center of the face bounding box as a reference relative to the distance between the centers of the two eyes, and is used to distinguish between overall head movements and local facial expression movements; the α is defined as a robust statistical measure of the magnitude of the dense optical flow vector within the face region, and the median or quantile statistics are often used to suppress local anomalous vectors.

[0033] Each of the above quantities first calculates the segment mean for the included frame-level values ​​within the time slice to give the stability level of that time slice, and then calculates the median of the segment change rate on the change sequence between adjacent frames to give the typical change trend, thus forming the video-side feature vector. The purpose of using deep learning in face localization and keypoint tracking is not to directly output pain scores, but to provide initial keypoint values ​​and quality scores that are robust to occlusion, lighting fluctuations, and rapid pose changes, ensuring that the normalization of the binocular center distance as a scale benchmark remains effective in different scenarios. On the speech side, near-field speech is framed with a fixed window length and step size. A metric for constructing the speech side feature vector is obtained for each frame. Short-time energy is defined as the average of the squared waveform amplitudes within the frame, characterizing sound intensity. Zero-crossing rate is defined as the normalized count of waveform symbol changes within the frame, reflecting the proportion of high-frequency components and noise roughness. Spectral centroid is defined as the frequency center position weighted by spectral amplitude; a larger value indicates a higher proportion of high-frequency energy. Bandpass energy ratio is defined as the proportion of energy within a pre-defined speech-related frequency band to the total energy, used to highlight frequency band activities related to vocalization. Spectral flatness is defined as the flatness of the spectrum at each frequency point; a higher value indicates a closer approximation to noise patterns, while a lower value indicates a closer approximation to harmonic structures. Fundamental frequency presence index is defined as the probability of sound or fundamental frequency reliability obtained based on autocorrelation or harmonic tracking methods, used to determine the presence of stable glottal excitation. Short-time jitter index is defined as a robust statistic of the fundamental frequency period length fluctuation between adjacent cycles, used to describe vocalization stability and tension.

[0034] Consistent with the video side, these frame-level metrics calculate segment means within the same time slice to obtain the stable acoustic level of that time slice, while simultaneously calculating the median segment rate of change to reflect the typical speed and direction of inter-frame changes, thus forming a speech-side feature vector at the time slice granularity. The combination of segment means and median segment rate of change is used because, on the one hand, the segment means can smooth out random noise and single-frame anomalies within a time slice containing multiple frames; on the other hand, the median segment rate of change can characterize continuous trends without being affected by extreme values. When both are incorporated into the video-side and speech-side feature vectors, they reflect both intensity levels and dynamic characteristics, facilitating subsequent short-window scoring to identify common pain-related anomalies within the local context of the five time slices.

[0035] Furthermore, the physiological process of blinking is characterized by the rapid closure of the eyelid opening within a very short period of time, followed by a swift return to the baseline. Under a frame rate of 30 frames per second, the duration corresponding to no more than 3 frames is approximately on the order of 100 milliseconds, which can cover the onset phase of a typical blink. The requirement that the eyelid opening decreases by 0.2 relative to the preset baseline within no more than 3 frames and then recovers by 0.2 within the following 3 frames essentially constrains a "symmetrical rapid decrease and rapid increase" geometric shape, excluding non-instantaneous events such as slow squinting or continuous eye closure. At the same time, the requirement that the median rate of change of the eyebrow-eye distance in this time slice is positive is to use the accompanying action of slightly raising the forehead-eye area to distinguish misjudgments caused by looking down, frowning, or occlusion, making the blinking starting point anchor point more specific. The acoustic mechanism of plosive sounds is the instantaneous release of broadband high-energy short pulses after the glottis is fully closed. This is characterized by a significant jump in short-time energy, an increase in the proportion of high-frequency energy, and a short-time transition of the spectrum from harmonic dominance to an approximate noise pattern. Therefore, it is stipulated that the average value of the short-time energy segment increases by a factor of 2 or more relative to the preset baseline, the average value of the bandpass energy ratio segment increases by 0.3 or more, and the average value of the spectral flatness segment increases by 0.2 or more. If all three conditions are met simultaneously, this broadband burst event can be stably characterized at the time slice granularity and distinguished from continuous vowels, groans, or environmental steady-state noise.

[0036] Furthermore, the short-window scoring unfolds around the target time slice within a window of 5 time slices in length. This setting ensures that the scoring covers both the immediate changes before and after the target slice without overly smoothing out real fluctuations. Within this window, for each feature quantity listed in the fixed feature list, the historical median is first calculated in the historical interval 30 seconds before the target time slice as a baseline to accommodate individual differences and device gain variations. Then, the median absolute deviation of this historical interval from the baseline is used as the primary scale, and the larger of this median and a minimum lower limit is taken as the final scale. The introduction of the lower limit is to avoid false amplification caused by an excessively small scale in long-term stable or almost fluctuation-free scenarios.

[0037] Subsequently, the deviation of the current value of this feature quantity from the baseline in each time slice within the short window is made dimensionless according to the scale to obtain the deviation ratio sequence, and the median is used as the short window representative deviation ratio of this feature quantity. The median is selected to maintain robustness in the presence of single-frame blur, momentary silence or brief mechanical vibration, and not to be pulled by extreme values. When the representative deviation ratio reaches or exceeds 2, it indicates that the feature quantity has deviated significantly from its own recent fluctuation boundary, and is marked as abnormal; otherwise, it is normal. Because some features on the video and audio sides can respond synchronously to physiological blinking or vocal plosives, the short-window scoring immediately performs consistency verification with the anchor set after generating the initial judgment: if there is a blinking origin anchor within the short window, and the features associated with blinking are abnormal in at least 3 time slices, while the features associated with plosive sounds are abnormal in only 1 time slice, then the abnormality in that single time slice is considered transient interference and treated as normal, in order to avoid mistaking the rapid peak of eyelid opening and closing and the amplitude of optical flow in the facial area caused by blinking as pain; if there is a plosive sound origin anchor within the short window, and the features associated with plosive sounds are abnormal in at least 3 time slices, while the features associated with blinking are abnormal in only 1 time slice, then a symmetrical re-judgment is performed, in order to avoid mistaking the synchronous increase of short-term energy, bandpass energy ratio and spectral flatness caused by consonant plosives as pain-related vocalization.

[0038] After reclassification, each target time slice is linearly mapped to a number line from 0 to 10 using the ratio of the number of anomalies to the total number of features to obtain a pain score. This mapping directly reflects the "common anomaly coverage," unifying the absolute dimensional differences among different subjects, devices, and scenarios into a comparable pain score expression. Since the anomaly judgment of each feature is based on the dimensionless bias formed by the same historical baseline and scale, the score does not depend on the absolute magnitude of a fixed threshold but on the degree of relative deviation, thus exhibiting stability against changes in illumination, microphone gain drift, and individual facial expression intensity differences. The key to the above process lies in the simultaneous effect of three constraints: a unified time slice sequence ensures consistent sampling and computational boundaries across modalities in the time dimension; the baseline and scale formed by the historical median and the median of absolute deviation ensure an adaptive reference frame for the same subject in the short term; and the anchor set ensures structured recognition and reclassification of high-frequency instantaneous events such as physiological blinking and vocal plosives. In engineering, this allows short-window scoring to complete local judgments within a 1-second window and output a pain score corresponding to a timestamp in each time slice, meeting the latency constraints of real-time applications. Statistically... The deviation ratio uses the median instead of the mean to suppress the influence of extreme samples. The anomaly threshold uses a fixed dimensionless standard of 2 to ensure that the contributions of different feature quantities are within the same judgment framework. The re-judgment condition requires consistent anomalies across no less than 3 time slices to exclude occasional spikes, thus making the score more sensitive to continuous, multi-feature common deviations and less sensitive to isolated, single-feature short spikes. In terms of multimodal consistency, anomalies on the video side and the audio side need to independently meet the judgment in the same short window and be cross-validated through the anchor point set to avoid false high scores triggered by single-modal noise.

[0039] Furthermore, uncertainty is measured by both "cross-feature consistency" and "temporal continuity" to assess the reliability of short-window scoring results. Instead of directly comparing magnitudes, the system counts features that are consistent with the target time slice's judgment and appear consecutively at least three times within the short window. The higher the proportion of features that are consistent with the target time slice's judgment and appear consecutively at least three times within the short window, the more features and adjacent time slices support the judgment of abnormality or normality. This means that random noise or single spikes have less impact on the result, resulting in lower uncertainty. Conversely, when only a small number of features maintain the same judgment as the target time slice within the short window and appear consecutively at least three times, it indicates that the judgment lacks stable cross-feature and cross-temporal basis and is more likely caused by momentary occlusion, short-term silence, or equipment micro-vibration, resulting in higher uncertainty.

[0040] The threshold of "at least three consecutive occurrences" is used to distinguish sporadic, single-piece anomalies from temporally stable, fragmented patterns. The requirement of three occurrences ensures that the judgment spans at least several adjacent time slices before and after the target time slice, guaranteeing that the consistency has the most basic temporal continuity. The constraint of "consistency with the judgment of the target time slice" focuses the measurement on the supporting evidence of the current output result, avoiding the introduction of contradictory evidence unrelated to the current judgment. Based on a proportion, the uncertainty is obtained by subtracting the proportion from 1, which can normalize the result to the range of 0 to 1. This intuitively corresponds to the interpretation paradigm of "the closer to 0, the more credible; the closer to 1, the less credible," facilitating its alignment with the definition of real-time confidence and its use in conjunction with the entry and exit thresholds of the subsequent dual-threshold rule. In engineering implementation, writing pain scores and uncertainties into the score timeline in chronological order according to timestamps ensures that the temporal correspondence between the output of any time slice and its input data in a unified time slice sequence is completely preserved. This supports windowed review and anchor neighborhood retrieval for review and consistency verification in subsequent steps, and also facilitates rapid location and tracing using time as the key when environmental changes occur over long periods. At the same time, the monotonic time organization of the score timeline ensures that the real-time module can read, update, and confirm sequentially within a fixed delay budget, avoiding the segment boundary drift problem caused by out-of-order writing. Uncertainty can be used as a real-time scheduling signal to directly drive subsequent pending state maintenance or review triggering strategies, ultimately forming a continuously plottable and verifiable time series output together with the pain scores.

[0041] Furthermore, the dual-threshold rule and anchor-point neighborhood retrieval labeling organize continuous numerical changes on the score timeline into candidate segments with clear beginnings and ends. Explainable events from the anchor-point set are used to causally screen the neighborhood of the segment's starting point, thereby rigorously distinguishing sudden increases caused by instantaneous physiological movements or speech sounds from truly pain-related multimodal anomalies. The score timeline is a sequence of pain scores written in timestamp order; its fluctuations include both slow rises or falls driven by both video-side and speech-side feature vectors, as well as spikes caused by short-term peaks in a single modality.

[0042] The dual-threshold rule of entry and exit thresholds essentially introduces a hysteresis-based opening and closing criterion onto a continuous curve: when the pain score rises from a low level and exceeds the entry threshold for the first time, the time slice is marked as the candidate start; in subsequent time slices, as long as the score does not fall below the exit threshold, the candidate state remains open, thus avoiding frequent starts and stops caused by minor fluctuations in the score near the threshold; when the score first falls below the exit threshold, the time slice is marked as the candidate end. The design of the entry threshold being higher than the exit threshold separates the "entry" and "exit" channels, ensuring that the same fluctuation is not fragmented into multiple pieces due to a single drop and subsequent rise; at the same time, the minimum duration of the candidate segment is set to be no less than two time slices, which can automatically filter out occasional spikes of less than 400 milliseconds under a time resolution of 200 milliseconds, because such spikes are more likely to be caused by single-frame motion blur, momentary silence, or device micro-jitter, rather than physiologically significant pain responses.

[0043] Through the above mechanism, the dual-threshold rule transforms the numerical up-and-down traversal into structured candidate segments, providing clear objects for subsequent anchor neighborhood consistency verification. For each candidate segment, the system searches for anchor points in the anchor neighborhood near its starting point. The anchor neighborhood is defined as a fixed symmetrical range in units of time slices. The principle behind this is that the triggering of the candidate segment's start often occurs within a short period after a precipitating event. Therefore, searching around the starting point can maximize the capture of blink start anchor points or plosive start anchor points related to the cause of the segment. The retrieved anchor points are arranged in ascending order of time slice index, ensuring a deterministic processing order when multiple anchor points exist, and keeping the verification logic consistent during repeated runs and cross-device reproduction. The labeling strategy uses "whether there is a unique dominant anchor type" and "whether another modality presents a common pattern consistent with the anchor" as the core criteria: if only the blinking start anchor is found in the neighborhood of the anchor, and the short-time energy and spectral flatness of the speech-side feature vector in the candidate segment do not show a common rising pattern consistent with the plosive, it means that the score increase of the segment is likely caused by the rapid change in eyelid opening and closing degree and facial optical flow amplitude caused by blinking, rather than being driven by speech plosives or pain-related vocalizations, and is therefore labeled as a blinking interference candidate.

[0044] If only the plosive sound origination anchor point is found in the anchor point neighborhood, and the eyelid opening / closing degree and facial optical flow amplitude of the video-side feature vector within the candidate segment do not show a common change pattern consistent with painful expressions, then the score increase of this segment is mainly driven by the synchronous transition of short-term energy and spectral morphology caused by the plosive sound, while the video side does not provide evidence of accompanying facial pain expressions; therefore, it is marked as a plosive sound interference candidate. If both the blink origination anchor point and the plosive sound origination anchor point are found in the anchor point neighborhood, it is difficult to determine a single main cause, indicating that the segment opening / closing degree and facial optical flow amplitude of the candidate segment do not show a common change pattern consistent with painful expressions. Initially, there is a superposition or successive occurrence of visual and acoustic transient events. In the subsequent rapid verification, the retention condition is "whether there are common anomalies on both the video and speech sides at the same time". Therefore, it is marked as a mixed interference candidate at this stage. If no anchor point is found in the neighborhood of the anchor point, it means that the formation of the candidate segment is not triggered by the two predefined transient events of blinking or plosive sounds. If the segment is supported by multiple feature abnormalities in the short window score, it is more consistent with the slow change pattern driven by pain-related expressions and vocalization. Therefore, it is marked as an anchorless consistency candidate. The key to the above classification principle lies in cross-validating the "existence and type of event anchors" with the "cross-modal common patterns": the blinking origin anchor should be consistent with the short, abrupt changes in eyelid opening and closing and facial optical flow amplitude on the video side, and should not be accompanied by a joint increase in short-term energy and spectral flatness on the speech side; the plosive origin anchor should be consistent with the synchronous transition of short-term energy and spectral flatness on the speech side, and should not be accompanied by pain-related joint changes in eyelid opening and closing and facial optical flow amplitude on the video side; when both types of anchors appear simultaneously, only when both the video side and the speech side show their respective common abnormalities can it be retained as a complex scene related to pain, otherwise it should be removed in subsequent rapid verification.

[0045] Furthermore, the system receives a set of candidate segments generated and labeled according to a dual-threshold rule; merges adjacent candidate segments with an interval of no more than one time slice; performs rapid verification: segments labeled as blink interference candidates or plosive interference candidates are directly eliminated; segments labeled as mixed interference candidates are retained only when both video and speech anomalies occur simultaneously within the segment; segments labeled as anchorless consistency candidates are retained; boundary determination is performed on the retained candidate segments: based on the score timeline, the time slice that first reaches the entry threshold is taken as the confirmation start point, and the last time slice before the first two consecutive time slices below the exit threshold are taken as the confirmation end point; if the duration between the confirmation start and end is less than two time slices, the segment is canceled; the system outputs the real-time pain score and real-time confidence level for each time slice, with the real-time confidence level being 1 minus the uncertainty; for each confirmed pain segment, the system outputs the confirmation start and end times, and uses the arithmetic mean of the segment's built-in confidence level as the segment confidence level; when the interval between two confirmed pain segments is no more than one time slice, they are merged and the start and end times and segment confidence levels are updated.

[0046] Furthermore, after transforming numerical decision results into stable event-level outputs on the fractional timeline, structured verification and boundary determination ensure that each output possesses temporal and semantic consistency. Upon receiving the set of candidate segments generated and labeled according to the dual-threshold rule, adjacent candidate segments with an interval of no more than one time slice are first merged. This is based on the fact that the fractional timeline exhibits brief dips and subsequent rises near the threshold. These gaps typically do not represent a genuine relief of the physiological state but rather fragmentation caused by noise, transient interference, or feature drift. Allowing merging within a range of no more than one time slice maintains the temporal continuity of segments and reduces missegmentation in subsequent confirmations. The principle of rapid verification is to prioritize the elimination of non-painful elevations whose causes are clearly pointed to by anchor points. Therefore, segments marked as blink interference candidates or plosive interference candidates are directly eliminated. Segments marked as mixed interference candidates are retained only when both video and speech abnormalities are present in the segment, to avoid single modal peaks being mistakenly identified as valid events in the case of anchor point superposition. Segments marked as anchorless consistency candidates are retained because their origin neighborhood lacks triggering evidence of blink origin anchor points and plosive origin anchor points, which is more consistent with the pain-related multimodal consistent elevation pattern.

[0047] When performing boundary determination on the retained candidate segments, the time slice that first reaches the entry threshold is selected as the confirmation start point based on the fractional timeline, and the last time slice before the first occurrence of two consecutive time slices below the exit threshold is selected as the confirmation end point. The segment is canceled when the duration between the confirmation start and end is less than two time slices. The principle is to use the hysteresis characteristics of the entry and exit thresholds and the minimum duration constraint to jointly suppress short spikes and boundary jitter, so that the output event has a minimum perceptible duration and matches the continuous fluctuation. The system outputs the real-time pain score and real-time confidence level for each time slice. The real-time confidence level is 1 minus the uncertainty, providing a directly usable estimate and reliability comparison for subsequent modules at the time slice granularity. For each confirmed pain segment, the system outputs the confirmation start and end times, and uses the arithmetic mean of the segment's built-in confidence level as the segment's confidence level. The arithmetic mean is chosen because a uniform time slice sequence is evenly spaced in time, and the average reflects the overall reliability of the entire segment and maintains a stable response to short-term fluctuations. When the interval between two confirmed pain segments does not exceed one time slice, they are merged, and the start and end times and segment confidence levels are updated. This aims to prevent a single continuous physiological event from being split into multiple independent events due to slight declines. The confidence level of the merged segment is recalculated according to the same rules to maintain consistency with the score timeline statistics. Through this sequential mechanism of merging, rapid verification, and boundary determination, the system compresses the candidate segment set into a final output with clearly defined confirmation start and end times, real-time pain scores, real-time confidence levels, and segment confidence levels, improving the stability and verifiability of event-level output while ensuring real-time performance.

[0048] Furthermore, the entire process from data acquisition to the output of the two-stage consistent inference device is fully realized according to the constraints of a unified time-slice sequence and anchor point set, and specific values ​​and calculations are given at each step. Assuming a time-slice length of 200 milliseconds, facial video is continuously acquired at 30 frames per second, and near-field speech is continuously acquired at 16000 Hz. A unified time reference is used to timestamp both channels, and the time slice advances in multiples of 200 milliseconds starting from t=120.0 seconds. The target time-slice index is denoted as... (This example takes) The corresponding time interval is [t=120.0s, t=120.2s), and the short window is used. (i.e., [t=119.6s, t=120.6s)). For ease of writing, let a certain feature quantity in the fixed feature list be denoted as... When a symbol first appears, its meaning should be explained uniformly: Indicates time slice upper features The current value (i.e., the "segment mean" within this time slice), Indicates the target time slice Features within the previous 30-second historical interval The historical median (baseline). Indicates the same historical interval The larger of the median absolute deviation and 0.001 is taken (scale). The deviation ratio of each time slice within the short window is defined as follows: ; Representation of features In the target time slice The short window represents the deviation ratio, defined as .

[0049] when The characteristic is recorded as "abnormal" in the target time slice, otherwise as "normal". The pain score in the target time slice is recorded as follows: Defined as ,in The number of features that are judged to be abnormal in the target time slice. The total number of features (7 from the video side and 7 from the speech side in this example, totaling...) Uncertainty is denoted as . Defined as ,in In short window The number of features that are consistent with the target time slice and appear consecutively at least 3 times. Real-time confidence is denoted as... .

[0050] Create a unified time-slice sequence and baseline / scale. The following historical baseline and scale are provided for both the video (normalized by binocular center distance) and audio sides: eyelid opening / closing. Distance between eyebrows and eyes ; Lip gap The angle of the corners of the mouth turning upwards (unified) ; Nasal wing expansion range Head translation range Facial region optical flow amplitude Voice side: Short-term energy Zero crossing rate ; Spectral centroid (normalized) Bandpass energy ratio Spectral flatness Fundamental frequency presence index Short-term jitter index Current value of each time slice within the short window. The following (units have been normalized or indexed according to their respective definitions): Video side in The values ​​are: eyelid opening and closing degree Distance between eyebrows and eyes lip gap The angle of the upturned corners of the mouth nasal wing expansion range Head translation range Facial region optical flow amplitude On the voice side The values ​​for are: short-time energy Zero crossing rate Spectral center Bandpass energy ratio Spectral flatness Fundamental frequency presence index Short-term jitter index Among them, in Within the video frame sequence, a blink initiation anchor point was detected (the eyelid opening / closing degree decreased by ≥0.2 relative to the baseline within no more than 3 frames, then recovered by ≥0.2 within no more than 3 frames, and the median rate of change of eyebrow-eye distance segment was positive at that time). On the speech side... The plosive start point anchor point was not triggered (although there was an energy spike, the bandpass energy ratio and spectral flatness did not simultaneously meet the threshold rise).

[0051] Short window scoring calculation and anchor point consistency reassessment: Taking eyelid opening and closing as an example, the calculation... ,have to Therefore This was identified as an anomaly. The same method was used to obtain the features at the target time slice. The short window represents the deviation ratio: eyebrow-eye distance (Abnormality); Lip gap (Normal); Angle of upturn at the corners of the mouth (Abnormal); Nasal wing expansion range (Abnormal); Head translation range (Abnormal); Facial region optical flow amplitude (Abnormal). Voice side: Short-time energy (Normal); Zero crossing rate (Normal); Spectral centroid (Normal); Bandpass energy ratio (Normal); Spectral flatness (Normal); Fundamental frequency presence index (Abnormal); Short-term jitter index (Abnormality). In the consistency reclassification, a blink initiation anchor point exists within the short window, and blink-related features (eyelid opening / closing, eyebrow-eye distance, facial region optical flow amplitude) show abnormalities in at least 3 time slices, while features related to plosive sounds (short-time energy, bandpass energy ratio, spectral flatness) only show abnormalities in 3 time slices. A single-piece anomaly was detected, therefore it was treated as normal. Since the median was used to represent the deviation ratio, the above three items were already normal in the target time slice, and the judgment remained unchanged. Therefore, the anomaly count for the target time slice was 6 items on the video side (all abnormal except for the lip gap) and 2 items on the speech side (fundamental frequency presence index and short-term jitter index), totaling... ,get .

[0052] Uncertainty and real-time confidence calculation: For each feature, the confidence level is calculated within a short window. The sequences were categorized as abnormal / normal based on a threshold of 2, and the number of features that "consistent with the target time slice and appearing consecutively at least 3 times" was counted. Eyelid opening and closing degree was... to Continuous abnormalities, the angle of the corners of the mouth turning upwards is to Continuous abnormalities, nasal alar expansion range to Continuous abnormalities, head translation range to Continuous abnormalities, with optical flow amplitude in the facial area at to Continuous abnormalities, distance between eyebrows and eyes to Continuous abnormalities; lip gap in to Continuous normal; speech-side short-time energy, bandpass energy ratio, and spectral flatness are within acceptable limits. All external components are normal and contain continuous normal segments of length ≥3; the fundamental frequency presence index is... to Continuous anomalies, short-term jitter indicators to Continuous anomalies, zero-crossing rate and spectral centroid are normal or covered throughout the window. A continuous normal segment with a length ≥ 3. Therefore ,

[0053] .

[0054] Candidate fragment generation and confirmation: Setting entry threshold Exit threshold ( The minimum duration is Seconds. Press and The same method was used to calculate the results for adjacent time slices as target slices. , , , , .when When the entry threshold is first reached, the candidate status is recorded as starting; thereafter until... Previously, the threshold was not lower than the exit threshold, and the duration covered... to There are 3 time slices in total, satisfying the minimum duration. The candidate is considered within the anchor neighborhood of its starting point (in this example, we take...). A time-slice search revealed only a blink initiation anchor point, and the short-term energy and spectral flatness on the speech side did not show a common rising pattern consistent with plosives. This candidate was marked as a blink interference candidate. Entering the rapid verification stage, since blink interference candidates were directly eliminated, this initiation candidate was not confirmed. Continuing to observe the score timeline, due to... and If the candidate is still above the exit threshold but has no anchor in the anchor neighborhood or is inconsistent with the pattern, the system generates a new candidate (anchor-free consistent candidate) in the next window. When the boundary is determined, the time slice in which the entry threshold is first reached is used. As the confirmation starting point, the last time slice before the first occurrence of two consecutive time slices falling below the exit threshold is taken as the confirmation ending point (assuming that in and If the threshold is lower than the exit threshold, then the endpoint is confirmed. ), confirm fragment coverage That is, [t=119.8s, t=120.6s). The real-time confidence level within this segment is... , , Fragment confidence is defined as the arithmetic mean of the built-in confidence of a fragment, denoted as . .

[0055] If the interval between two confirmed pain segments does not exceed one time slice, they are merged and the start and end times and segment confidence are updated according to the same rule; in this example, only one confirmed segment is generated, so merging is not necessary. At this point, the system outputs real-time pain scores and real-time confidence for each time slice, and outputs confirmed start and end times and segment confidence for confirmed segments, completing a full implementation process from a unified time slice sequence and anchor point set to the output of a two-segment consistent inference device.

[0056] Figure 2 A schematic diagram of the experimental curve for blink initiation anchor point detection is shown, detailing the technical principle of blink initiation anchor point identification based on changes in eyelid opening and closing in this invention. The graph uses time as the horizontal axis, with units in milliseconds, ranging from 0 milliseconds to 1400 milliseconds, fully covering the complete time cycle of a typical blinking action. The vertical axis represents eyelid opening and closing, expressed proportionally, with values ​​ranging from 0.0 to 1.0, where 1.0 represents a fully open eyelid and 0.0 represents a fully closed eyelid. The horizontal dashed line in the graph marks a preset baseline located at an eyelid opening and closing degree of 0.8, serving as a reference standard for judging abnormal changes in eyelid opening and closing. According to the technical solution of claim 4, when the eyelid opening and closing degree changes significantly relative to the preset baseline, the system will initiate the blink initiation anchor point detection procedure. The experimental curve clearly shows two typical blinking processes. During the first blink, at approximately 280 milliseconds, the eyelid opening angle dropped sharply from the preset baseline to about 0.2, a decrease of 0.6, far exceeding the 0.2 threshold requirement specified in the claims. Subsequently, at around 300 milliseconds, the eyelid opening angle rapidly recovered to near the preset baseline, with the recovery also exceeding the 0.2 criterion. The entire decrease and recovery process was completed within a duration of no more than 3 frames, fully meeting the detection conditions for the blink initiation anchor point. The second blink occurred between 550 and 570 milliseconds, exhibiting a similar change pattern to the first blink: the eyelid opening angle first rapidly decreased to below 0.2, and then recovered to a normal level within a short period. After detecting that the changes in eyelid opening angle at these two time points met all the judgment conditions, the system generated blink initiation anchor points in the corresponding time slices, indicated by the "blink" symbol marked with a box in the figure. These anchor points will serve as important time references for subsequent multimodal data soft alignment processing, ensuring accurate synchronization of video-side feature vectors and speech-side feature vectors in the temporal dimension.

[0057] Figure 3The experimental curve diagram for plosive start point anchor point detection is presented, systematically demonstrating the core technical mechanism of this invention based on short-term energy changes to identify plosive start point anchor points. The graph uses a standard two-dimensional coordinate system. The horizontal axis represents time in milliseconds, with a measurement range from 0 to 1400 milliseconds, ensuring the complete occurrence process of plosives in the speech signal can be captured. The vertical axis represents short-term energy, marked as multiples, with values ​​ranging from 0 to 8 times, fully covering the energy change amplitude under normal speech and plosive states. The graph includes two important horizontal reference lines: a preset baseline and a 2x threshold line. The preset baseline is located at 1x short-term energy, representing the short-term energy level of near-field speech under normal conditions. This baseline is calculated from historical data within 30 seconds prior to the target time slice. The 2x threshold line is located at 2x short-term energy. According to the technical specification of claim 4, when the average value of the short-term energy segment increases by a multiple of 2 relative to the preset baseline, the primary condition for plosive start point anchor point detection is met. The experimental curve records two distinct plosive events. The first plosive sound occurred at time 290 milliseconds, with the short-term energy surging dramatically from 1x to approximately 6x the preset baseline, exceeding the detection threshold of 2x. The second plosive sound appeared around 550 milliseconds, with the short-term energy also exhibiting a significant surge, rising to approximately 5x the baseline. At these two time points, the system not only detected the large increase in short-term energy but also simultaneously verified whether the segment mean of the bandpass energy ratio increased by more than 0.3 relative to its preset baseline, and whether the segment mean of the spectral flatness increased by more than 0.2 relative to its preset baseline. When all detection conditions were met simultaneously, the system generated plosive sound initiation anchor points within the corresponding time slices, marked with a box labeled "plosive" in the figure. These anchor points, together with the blink initiation anchor points, constitute the anchor point set, providing crucial time synchronization information for the two-stage consistent inference device. This ensures accurate soft alignment of multimodal features within the anchor point neighborhood, thereby improving the accuracy and reliability of pain assessment.

[0058] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. A multimodal real-time intelligent pain assessment and assistance system based on deep learning, characterized in that, The system includes: a multimodal data acquisition device and a two-stage consistency inference device; wherein, the multimodal data acquisition device synchronously acquires facial video and near-field speech under a unified time reference, segments the time axis to form a unified time-slice sequence; generates video-side feature vectors and speech-side feature vectors respectively according to a fixed feature list; generates blink start anchors and plosive start anchors according to fixed judgment rules to form an anchor set; performs soft alignment in the neighborhood of anchor points based on the anchor set and performs linear interpolation to fill in missing time slices, so as to output the aligned unified time-slice sequence and anchor set; the two-stage consistency inference device, under the constraints of the aligned unified time-slice sequence and anchor set, first performs short-window scoring on each time slice to obtain a pain score and a non-pain score. The determination is evaluated, and when the uncertainty exceeds a threshold, a windowed verification is performed to update the pain score and uncertainty. Subsequently, a consistency check is performed based on the anchor point set to obtain candidate segments. The candidate segments are confirmed on the score timeline, and the real-time pain score, confidence level, and start and end times of the confirmed pain segments are output. The specific steps of the two-stage consistency inference device to perform short-window scoring to obtain the pain score and uncertainty include: for each target time slice, a short window of length 5 time slices is taken on the unified time slice sequence centered on that time slice; for each feature quantity listed in the fixed feature list, the historical median is calculated as the baseline within 30 seconds before the target time slice, and the median absolute deviation of the historical interval relative to the baseline is calculated and compared with 0.

001. Take the larger value as the scale; divide the absolute value of the difference between the current value of the feature quantity and the baseline in each time slice within the short window by the scale to obtain the deviation ratio sequence, and take the median of the sequence as the short window representative deviation ratio of the feature quantity. If the representative deviation ratio is greater than or equal to 2, it is considered abnormal; otherwise, it is considered normal. If there is a blink start anchor point within the short window and the feature quantity associated with blinking shows abnormality in no less than 3 time slices, while the feature quantity associated with plosive sounds shows abnormality in only 1 time slice, then the abnormality in that single time slice is considered transient interference and treated as normal. If there is a plosive sound start anchor point within the short window and the feature quantity associated with plosive sounds shows abnormality in no less than 1 time slice, then the abnormality in that single time slice is considered transient interference and treated as normal. If anomalies occur in three time slices, but the feature quantity associated with blinking only shows anomalies in one time slice, a symmetrical reclassification is performed. After reclassification, the number of anomalies in all feature quantities for the target time slice is counted against the total number of feature quantities. The pain score is set to 10 multiplied by the number of anomalies divided by the total number of feature quantities. The uncertainty is set to 1 minus the ratio of the number of feature quantities that are consistent with the target time slice judgment and occur consecutively at least three times within the short window to the total number of feature quantities. The resulting pain score and uncertainty are written into the score timeline in chronological order of timestamps. The two-stage consistency inference device performs consistency verification based on the anchor point set to obtain... The specific steps for generating candidate segments include: generating candidate segments on the score timeline using a dual-threshold rule, which includes an entry threshold and an exit threshold. The entry threshold is used to determine the start of a candidate segment when the pain score rises, and the exit threshold is used to determine the end of a candidate segment when the pain score falls back. The entry threshold must be higher than the exit threshold, and the minimum duration of a candidate segment must be no less than two time slices. For each candidate segment, anchor points are retrieved within the anchor point neighborhood near its starting point. The anchor point neighborhood is defined as a fixed symmetrical range in units of time slices, and the retrieved anchor points are sorted in ascending order by time slice index. If only the blink start anchor point is retrieved within the anchor point neighborhood and the speech-side features are... If the short-time energy and spectral flatness of the vector within a candidate segment do not exhibit a common rising pattern consistent with plosive sounds, the candidate segment is marked as a blink interference candidate. If only the plosive sound origination anchor is found in the anchor point neighborhood and the eyelid opening and closing degree and facial optical flow amplitude of the video-side feature vector within the candidate segment do not exhibit a common change pattern consistent with painful expressions, the candidate segment is marked as a plosive sound interference candidate. If both blinking origination anchors and plosive sound origination anchors are found in the anchor point neighborhood, the candidate segment is marked as a mixed interference candidate. If no anchor point is found in the anchor point neighborhood, the candidate segment is marked as an anchorless consistency candidate.

2. The deep learning-based multimodal real-time intelligent pain assessment and assistance system as described in claim 1, characterized in that, Facial video is continuously captured at a fixed frame rate of 30 frames per second, and near-field speech is continuously captured at a fixed sampling rate of 16,000 Hz. A timestamp in milliseconds is written for each frame of facial video and each frame of near-field speech, and the timestamps come from the same time base. The time axis is divided into equal-length slices, each 200 milliseconds long, with the start of each slice advancing in multiples of 200 milliseconds from the start of the capture. Within each slice, all facial video and near-field speech within its time range are collected, and a slice index, a start timestamp, and an end timestamp are added to the slice, thus forming a unified time slice sequence.

3. The deep learning-based multimodal real-time intelligent pain assessment and assistance system as described in claim 2, characterized in that, The process by which the multimodal data acquisition device generates video-side feature vectors and speech-side feature vectors based on a fixed feature list includes: performing face localization and key point tracking in each frame of facial video, using the distance between the centers of the two eyes as the scale reference; calculating the segment mean and median rate of change of the following quantities in each time slice: eyelid opening and closing, eyebrow-eye distance, lip gap, corner of mouth upward angle, nasal wing expansion amplitude, head translation amplitude, and facial region optical flow amplitude, to form the video-side feature vector; dividing the near-field speech into frames according to a fixed window length and step size; calculating the segment mean and median rate of change of the following quantities in each time slice: short-time energy, zero-crossing rate, spectral centroid, bandpass energy ratio, spectral flatness, fundamental frequency presence index, and short-time jitter index, to form the speech-side feature vector.

4. The deep learning-based multimodal real-time intelligent pain assessment and assistance system as described in claim 3, characterized in that, The process by which a multimodal data acquisition device generates blink initiation anchor points and plosive sound initiation anchor points according to fixed judgment rules and forms an anchor point set includes: On a uniform time-slice sequence, only two types of anchor points are generated and written into the anchor point set according to the following rules: Blink initiation: When, within a certain time slice, the eyelid opening and closing degree is detected to decrease by 0.2 relative to the preset baseline of the metric for a duration of no more than 3 frames, and subsequently increases by 0.2 relative to the preset baseline of the metric for a duration of no more than 3 frames, while the eyebrow-eye distance changes within the segment of that time slice... When the median of the conversion rate is positive, the time slice is marked as the blink start anchor point; plosive start point: when the average value of the short-time energy segment is detected to increase by a factor of 2 or more relative to the preset baseline of the metric within a certain time slice, and the average value of the bandpass energy ratio segment increases by an amount of 0.3 or more relative to its preset baseline, and the average value of the spectral flatness segment increases by an amount of 0.2 or more relative to its preset baseline, the time slice is marked as the plosive start anchor point; each record in the anchor point set contains the anchor point type, time slice index, start timestamp, and end timestamp.

5. The deep learning-based multimodal real-time intelligent pain assessment and assistance system as described in claim 1, characterized in that, The system receives a set of candidate segments generated and labeled according to a dual-threshold rule; merges adjacent candidate segments with an interval of no more than one time slice; performs rapid verification: segments labeled as blink interference candidates or plosive interference candidates are directly eliminated; segments labeled as mixed interference candidates are retained only if both video and speech anomalies occur simultaneously within the segment; segments labeled as anchorless consistency candidates are retained; performs boundary determination on the retained candidate segments: based on the score timeline, the time slice that first reaches the entry threshold is taken as the confirmation start point, and the last time slice before the first two consecutive time slices below the exit threshold are taken as the confirmation end point; if the duration between the confirmation start and end is less than two time slices, the segment is canceled; the system outputs the real-time pain score and real-time confidence level for each time slice, with the real-time confidence level being 1 minus the uncertainty; for each confirmed pain segment, the system outputs the confirmation start and end times, and uses the arithmetic mean of the segment's built-in confidence level as the segment confidence level; when the interval between two confirmed pain segments is no more than one time slice, they are merged and the start and end times and segment confidence levels are updated.