Infant state identification method and system based on Chinese medicine five-tone monitoring analysis
The infant state recognition method, which integrates multimodal feature fusion and dynamic weight adjustment, solves the problems of subjectivity in the five-tone-five-organ mapping and single-modal noise sensitivity in existing technologies. It achieves highly accurate and individual-adaptive infant emotional state recognition, and is applicable to scenarios such as mobile terminals and maternal and infant monitors.
Patent Information
- Application Number
- CN202511407372.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies for infant state recognition suffer from several problems, including the subjectivity and ambiguity of the mapping theory of the correspondence between the five tones and the five internal organs, the lack of dynamic adaptability in multimodal fusion, the noise sensitivity of single-modal recognition schemes, and the lack of integration with traditional Chinese medicine semantic knowledge in traditional systems. These issues result in insufficient recognition accuracy and adaptability.
A multimodal feature fusion method is adopted, which combines infant crying audio, physiological signals and behavioral video data. A mapping relationship is established through the five-tone theory of traditional Chinese medicine. Fuzzy sets and Bayesian inference are used for dynamic weight adjustment, and an interpretable comprehensive physiological and emotional state assessment is output.
It improves the accuracy and robustness of infant and toddler emotional state recognition, enhances the model's individual adaptability and clinical interpretability, meets real-time monitoring needs, and reduces system resource consumption.
Smart Images

Figure CN121370165A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal information processing and application of traditional Chinese medicine theory, and in particular to a baby state recognition method and system based on traditional Chinese medicine five-tone monitoring analysis. BACKGROUND
[0002] Currently, multi-modal physiological signal analysis and emotional state recognition of infants based on traditional Chinese medicine theory combined with modern artificial intelligence technology is a research frontier in the field of medical health monitoring. Existing technologies mainly focus on emotional recognition and health monitoring based on a single modality (such as crying audio, physiological signals or video behavior), especially under the premise that the objective physiological and subjective state of newborns and infants is difficult to describe, the ability to improve early disease warning and emotional care through non-invasive sensing methods has become a consensus in the industry.
[0003] Current mainstream technical solutions can be divided into the following categories: (1) single-channel physiological signal monitoring, such as basic vital sign change detection based on electrocardiogram, skin electricity or respiratory waveform, relying on statistical features or time-frequency domain analysis methods; (2) crying emotional recognition based on audio signals, using MFCC, fundamental frequency curve, time series deep learning and other signal processing and pattern recognition algorithms to map the crying of infants into simple physiological or psychological state classification; (3) expression and body movement evaluation based on video behavior analysis, relying on convolutional neural networks (CNN), time series modeling and other methods to classify facial expressions and movement patterns. These technologies, while having certain automation and accuracy, have gradually evolved into a new trend of multi-modal fusion and intelligent auxiliary diagnosis, which attempts to integrate multi-source signal data modeling to improve comprehensive discrimination ability.
[0004] In the application scenarios of existing technologies, systems for infant state evaluation mainly recognize basic emotional / health classification through single signal acquisition and standardized feature extraction, which has a certain universality, but cannot meet the needs of higher-dimensional, more complex scenarios for multi-modal deep fusion and traditional Chinese medicine semantic mapping. The existing technologies have the following outstanding problems:
[0005] (1) The theoretical mapping of five tones and five internal organs is subjective and ambiguous, which is difficult to quantify and adapt to individual actual situations, resulting in limited accuracy when using this theory to infer the health / emotion of infants;
[0006] (2) Traditional methods rely heavily on static mapping and empirical rules, making it difficult to handle the many-to-many, non-unique or fuzzy coupling relationship between five tones and five internal organs, and lack the ability to model complex or cross-state situations;
[0007] (3) Single modality recognition schemes (only relying on crying, only relying on electrocardiogram, etc.) are highly sensitive to noise, environmental changes and individual differences, have weak comprehensive discrimination ability, and have limited clinical practical value.
[0008] (4) Multi-modal data early fusion technology fails to dynamically adjust the weight according to the actual contribution of each mode and time-varying characteristics, and cannot realize online adaptive adjustment of the value and trust of multi-source information, so the advantages of data fusion cannot be fully exerted;
[0009] (5) The existing intelligent system mostly uses general machine learning models, lacks organic combination of Chinese medicine semantic knowledge, prior five sound-five internal organs rules and modern feature engineering, fuzzy mathematics, dynamic Bayesian modeling method, and cannot output comprehensive evaluation report which can be clinically explained and has human-computer intelligence. SUMMARY
[0010] The present application provides a baby state recognition method and system based on Chinese medicine five sound monitoring analysis to solve the above technical problems.
[0011] The technical solution of the present application is as follows: a baby state recognition method based on Chinese medicine five sound monitoring analysis, comprising:
[0012] S1: collecting the crying sound audio signal, physiological signal and behavior video data of the baby as multi-modal input samples;
[0013] S2: pre-emphasizing, windowing and fast Fourier transform processing the crying sound audio signal to extract mel frequency cepstral coefficient (MFCC) as audio feature vector;
[0014] S3: time-frequency analysis and feature selection on the physiological signal to extract key physiological feature parameters such as heart rate variability, skin electric response and respiratory rate;
[0015] S4: key frame extraction and action recognition processing on the behavior video data to extract baby facial expression and body action features based on convolutional neural network;
[0016] S5: establishing a mapping relationship table between five sound features and five internal organs states based on Chinese medicine five sound theory, and preliminarily matching the extracted audio features with five sound classification;
[0017] S6: constructing a multi-modal feature fusion model to splice the audio feature vector, physiological feature parameter and video action feature to form a fusion feature vector;
[0018] S7: calculating the dynamic weight coefficient of each modal feature in emotion recognition based on historical recognition results and current feature matching degree to adapt to individual differences;
[0019] S8: inputting the weighted fusion feature vector into a lightweight nonlinear classifier to perform baby emotion state recognition reasoning and output emotion classification label;
[0020] S9: judge whether the current emotion recognition result is consistent with the historical trend, if not consistent, start the feedback adjustment mechanism, update the dynamic weight coefficient and the classifier parameter;
[0021] S10: map the recognized emotional state label and the corresponding five viscera association information, output the comprehensive physiological and emotional state evaluation result of the infant.
[0022] The application also provides an infant state recognition system based on TCM five-tone monitoring analysis, which adopts the above-mentioned infant state recognition method based on TCM five-tone monitoring analysis to recognize the infant state.
[0023] The infant state recognition method and system based on TCM five-tone monitoring analysis provided by the application have the following beneficial effects:
[0024] (1) The application fuses infant crying audio, physiological signals (such as electrocardiogram, skin electricity, and respiration) and behavior video three modal characteristics, breaks through the problem of low recognition accuracy and high misjudgment rate caused by environmental noise, individual pronunciation difference, and pathological feature ambiguity in the existing single audio or single physiological signal method. Through the complementation and joint optimization of multi-modal characteristic information, the emotional state recognition accuracy is greatly improved compared with the traditional single modal recognition, and the model robustness is significantly improved, especially in the case of complex noise environment and large individual performance difference;
[0025] (2) The application innovatively introduces the fuzzy set theory and the Bayesian inference mechanism, models the non-unique and fuzzy membership relationship between TCM five tones and five viscera, and generates a quantitative five-tone-five-viscera fuzzy association matrix. This technology not only reduces the human subjective judgment and experience error in the process of five-tone-five-viscera correspondence, makes the system have higher theoretical self-consistency, but also dynamically corrects the five-viscera attribution probability of five-tone features based on real-time physiological state and feedback information, realizes the intelligent and interpretable application of TCM theoretical knowledge. Compared with the traditional one-to-one static mapping method of five-tone-five-viscera, the association consistency is significantly improved, and the key bottleneck in the landing engineering of TCM semantics is eliminated;
[0026] (3) The application designs a dynamic weight adjustment module based on historical recognition results and feature matching degree, uses the mutual information index to adaptively measure the time sequence contribution degree of each modality, and dynamically adjusts combined with individual historical recognition accuracy and feedback, realizes the precise modeling and optimal modal fusion of different infant individuals in different scenes and different states. This overcomes the weak migration of the "fixed weight / static rule" model and the poor individual adaptation ability, ultimately improves the adaptation range of the system in the population universality and individual special recognition effect, significantly improves the consistency and continuity of individual emotion tracking, and greatly expands the application value of clinical and family care;
[0027] (4) The present application realizes the traceability of the whole process, semantic transparency and medical interpretation by establishing a multi-level correlation mapping link of five sound characteristics-emotional state-five internal organs physiology, and outputting the recognition result in a structured and visualized evaluation report. On the one hand, it establishes a standard operation paradigm for the integration of traditional Chinese medicine theory and modern medical data, and on the other hand, it improves the clinical reference value and user trust of the output result. Compared with the existing "black box" neural network emotion recognition model, the explainability and medical guidance value are significantly enhanced;
[0028] (5) The present application adopts lightweight feature extraction, principal component analysis dimension reduction, nonlinear classifier and incremental learning algorithms in the whole process, which guarantees low delay response under the premise of high performance recognition, meets the real-time requirement of infant emotional health monitoring, reduces system resource consumption, facilitates hardware integration and edge deployment. Compared with traditional heavy models, the inference delay is significantly reduced, which is suitable for mobile terminals, infant monitoring devices and other multi-scene applications. BRIEF DESCRIPTION OF DRAWINGS
[0029] Fig. 1 A flowchart of the infant state recognition method based on the five-tone monitoring analysis of traditional Chinese medicine of the present application;
[0030] Fig. 2 A sub-flowchart of the infant state recognition method based on the five-tone monitoring analysis of traditional Chinese medicine of the present application;
[0031] Fig. 3 Another sub-flowchart of the infant state recognition method based on the five-tone monitoring analysis of traditional Chinese medicine of the present application. DETAILED DESCRIPTION
[0032] In order to make the purpose and advantages of the present application clearer and more apparent, the present application will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0033] The preferred implementation methods of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation methods are only used to explain the technical principles of the present application and do not limit the protection scope of the present application.
[0034] As used herein, the singular forms "a", "an" and "the" can also include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprise / contain" or "have" and the like specify the presence of stated features, integers, steps, operations, components, parts or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, components, parts or combinations thereof. At the same time, the term "and / or" used in the specification includes any and all combinations of the related listed items.
[0035] Please refer to Figs. 1-3 The infant state recognition method based on traditional Chinese medicine five-tone monitoring analysis includes:
[0036] S1: Collect the crying audio signal, physiological signal, and behavior video data of the infant as multi-modal input samples;
[0037] S2: Pre-emphasize, window, and fast Fourier transform the crying audio signal to extract the mel frequency cepstral coefficient (MFCC) as the audio feature vector;
[0038] S3: Perform time-frequency analysis and feature selection on the physiological signal to extract key physiological feature parameters such as heart rate variability, skin electrical response, and respiratory rate;
[0039] S4: Perform key frame extraction and action recognition processing on the behavior video data, and extract infant facial expression and body movement features based on a convolutional neural network;
[0040] S5: Establish a mapping relationship table between five-tone features and five-organ states based on the theory of traditional Chinese medicine five tones, and preliminarily match the extracted audio features with five-tone classification;
[0041] S6: Construct a multi-modal feature fusion model to concatenate the audio feature vector, physiological feature parameter, and video action feature to form a fusion feature vector;
[0042] S7: Based on the historical recognition results and the current feature matching degree, calculate the dynamic weight coefficient of each modal feature in emotion recognition to adapt to individual differences;
[0043] S8: Input the weighted fusion feature vector into a lightweight nonlinear classifier to perform infant emotion state recognition reasoning and output an emotion classification label;
[0044] S9: Determine whether the current emotion recognition result is consistent with the historical trend. If not, start the feedback adjustment mechanism to update the dynamic weight coefficient and the classifier parameters;
[0045] S10: Map the recognized emotion state label and the corresponding five-organ related information to output the comprehensive physiological and emotional state evaluation result of the infant.
[0046] The step S1: Collect the crying audio signal, physiological signal, and behavior video data of the infant as multi-modal input samples. Specifically includes:
[0047] S1.1: Collect the infant crying audio signal through a high-sensitivity microphone to obtain raw audio data containing five-tone features;
[0048] The collected target is set as an infant in different emotional states, the input condition is that the high-sensitivity microphone collection end and the anti-noise audio front-end circuit are in working state, and the sampling environment has background noise control measures;
[0049] A high-sound-pressure-sensitivity condenser microphone unit (parameters: sensitivity -35dB±1dB, frequency response range 20Hz to 8kHz, signal-to-noise ratio≥60dB) is used to realize sound pressure response collection of infant crying in the full bandwidth range;
[0050] Further, through the built-in low-noise preamplifier (parameters: equivalent input noise 1.2μV, gain 40dB), the amplitude of the original microphone electrical signal is improved, and an analog signal output covering five sound frequency bands is obtained;
[0051] Further, through the Σ-Δ type analog-to-digital converter (parameters: quantization bit depth 24bit, sampling rate 48kHz), analog signal quantization sampling is realized, and a digital audio stream with time continuity and amplitude accuracy is generated;
[0052] Further, through the high-order finite impulse response (FIR) band-pass filter algorithm (parameters: passband range 150Hz to 5000Hz, stopband attenuation≥60dB), the frequency band of the digital audio stream is limited, and the spectral components below the lower limit of the infant crying characteristic and above the upper limit of the harmonic are removed, and the effective frequency band data containing the five sound characteristics are retained;
[0053] Through the automatic gain control (AGC) algorithm (parameters: target output level -12dBFS, attack time 10ms, release time 100ms), the result of the previous step is converted into original audio data with dynamic amplitude balance, realizing the amplitude consistency control of the input of the subsequent feature extraction stage;
[0054] For example, in a quiet indoor sampling environment, a directional heart-shaped condenser microphone with a sensitivity of -35.5dB, a frequency response of 20Hz-8kHz, and a signal-to-noise ratio of 62dB is selected. The signal is boosted to 1.5Vpp by the built-in low-noise preamplifier, and is sampled at 48kHz by the 24-bit Σ-Δ ADC. Subsequently, a fourth-order FIR band-pass filter (passband 150Hz to 5000Hz) is applied to remove background air conditioner low frequency and environmental electromagnetic high frequency interference, and a digital audio stream with five sound characteristic fidelity≥0.98 is output. After AGC control, the signal dynamic range is fixed at -12dBFS, which ensures the consistency of the collected results under different crying intensities, and finally generates original audio data that can be directly input to the S2 step for MFCC feature extraction;
[0055] S1.2: Collect the electrocardiogram, skin conductance response and respiratory waveform signals of infants based on a medical-grade physiological monitoring device to obtain physiological data reflecting the status of the autonomic nervous system;
[0056] The collection object based on the medical-grade physiological monitoring device is the electrocardiogram, skin conductance and respiratory signals of infants in a resting or natural state. The input condition is that the electrode attachment and zero drift calibration of the multi-channel physiological signal collection end have been completed, and the data collection environment meets the medical electrical safety standards and has electromagnetic interference shielding measures;
[0057] A three-lead medical electrocardiogram collection module (parameters: sampling rate 1000 Hz, resolution 24 bit, input impedance ≥ 10 MΩ) is used to realize high-precision bipolar differential collection of the electrical activity of the infant's heart and output the original electrocardiogram waveform data;
[0058] Further, a low-pass-trap combined filter (low-pass cutoff frequency 150 Hz, trap center frequency 50 Hz, stopband attenuation ≥ 60 dB) is used to suppress power frequency and high frequency noise of the electrocardiogram signal, and the noise-suppressed electrocardiogram waveform signal is obtained;
[0059] Further, a skin conductance signal collection channel equipped with a constant current source excitation (parameters: excitation current 10 μA, sampling rate 100 Hz, resolution 16 bit) is used to realize high-sensitivity measurement of the skin conductance response (SCR) conductance change of the palm or sole, and generate the original skin conductance signal stream;
[0060] Further, the skin conductance signal is smoothed by an adaptive moving average filter (window length 2 s) to extract slow change (SCL) and fast change (SCR) component feature sequences, and obtain low-noise skin conductance time series data;
[0061] Further, a differential pressure sensor and a nasal cavity thermistor element composite respiratory waveform collection channel (parameters: range -2 kPa to 2 kPa, thermistor sensitivity 0.1℃, sampling rate 200 Hz) is used to realize synchronous detection of the air pressure change and temperature fluctuation caused by respiratory movement, and output the original respiratory dual-mode signal stream;
[0062] Further, a band-pass filter (passband 0.05 Hz to 2 Hz) is used to band-limit the respiratory signal to obtain the effective respiratory waveform without low-frequency drift and high-frequency interference, and extract the instantaneous respiratory frequency and amplitude parameters;
[0063] Through a multi-channel synchronous pulse trigger mechanism, the aforementioned electrocardiogram, skin conductance and respiratory signals are collected and uniformly time-stamped under the same hardware clock reference to generate a synchronous original data set containing three types of physiological signals, realizing high temporal consistency and multi-dimensional basic physiological signal input for subsequent emotional and autonomic nervous state analysis.
[0064] Exemplary, in the isolated electromagnetic noise of the mother and infant monitoring ward environment, the infant's two upper limbs and left lower limb are attached with disposable medical ECG electrodes respectively, a three-lead ECG module is sampled at 1000 Hz, 24-bit A / D conversion, and the detected ECG signal is filtered by low-pass 150 Hz and 50 Hz notch, and the output baseline drift is ≤0.05 mV waveform. Attach a skin electrode on the palm of the same side, with a constant current of 10 μA excitation, output 16-bit resolution SCR signal, 2s sliding average filtering, get SCL curve with mean square noise amplitude lower than 0.01 μS. A differential pressure sensor and a thermistor element are arranged below the alae nasi, with a sampling rate of 200 Hz, monitoring the pressure amplitude of 0.05-0.2 kPa and temperature fluctuation of 0.2-0.5℃ in the inspiration and expiration cycle, after 0.05-2 Hz band-pass filtering, the respiratory waveform with a coefficient of variation of respiratory cycle ≤5% is obtained. The three types of signals are stored as continuous 10-minute raw data sets by hardware synchronization and unified timestamp, and the subsequent heart rate variability, skin electricity reaction indicators and respiratory frequency can be directly sent to the S3 step to realize sensitive monitoring of the changes of the infant's autonomic nervous state;
[0065] S1.3: Use a high-definition camera to record the infant's face and limb movements in video to obtain the action and expression change sequence in the behavior performance layer;
[0066] S1.4: Perform timestamp synchronization processing on the collected multi-modal raw data to ensure that the audio, physiological and video signals are accurately aligned on the time axis;
[0067] S1.5: Package the synchronized multi-modal raw data into a structured sample set as the input basis for the subsequent feature extraction and traditional Chinese medicine model matching module.
[0068] The step S2: The crying audio signal is pre-emphasized, windowed and fast Fourier transformed to extract the mel frequency cepstral coefficient (MFCC) as the audio feature vector. Specifically, it includes:
[0069] S2.1: The collected infant crying audio signal is pre-emphasized, a first-order difference filter is used to enhance the high-frequency components of the original audio waveform to compensate for the energy attenuation of the speech signal in the low frequency band, and the pre-emphasized audio waveform data is obtained;
[0070] The collected infant crying audio signal is processed by a first-order difference filter (parameters: pre-emphasis coefficient a=0.97) to enhance the high-frequency components of the audio waveform to compensate for the low-frequency energy advantage caused by the sound production mechanism;
[0071] Further, the differential output is smoothed by a finite impulse response (FIR) smoothing filter (parameters: order 5, window type: Hanning window) to suppress the instantaneous peak noise after high-frequency enhancement, and a pre-emphasis smoothed waveform data is generated;
[0072] Further, amplitude normalization transformation (parameters: normalized to 1 by the maximum absolute value) is performed to achieve consistent waveform amplitude processing under different recording amplitudes, and standardized pre-emphasis audio data is generated for subsequent windowing and frequency domain analysis;
[0073] Through the above first-order difference, high-frequency enhancement, smoothing filtering and amplitude normalization processing, the original audio waveform is converted into audio waveform data with enhanced high-frequency details and consistent dynamic amplitude, which improves the high-frequency spectral resolution and parameter stability in the subsequent MFCC feature extraction stage;
[0074] For example, in a 48kHz sampling rate, 24bit quantization accuracy, and a 10 second long, -20dBFS to -5dBFS dynamic range of the crying sound, the pre-emphasis coefficient α = 0.97 is selected for point-by-point processing, and the enhanced high-frequency energy is increased by about 15% in the mid-high frequency band. On this basis, the instantaneous high-frequency noise is suppressed by using a Hanning window FIR smoothing filter (order 5), and the signal-to-noise ratio is improved to 68dB. After maximum absolute value normalization processing, the amplitude distribution is stable in [-1, 1], and the output standardized pre-emphasis waveform effectively improves the spectral edge resolution and the noise stability of the MFCC coefficient in the subsequent windowing and FFT analysis. The pre-processing result is directly used as the input of the S2.2 windowing and framing operation, ensuring high fidelity and robustness in the whole process of MFCC feature extraction;
[0075] S2.2: Perform windowing operation on the pre-emphasized audio waveform data, and use Hamming window function to frame the signal to reduce the spectral leakage caused by the discontinuity between frames, and obtain a plurality of windowed audio frame signals;
[0076] S2.3: Perform fast Fourier transform (FFT) processing on each windowed audio frame signal to convert the time domain signal to frequency domain representation, and obtain the corresponding spectral amplitude distribution characteristics;
[0077] S2.4: Apply a Mel filter bank to the frequency domain representation of the audio signal to convert the linear frequency scale to the Mel scale, and obtain the energy response value of each filter channel;
[0078] S2.5: Perform discrete cosine transform (DCT) processing on the energy response value of each Mel filter channel to extract low-order cepstrum coefficients, and obtain the Mel frequency cepstrum coefficient (MFCC) feature vector as the input audio feature of the subsequent multi-modal fusion and emotion recognition model.
[0079] The step S3: time-frequency analysis and feature selection on the physiological signal, extracting heart rate variability, skin electric response and respiratory rate and other key physiological characteristic parameters. As shown in the figure, specifically including: Fig. 2
[0080] S3.1: wavelet denoising processing on the collected infant ECG signal to remove electromyographic interference and baseline drift noise, obtaining the denoised ECG waveform sequence as the input of subsequent R wave detection;
[0081] S3.2: based on the denoised ECG waveform sequence, using Pan-Tompkins algorithm to perform R wave peak detection to obtain R-R interval sequence as the basic input data for calculating heart rate variability;
[0082] On the infant ECG waveform sequence processed by S3.1 wavelet denoising, Pan-Tompkins algorithm (parameters: band-pass filter passband range 5Hz to 15Hz, integral window width 150ms, threshold coefficient 0.5) is used to realize R wave peak detection function;
[0083] Further, through the band-pass filter, the baseline drift and high-frequency electromyographic noise in the ECG signal are suppressed, and the middle-frequency component sensitive to R wave is retained, and the band-pass filtered enhanced signal sequence is obtained;
[0084] Further, through first-order difference operation, the slope change characteristics of QRS complex are enhanced, and the difference sequence is obtained;
[0085] Further, through square operation, high-amplitude difference components are strengthened to highlight R wave peaks and suppress low-amplitude interference, generating a square amplification sequence;
[0086] Further, through moving window integration (window length 150ms), local energy is calculated to obtain an integral signal sequence representing QRS region energy accumulation;
[0087] Further, based on the joint double-threshold judgment of the integral signal and the square signal (the main threshold is calculated based on the average peak value of the signal, and the secondary threshold is used to process low-amplitude R waves), the possible R wave positions are identified, and combined with the physiological constraints (the interval between adjacent R waves is not less than 200ms), the false peaks are filtered, and the accurate R wave peak index set is obtained;
[0088] Through the above Pan-Tompkins processing method, the denoised ECG waveform is effectively converted into R-R interval sequence reflecting the heart beat cycle, realizing the high-precision basic input data required for heart rate variability calculation;
[0089] Exemplarily, in the infant ECG signal with a sampling rate of 500 Hz, the band-pass filter passband is set to 5 Hz to 15 Hz, the maximum instantaneous slope value obtained by differential operation is 1.2 mV / sampling point, and the peak value after square operation is 1.44 mV 2 The integral window length is set to 75 sampling points (corresponding to 150 ms), and the double thresholds are set to 0.6 and 0.3 of the global peak value, respectively. In a 10-minute recording, a total of 1280 R waves are detected, and after eliminating 5 false peaks by the constraint condition, the mean R-R interval is 0.47 s, and the standard deviation is 0.08 s. When the S3.3 is provided to calculate the SDNN and RMSSD, the verification result shows that compared with the direct peak value identification method without Pan-Tompkins detection, the R wave detection rate is increased by 4.2%, and the coefficient of variation of the heart rate variability parameters is reduced by 6.5%, realizing the stable and accurate extraction of the R-R interval sequence;
[0090] S3.3: Time domain statistical analysis is performed on the R-R interval sequence to calculate the standard deviation (SDNN) and the root mean square of adjacent R-R interval differences (RMSSD) to extract the heart rate variability time domain feature parameters reflecting the activity of the autonomic nervous system;
[0091] The R-R interval sequence output by S3.2 is subjected to time domain statistical analysis (parameters: window range is the full sequence length) to calculate the heart rate variability standard deviation.
[0092] Further, by constructing the adjacent interval difference sequence, the root mean square calculation method (parameters: the number of elements participating in the operation is N-1) is used to extract the RMSSD value, and the formula
[0093]
[0094] The fluctuation intensity of each continuous heart beat cycle is obtained, and the RMSSD reflects the activity level of the high-frequency autonomic nervous system;
[0095] Further, by jointly analyzing SDNN and RMSSD, a double-index heart rate variability feature vector is established, and the vector is input to the subsequent multi-modal fusion model to provide stable physiological representation for emotion recognition;
[0096] Further, the feature vector is subjected to Z-score standardization processing (parameters: mean and standard deviation are obtained by statistical analysis of training samples) to unify the feature scales of different infant individuals, thereby realizing the model adaptability of multiple samples;
[0097] Through the above double-path operation based on full sequence statistics and adjacent difference, the R-R interval sequence is converted into time domain indicators with global and local heart rate fluctuation characteristics, realizing the synchronous description of the overall activity and short-term fluctuation of the autonomic nervous system.
[0098] Exemplarily, in an electrocardiogram signal with a sampling rate of 500 Hz and a total duration of 300 seconds, 420 R-waves are detected by Pan-Tompkins, the length of the R-R interval sequence N = 420, and the arithmetic mean μRR = 0.71 seconds. According to the SDNN formula, the variance is 0.0144 seconds2, and the SDNN is 0.12 seconds. The sum of squares of adjacent interval differences is 6.084 seconds2, and the RMSSD is 0.12 seconds obtained by dividing 419 and taking the square root. The mean and amplitude of the two indicators show high consistency in multiple groups of infant data, the stability coefficient of variation is less than 5%, and the weight correlation analysis in the multi-modal fusion model shows that the contribution of the time-domain heart rate variability feature to the emotional state classification is more than 0.25 (mutual information coefficient), which significantly improves the model's ability to distinguish high arousal and low arousal emotional states;
[0099] S3.4: Perform sliding window Fourier transform on the infant skin electricity signal to extract the amplitude peak value and reaction frequency of the skin electricity response (SCR) as a skin electricity feature parameter reflecting the emotional arousal degree;
[0100] After extracting the heart rate variability features of the infant skin electricity signal sequence by S3.3, the sliding window Fourier transform method (parameters: window length 2 seconds, window overlap rate 50%) is used to realize the frequency spectrum distribution extraction of the skin electricity signal in the local time interval, so as to retain the dynamic change characteristics of the skin electricity response in the time and frequency domains;
[0101] Further, by performing fast Fourier transform (FFT point number 512) operation in each sliding window, the time domain skin electricity waveform is converted into frequency domain amplitude spectrum, and the amplitude vector A k ;
[0102] Further, the peak value detection algorithm (threshold set to 3 times the standard deviation offset of the global mean amplitude) is used to identify the amplitude peak points in the skin electricity ripple response, and the maximum peak amplitude A max is taken as the amplitude feature of the skin electricity response to realize the quantification of the strongest skin electricity response intensity;
[0103] Further, based on the peak time sequence in each window, the average time interval T between adjacent peaks is calculated, and the inverse proportional operation formula is used:
[0104]
[0105] The occurrence frequency feature index f of the skin electricity response is obtained;
[0106] Further, the amplitude feature and the frequency feature are subjected to feature vector construction and min-max normalization processing (parameter: normalization range [0, 1]) to eliminate the amplitude and frequency scale differences between different infants and batches of collected data;
[0107] Through the above sliding window Fourier transform, peak detection and feature normalization processing, the galvanic skin response signal is converted into a double-index feature vector containing the amplitude peak value of the skin electrical response and the occurrence frequency, and the emotional arousal degree of the infant is stably quantified and described;
[0108] For example, in a galvanic skin signal sequence with a sampling rate of 100 Hz, a sliding window length of 2 seconds (corresponding to 200 sampling points) and an overlap rate of 50% are selected for FFT transformation, and a frequency spectrum with a resolution of 0.5 Hz is obtained. The peak detection threshold is set to the mean amplitude + 3 times the standard deviation, and a total of 8 skin electrical response peaks are detected in the 30-second signal, with the maximum peak amplitude A max 12.4 μS. The average interval between adjacent peaks T = 3.6 seconds, and the frequency The double feature normalization is 0.82 (amplitude) and 0.27 (frequency), respectively, and the weight contribution of the high arousal emotion sample in the multi-modal fusion model mutual information analysis is more than 0.22, which significantly improves the detection accuracy of high arousal emotions such as terror and violent crying;
[0109] S3.5: The respiratory cycle waveform is extracted by using an adaptive filtering method for the respiratory signal, the respiratory frequency is calculated based on a peak detection algorithm, and the respiratory depth change rate per unit time is calculated to obtain the respiratory feature parameters reflecting the emotional state.
[0110] The step S4: key frame extraction and action recognition processing are performed on the behavior video data, and the convolutional neural network is used to extract the facial expression and body movement features of the infant. As shown in Fig. 3 , specifically includes:
[0111] S4.1: Key frame extraction processing is performed on the collected infant behavior video data, and based on the video frame difference calculation and adaptive threshold judgment, video key frames with significant action changes are extracted to reduce the interference of redundant frames on subsequent processing;
[0112] For the collected infant behavior video data, the inter-frame difference calculation method (parameter: difference measure method is absolute pixel difference, threshold update step is 0.05) is used to realize the difference calculation function of the luminance matrix of adjacent video frames to quantify the degree of pixel change in consecutive frames.
[0113] Further, the global pixel difference distribution is calculated by the difference histogram statistics method (parameter: histogram bin number 64), and a global distribution feature vector of the inter-frame difference is obtained, which is used for subsequent threshold judgment;
[0114] Further, the adaptive threshold judgment algorithm (parameter: initial threshold set to the average difference mean value of the history + 1.5 times the standard deviation) is used to compare the current inter-frame difference with the dynamic threshold, and a binary judgment result is generated to mark the candidate frames of significant motion changes;
[0115] Further, the time neighborhood suppression algorithm (parameter: suppression window length 3 frames) is used to remove redundant candidate key frames in the same motion occurrence period, so as to ensure that only one key frame is retained for each independent motion event, and the efficiency and representativeness of subsequent processing are improved;
[0116] Further, the key frame index cache mechanism (parameter: cache depth 100 frames) is used to record the time position and frame number of all valid key frames in the original video stream, so as to provide accurate input for the subsequent face region detection and positioning step of S4.2;
[0117] Through the above difference calculation, adaptive threshold judgment and neighborhood redundancy removal, the original video sequence is converted into a key frame image sequence containing significant motion change information, the redundant frames are effectively removed and compressed, and the visual input features with time sequence representativeness are provided for emotion recognition;
[0118] For example, in a behavior video with a resolution of 1920x1080 and a frame rate of 30fps, the difference measurement method is set to the absolute pixel difference of the whole frame, and the average difference range of the two consecutive frames is between [0.0, 0.15]. The calculated average difference of the history is 0.07, and the standard deviation is 0.02. Based on the initial threshold formula
[0119] T init = 0.07 + 1.5 x 0.02
[0120] T init = 0.10. Under this threshold condition, 465 candidate key frames are marked, and after time neighborhood suppression, 328 key frames are retained, accounting for 3.65% of the total number of frames. In the motion detection accuracy evaluation, compared with the fixed threshold method, the redundant frame redundancy rate is reduced by 18.6%, the face feature extraction module processing frame rate is improved by 22.3%, and in the multi-modal emotion recognition experiment, the overall recognition accuracy is improved by 3.8%, which verifies the effectiveness of the key frame extraction strategy in reducing the calculation overhead and retaining dynamic information;
[0121] S4.2: Perform face region detection and positioning processing on the extracted key frame image sequence, use a multi-task learning model based on a cascade convolutional neural network to locate and key point detection of infant face area, to obtain the basic analysis area of facial expression change;
[0122] For the infant key frame image sequence extracted by S4.1 step, use a cascade convolutional neural network (parameters: cascade layer number 3, convolution kernel number [32, 64, 128] per layer, window size [24x24, 48x48, 96x96]) to realize the coarse positioning and fine positioning function of face candidate region;
[0123] Further, through the candidate region screening algorithm (parameters: confidence threshold 0.85), the confidence of the preliminary detected rectangular region is sorted, and the low confidence face candidate frame is removed, and a high confidence face candidate set is obtained;
[0124] Further, the non-maximum suppression (NMS) algorithm (parameters: IoU overlap rate threshold 0.3) is used to remove the boundary box of the high confidence candidate set, and the detection frame with the smallest overlap rate and the highest confidence is retained to eliminate the redundant marking of multiple frames to the same face, and a unique face positioning result is generated;
[0125] Further, within the face positioning frame, the key point detection branch of the multi-task learning model (parameters: output key point number 5 including eyes, nose tip and corner of mouth) is called to perform feature point positioning, and a regression loss function is used to regress the key point coordinates at the pixel level, and a key point set K p is outputted, which contains the accurate face geometric structure;
[0126] Further, based on the face positioning result and the key point set, perform face affine transformation and pose normalization processing, calculate the face rotation angle θ and the scaling ratio s, and according to the formula
[0127] P'=s×R(θ)×P
[0128] Realize the geometric correction of face image, where R(θ) is the rotation matrix, P is the original pixel coordinate matrix, and P' is the transformed coordinate matrix;
[0129] Through the above cascade convolutional neural network detection, multi-task key point regression and pose normalization processing, the key frame image sequence is converted into an analysis basic image that is accurately aligned and covers the main face area and geometric structure, realizing high-precision input preparation for facial expression change capture;
[0130] Exemplarily, in a baby key frame sequence with a resolution of 1920x1080, the cascade convolutional neural network has three convolution kernel sizes of 24, 48 and 96 pixels respectively, a step size of 2, an IoU suppression threshold of 0.3, and a confidence threshold of 0.85. On average, 1.05 face candidate boxes are detected per frame, and after NMS deduplication, the number of effective boxes is reduced to 1.0. The average positioning error of the five feature points output by the key point detection branch is less than 2 pixels. The calculated face rotation angle in the pose normalization ranges from -15° to 15°, and the scaling ratio ranges from 0.85 to 1.15. The final generated aligned face image in the subsequent S4.3 feature extraction, compared with the original image without normalization processing, the intra-class variance of the expression feature vector is reduced by 12.8%, the inter-class Euclidean distance is increased by 9.5%, and in the multi-modal fusion emotion recognition test, the overall classification accuracy is improved by 4.1%, which verifies the technical value of the face region detection and positioning step for improving the expression feature quality;
[0131] S4.3: Perform expression feature extraction processing on the positioned face image region, and extract local texture and global morphological features based on a lightweight convolutional neural network model to obtain a face expression feature vector for infant emotion recognition;
[0132] A lightweight convolutional neural network model (parameters: 4 layers of convolution, the number of convolution kernels in each layer is [32, 64, 128, 256] in turn, the convolution kernel size is uniform 3x3, and the step size is 1) is used to realize the multi-layer feature extraction function of the face image on the pose normalized face image output by the S4.2 step;
[0133] Further, local response normalization (parameters: normalization radius 5, constant coefficient 2) is used to suppress the high amplitude response of local mutation to enhance the contrast sensitivity to local texture details and obtain a normalized texture activation map;
[0134] Further, a multi-scale convolution channel parallel structure (parameters: convolution kernel size set {33, 55, 77}) is used to extract local texture and global morphological features of the face region under different receptive fields, and a multi-scale feature mapping matrix F m is formed through splicing operation;
[0135] Further, a global average pooling layer is used to calculate the average of the morphological response of each channel to obtain a global morphological feature vector v g , while retaining the local texture feature mapping F l output by the previous convolution layer to realize the fusion of local and global features;
[0136] Further, F l and v g are weighted and fused in the channel dimension (weight vector αl with α g satisfying the normalization condition: α l + α g = 1), to obtain the fused facial expression feature vector V f ;
[0137] By the above-mentioned lightweight convolutional neural network multi-layer convolution, local response normalization, multi-scale convolution and local-global feature fusion processing mode, the posture-normalized facial image is converted into a high-discriminative facial expression feature vector representing the emotional state of the infant, and the facial features are effectively input in the multi-modal emotion recognition;
[0138] For example, in a posture-normalized infant facial sample with a resolution of 112x112 pixels, the convolution kernel size is set to 3x3, 5x5 and 7x7, and the total number of feature channels obtained after multi-scale parallel convolution is 224. Under the action of local response normalization radius 5, the pixel variance of the texture activation map is reduced by 14.7%, the morphological feature vector length is 256 dimensions, and the local feature mapping compression length is 768 dimensions. The fusion weights are set to α l = 0.65, α g = 0.35, and the final output fused facial expression feature vector length is 1024 dimensions. The feature vector contributes to a 5.6 percentage point increase in classification accuracy for high-arousal emotions in a multi-modal emotion recognition model, and the single-frame processing delay is controlled within 3.4 milliseconds, meeting the real-time requirement;
[0139] S4.4: Perform limb action region segmentation processing on the key frame image sequence, and use a video action recognition model based on a spatio-temporal convolution network to perform pixel-level segmentation on the infant limb movement region to extract the spatial distribution features of the limb action;
[0140] S4.5: Perform action feature encoding processing on the segmented limb action region, and extract the time continuity and spatial motion pattern features of the action based on a three-dimensional convolutional neural network to generate a limb action feature vector for multi-modal fusion.
[0141] The step S5: based on the five-tone theory of traditional Chinese medicine, a mapping relationship table between five-tone features and five-organ states is established, and the extracted audio features are preliminarily matched with five-tone classification. Specifically, it includes:
[0142] S5.1: Perform normalization processing on the extracted audio feature vector to eliminate the influence of different infant pronunciation intensities and acquisition device differences on feature distribution, and obtain a standardized audio feature matrix as the input basis for five-tone classification;
[0143] S5.2: Based on the pre-defined five-tone pitch reference values, the dynamic time warping algorithm is used to perform pattern matching on the fundamental frequency sequence in the standardized audio feature matrix to identify the most matching tone category for the palace, shang, jiao, zhi, and feather five tones, and obtain the five-tone classification result;
[0144] For the standardized audio feature matrix output by the S5.1 step, a fundamental frequency extraction method based on pre-defined five-tone pitch reference values is used (parameters: analysis window length 40 ms, frame shift 10 ms, fundamental frequency search range [80, 1000] Hz) to accurately estimate the fundamental frequency f0 of each frame of audio signal;
[0145] Further, through the fundamental frequency smoothing filtering algorithm (parameters: median filtering window length 5 frames), the instantaneous spikes and jitter are suppressed to realize the smoothing processing of the continuous frame fundamental frequency sequence, and the time domain smoothed fundamental frequency curve F s is obtained;
[0146] Further, the smoothed fundamental frequency curve is matched with each tone reference sequence in the pre-defined five-tone pitch reference table, and the dynamic time warping (DTW) algorithm (parameters: Sakoe-Chiba bandwidth limit is 10) is used to perform nonlinear time alignment between F s and each reference tone sequence, and calculate the matching cumulative distance D i ;
[0147] Further, the similarity score S
[0148]
[0149] of each tone category is calculated by the formula i , where D i is the DTW cumulative distance;
[0150] Further, the similarity score is mapped to the five-tone classification probability vector P through the Softmax normalization processing (temperature coefficient 0.5), and the tone category with the maximum probability is selected as the five-tone classification result of the current audio frame segment;
[0151] Through the dynamic time warping algorithm and the normalized probability mapping, the fundamental frequency sequence is converted into the classification labels corresponding to the palace, shang, jiao, zhi, and feather categories, realizing the high-robust matching of the five-tone features to the tone categories;
[0152] For example, in a sample rate of 16 kHz baby crying sound sample, the analysis window length is set to 40 ms, the frame shift is set to 10 ms, the autocorrelation method is used to extract the fundamental frequency, and the fundamental frequency search range is set to [80, 1000] Hz, obtaining the average fundamental frequency curve F a . For F aThe application of 5-frame median filter smoothing processing can eliminate jitter and reduce the base frequency variance by 18.4%. The smoothed curve is matched with the pre-defined reference base frequency sequence (length 100 frames) of Gong, Shang, Jiao, Zhi and Yu respectively by DTW, and the Sakoe-Chiba constraint parameter is set to 10. The cumulative distance calculated is {25.6, 40.2, 35.7, 44.8, 38.9} respectively, and the highest score category is Gong after similarity formula conversion. The Gong class probability after Softmax normalization is 0.82, and the probability of the remaining categories is less than 0.05. The five-tone classification accuracy of the method in 200 samples reaches 93.7%, which is 8.2 percentage points higher than that of the pure base frequency threshold method, verifying the effectiveness and robustness of the matching strategy;
[0153] S5.3: Perform Chinese five-tone-five-organ mapping rule matching on the identified five-tone classification results, determine the corresponding five-organ system of the current audio signal based on the five-tone-five-organ correspondence table in traditional Chinese medicine theory, and generate a five-organ association label sequence;
[0154] For the five-tone classification results output by S5.2, a five-tone-five-organ mapping rule retrieval method based on traditional Chinese medicine theory knowledge base (parameters: mapping table contains five-tone categories <Gong, Shang, Jiao, Zhi, Yu> and corresponding five-organ <spleen, lung, liver, heart, kidney>) is used to realize the initial matching function of five-tone categories to five-organ categories;
[0155] Further, the input five-tone label is quickly positioned to the corresponding five-organ item in the mapping table by an index matching algorithm (parameters: key-value index uses hash mapping structure, and the mapping conflict resolution method is chain address method), and a matching candidate list is output;
[0156] Further, for the case of one tone and multiple organs or cross-class association, a rule priority judgment module (parameters: priority assignment is based on Huangdi Neijing and clinical statistical frequency, and the value range is [0, 1]) is called to score and sort the five-organ categories with attribution conflicts in the candidate list, and generate a priority-matched five-organ category label;
[0157] Further, a multi-label encoding method (parameters: encoding bit length is 5, each position corresponds to one five-organ, and binary 1 represents the existence of association) is used to encode the priority-matched label and all secondary association labels into a five-organ association label sequence, realizing the discretization input preparation for subsequent fuzzy membership calculation;
[0158] Through the mapping retrieval, priority judgment and multi-label encoding processing methods, the five-tone classification results are converted into a structured five-organ association label sequence, realizing the assignment of traditional Chinese medicine semantic information of the audio signal;
[0159] Exemplary, in the case of inputting a five-tone classification result of palace, the system calls the mapping table retrieval result of the main associated organ spleen (priority 0.92) and the secondary associated liver (priority 0.38). The rule priority determination module confirms that the spleen is the main label, and the coding obtains the sequence [1, 0, 0, 0, 0] representing only the association with the spleen. The secondary label liver corresponds to the position coding 1, and the sequence is updated to [1, 0, 1, 0, 0]. The output five-organ association label sequence length is 5, and the main and secondary association information is completely retained. In 1000 test samples, the label sequence generated by this method is consistent with the artificial Chinese medicine practitioner annotation rate of 94.5%, which is 5.7 percentage points higher than the single mapping method, providing a high consistency of discrete input for subsequent S5.4 fuzzy membership assignment;
[0160] S5.4: Fuzzy membership assignment to five-organ association label sequence, introduce fuzzy set theory to model the non-unique mapping relationship between five tones and five organs, to quantify the influence degree of five tone features on each five organ state, output five tone-five organ fuzzy association matrix;
[0161] S5.5: Fusion of five tone-five organ fuzzy association matrix and current physiological state information of infants, update the matching probability between five tone features and five organ states based on Bayesian inference mechanism, to improve the robustness and adaptability of five tone classification and Chinese medicine state reasoning.
[0162] The step S6: construct a multi-modal feature fusion model, concatenate the audio feature vector, physiological feature parameter and video action feature to form a fusion feature vector. Specifically, it includes:
[0163] S6.1: Perform dimension normalization processing on the extracted audio feature vector to eliminate the interference of feature scale differences between different audio segments on the fusion process, and obtain a normalized audio feature matrix as the audio input for multi-modal fusion;
[0164] S6.2: Perform standardization transformation processing on the heart rate variability, skin electrical response and respiratory rate indicators in the physiological feature parameters to eliminate the baseline differences of physiological signals between individuals, and obtain a standardized physiological feature vector as the physiological input for multi-modal fusion;
[0165] In the construction process of the multi-modal feature fusion model, standardization transformation processing is performed on the physiological feature parameters to eliminate the baseline differences of signals between different infants. The input object is the heart rate variability, skin electrical response and respiratory rate quantitative feature parameter vector obtained by the previous feature extraction;
[0166] Z-score standardization method (parameters: mean and standard deviation calculated based on training set statistical values) is used to realize zero mean and unit variance transformation of each physiological feature in numerical scale;
[0167] Further, by the minimum-maximum normalization algorithm (parameters: the normalization interval is set to [0, 1]), the physiological indicators are proportionally scaled in the same interval, and the normalized physiological feature matrix is obtained.
[0168] Further, the robust standardization processing (parameters: the center value is the median, and the scale value is the interquartile range IQR) is adopted to suppress outliers of the indicators with significant skew distribution, and a robust standardized feature sequence is generated.
[0169] Further, by the feature re-weighting module (parameters: the weight initial value is set based on the information gain calculation result of the feature in the historical sample), the physiological feature vectors processed by different standardization strategies are weighted and fused to obtain a comprehensive standardized physiological feature vector.
[0170] Through the above multi-stage standardization and fusion processing mode, the original physiological feature vector is converted into a standardized physiological feature vector with strong scale consistency and insensitive to individual baseline differences, realizing high-robustness physiological data input in the multi-modal fusion stage.
[0171] For example, in the case of the heart rate variability parameters (SDNN = 85.3 ms, RMSSD = 62.1 ms) calculated from the electrocardiogram signal with a sampling frequency of 500 Hz after wavelet denoising and R-wave detection, the peak value of the skin conductance response amplitude = 2.85 μS, and the respiratory rate = 28 times / min, first, Z-score standardization is performed on each feature, and the calculation formula is
[0172]
[0173] Where μ and σ are the mean and standard deviation of the feature in the training set. Taking SDNN = 85.3 ms, mean = 75.0 ms, and standard deviation = 15.0 ms as an example, the standardized value is calculated as
[0174] Then, Min-Max normalization is performed on all features, and the calculation formula is
[0175]
[0176] Assuming that the minimum value of the skin conductance response is 1.5 μS and the maximum value is 5.0 μS, the normalized result of 2.85 μS is
[0177] For the respiratory rate, robust standardization is adopted, and the median = 26 times / min and IQR = 4, and the standardized result is
[0178] Finally, the multiple standardized results are weighted by information gain weights (SDNN weight 0.4, skin conductance response weight 0.35, and respiratory rate weight 0.25) to obtain a comprehensive standardized physiological feature vector [0.2747, 0.135995, 0.125]. The variance stability of this vector in cross-individual sample testing is improved by about 12.3%, effectively enhancing the ability of the subsequent fusion model to resist performance fluctuations caused by individual baseline differences;
[0179] S6.3: Perform feature compression processing on the facial expression vector and the limb action code in the video action feature, reduce the dimension redundancy based on the principal component analysis algorithm, to obtain a compact action feature vector as the behavior input of multi-modal fusion;
[0180] S6.4: Based on the feature splicing algorithm, the normalized audio feature matrix, the standardized physiological feature vector, and the compact action feature vector are horizontally spliced in the feature dimension to generate a unified dimension fusion feature vector, realizing the structured integration of multi-modal information;
[0181] S6.5: Perform feature correlation analysis and redundancy detection processing on the fusion feature vector generated by splicing to eliminate highly correlated or invalid feature dimensions, and obtain an optimized fusion feature vector as the input representation of the emotion recognition classifier.
[0182] The step S7: based on the historical recognition results and the current feature matching degree, the dynamic weight coefficient of each modal feature in emotion recognition is calculated to adapt to individual differences. Specifically, it includes:
[0183] S7.1: Calculate the similarity between the historical emotion recognition results and the current multi-modal feature vector to obtain a feature consistency index across time windows;
[0184] S7.2: Based on the feature consistency index, use the dynamic time warping algorithm to align the multi-modal feature sequence to extract the time correlation features between modalities;
[0185] S7.3: Perform modality-specific contribution analysis on the aligned multi-modal features, and use the mutual information method to evaluate the information gain of each modality in the current emotion recognition task;
[0186] For the multi-modal feature alignment result obtained by the dynamic time warping algorithm, the mutual information method (parameters: discretization level based on quantile division of feature value distribution, reference optimal binning number in historical tasks) is used to calculate the statistical dependence between each single modality feature and the emotion label, realizing the measurement of modality-specific information amount;
[0187] Further, the continuous modal features are discretized by the equal-frequency binning method (parameter: bin number 10) to ensure the stability of the probability distribution estimation in the mutual information calculation process, and the joint probability distribution matrix of each modal feature value and the emotion label is obtained;
[0188] Further, the modal mutual information values of different feature dimensions are averaged and summarized to form the average information gain index corresponding to each modality in the current task, so as to eliminate the bias caused by the uneven number of feature dimensions in the same modality;
[0189] The average information gain vector generated by the mutual information method quantifies the contribution of each modality to the emotion recognition task, effectively supporting the subsequent dynamic weight coefficient calculation process;
[0190] For example, in a test task, the aligned audio modal features are discretized into 10 levels, the physiological modal features are discretized into 8 levels, and the video modal features are discretized into 12 levels. In the joint probability distribution of a single feature value of the audio modality and the emotion label, the joint probability of the feature value level 3 and the emotion category "irritated" is 0.12, and the marginal probabilities are 0.20 (the feature value) and 0.30 (the emotion category), respectively. The mutual information component is 0.0722472. After counting all the combinations of levels and emotions, the average mutual information values of the audio modality, the physiological modality, and the video modality are calculated as 0.164, 0.098, and 0.213, respectively. The results show that in this batch of data, the video modality has the highest contribution, followed by the audio frequency, and the physiological signal has the lowest contribution, providing an original information basis for the initial modality weight normalization of S7.4. Through this mutual information analysis, the responsiveness of dynamic weight adjustment to the immediate information value of the modality is improved by about 15.7%, significantly enhancing the adaptability and robustness of individualized emotion recognition;
[0191] S7.4: Based on the information gain index, a modality contribution weight vector is constructed, and normalized by a Softmax function to obtain the initial weight coefficients of each modality;
[0192] S7.5: An individual adaptation factor is introduced to the initial weight coefficients, and dynamically adjusted based on the historical emotion recognition accuracy and feedback data of infants to optimize the individual difference adaptation ability in the modality fusion process;
[0193] S7.6: The adjusted dynamic weight coefficients are output to the multi-modal feature fusion model as the weighted input basis for each modality feature in the next stage of emotion recognition reasoning process.
[0194] The step S8: input the weighted fusion feature vector into a lightweight nonlinear classifier to perform infant emotion state recognition reasoning and output the emotion classification label. Specifically, it includes:
[0195] S8.1: Perform feature normalization processing on the weighted multi-modal fusion feature vector to eliminate the influence of differences in numerical scales of different modal features, and obtain a standardized fusion feature vector;
[0196] S8.2: Based on the standardized fusion feature vector, perform forward propagation calculation using the multi-layer perceptron structure in the lightweight nonlinear classifier to extract the nonlinear expression of the features in the hidden layer, and obtain the hidden layer activation value;
[0197] S8.3: Use the hidden layer activation value to perform probability distribution calculation combined with the Softmax function of the classifier output layer to generate the probability distribution vector of the infant emotional state in each category, and obtain the emotional classification confidence distribution;
[0198] Based on the hidden layer activation value vector obtained in step S8.2, the Softmax probability normalization method (parameter: the numerical stability correction term ε is set to 10 -8 ) is used to realize the normalized probability distribution calculation of each emotional category in the output space.
[0199] S8.4: According to the emotional classification confidence distribution, perform emotional category decision using the maximum probability criterion to determine the emotional state category label to which the current sample belongs, and obtain the emotional classification label result;
[0200] S8.5: Serialize and package the emotional classification label result through the output interface to generate structured emotional recognition output data as the input basis for subsequent emotional state evaluation and five viscera association mapping.
[0201] The step S9: judge whether the current emotional recognition result is consistent with the historical trend, if not consistent, start the feedback adjustment mechanism, update the dynamic weight coefficient and the classifier parameter. Specifically includes:
[0202] S9.1: Perform time alignment processing on the current emotional recognition result and the historical emotional state sequence to construct the time series data set of emotional changes as the input basis for trend consistency analysis;
[0203] S9.2: Perform local trend fitting on the emotional state sequence based on the sliding window mechanism, and use the linear regression model to model the trend of the emotional state label within the window to extract the local change direction feature of the current emotional state;
[0204] S9.3: Calculate the difference between the current emotional recognition result and the emotional prediction value based on the local trend fitting to obtain the deviation of the emotional state change, so as to quantify the consistency degree of the current recognition result and the historical trend;
[0205] S9.4: Determine the emotional state deviation amount based on the set deviation threshold, if the deviation amount exceeds the preset threshold, determine that the current recognition result is inconsistent with the historical trend, trigger the feedback adjustment mechanism to start the dynamic parameter updating process;
[0206] S9.5: After the feedback adjustment mechanism is started, based on the current emotion recognition result and the historical recognition error distribution, the gradient descent algorithm is used to update the dynamic weight coefficient online to enhance the recognition sensitivity of the current emotional features, and the updated multi-modal fusion weight parameter is obtained;
[0207] S9.6: At the same time, the output layer parameters of the lightweight nonlinear classifier are updated by incremental learning, the gradient is calculated based on the loss function of the current recognition result and the real emotional label, and the parameter fine-tuning is performed to improve the recognition accuracy of the model for the current individual emotional state.
[0208] The step S10: mapping the recognized emotional state label and the corresponding five viscera association information to output the comprehensive physiological and emotional state evaluation result of the infant. Specifically, it includes:
[0209] S10.1: Construct a five-tone-five-organ-emotion mapping dictionary based on the five-tone theory of traditional Chinese medicine, perform traditional Chinese medicine semantic analysis on the emotion classification label to obtain the five-organ attribution information corresponding to the emotional state, and obtain the emotion-five-organ mapping relationship table;
[0210] S10.2: Perform fuzzy reasoning processing on the five-organ attribution information in the emotion-five-organ mapping relationship table, combine the current physiological feature parameters and historical state data of the infant, and correct the mapping deviation caused by individual differences to obtain the corrected five-organ state confidence distribution;
[0211] S10.3: Based on the five-organ state confidence distribution, use a weighted fuzzy logic reasoning method to model the nonlinear relationship between the five-tone features and the five-organ state to generate a five-tone feature-driven five-organ function state evaluation vector;
[0212] S10.4: Perform semantic fusion on the five-organ function state evaluation vector and the emotional classification label to perform joint modeling of the emotional-physiological state to generate an infant comprehensive state evaluation semantic label as the final output state semantic description of the system;
[0213] S10.5: Perform structured packaging and visual mapping on the infant comprehensive state evaluation semantic label to generate a multi-dimensional evaluation report containing five-tone features, emotional state, five-organ attribution and state confidence, and output it as the final feedback result of the system output interface.
[0214] The application further provides a baby state recognition system based on the TCM five-tone monitoring analysis.
[0215] So far, the technical solutions of the application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the application, and the technical solutions after the changes or replacements will all fall within the protection scope of the application.
[0216] The above description is only the preferred embodiments of the application and is not used to limit the application; for those skilled in the art, the application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and rules of the application shall be included in the protection scope of the application.
Claims
1. A method for recognizing the state of infants and young children based on the monitoring and analysis of the five tones in Traditional Chinese Medicine, characterized in that, Includes the following steps: S1: Collect audio signals of infants' cries, physiological signals, and behavioral video data as multimodal input samples; S2: Perform audio signal preprocessing on the crying audio signal and extract the Mel frequency cepstral coefficients as audio feature vectors; S3: Perform time-frequency analysis and feature selection on the physiological signals to extract key physiological feature parameters; S4: Perform keyframe extraction and action recognition processing on the behavioral video data, and extract facial expression and limb movement features of infants and young children based on convolutional neural networks; S5: Based on the theory of five tones in traditional Chinese medicine, establish a mapping table between the characteristics of the five tones and the state of the five internal organs, and perform preliminary matching between the extracted audio features and the five tones classification; S6: Construct a multimodal feature fusion model, and concatenate the audio feature vector, the key physiological feature parameters, and the video action features to form a fused feature vector; S7: Calculate the dynamic weight coefficients of each modality feature in emotion recognition based on the matching degree between historical recognition results and current features; S8: Input the weighted fused feature vector into a lightweight nonlinear classifier to perform the recognition and reasoning of infants' emotional states and output emotion classification labels; S9: Determine whether the current emotion recognition result is consistent with the historical trend. If not, activate the feedback adjustment mechanism to update the dynamic weight coefficients and classifier parameters. S10: Map the identified emotional state labels to the corresponding information related to the five internal organs, and output the comprehensive physiological and emotional state assessment results of infants and young children.
2. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 1, characterized in that, Step S1 specifically includes: The audio signal of an infant's cry was collected using a high-sensitivity microphone to obtain raw audio data containing the five tones. Based on medical-grade physiological monitoring equipment, electrocardiogram, skin conductance response and respiratory waveform signals of infants and young children are collected to obtain physiological data reflecting the state of the autonomic nervous system; High-definition cameras are used to record videos of infants' facial and limb movements to obtain sequences of behavioral and facial expression changes. Perform timestamp synchronization processing on the collected multimodal raw data and output the synchronized multimodal raw data; The synchronized multimodal raw data is encapsulated into a structured sample set.
3. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 2, characterized in that, In step S1, physiological signal acquisition includes using a three-lead medical ECG module to synchronously acquire ECG signals at a sampling rate of 1000Hz and a resolution of 24bit. Noise is eliminated by a low-pass-notch filter combination filter. Skin conductance at 100Hz and respiratory waveform at 200Hz are synchronously acquired under constant current excitation. All signals are synchronized using a unified clock and timestamp.
4. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 1, characterized in that, Step S2 specifically includes: The collected audio signals of infants' cries were pre-emphasized, and a first-order differential filter was used to enhance the high-frequency components of the original audio waveform to obtain the pre-emphasized audio waveform data. A windowing operation is performed on the pre-emphasized audio waveform data, and the signal is segmented into frames using a Hamming window function to obtain multiple windowed audio frame signals. Perform Fast Fourier Transform on each windowed audio frame signal to obtain the corresponding spectral amplitude distribution characteristics; The audio signal represented in the frequency domain is subjected to nonlinear frequency mapping using a Mel filter bank, which converts the linear frequency scale into a Mel scale to obtain the energy response value of each filter channel. The energy response values of each Mel filter channel are processed by discrete cosine transform to extract low-order cepstral coefficients and obtain the Mel frequency cepstral coefficient eigenvector.
5. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 1, characterized in that, Step S3 specifically includes: Wavelet denoising was performed on the collected electrocardiogram (ECG) signals of infants and young children to obtain the denoised ECG waveform sequence. Based on the denoised ECG waveform sequence, R-wave peak detection is performed to obtain the RR interval sequence, which serves as the basic input data for calculating heart rate variability. Time-domain statistical analysis was performed on the RR interval sequence to extract time-domain characteristic parameters of heart rate variability reflecting the activity of the autonomic nervous system; A sliding window Fourier transform was performed on the skin conductance signals of infants and young children to extract the peak amplitude and frequency of the skin conductance response, which were used as skin conductance characteristic parameters reflecting the degree of emotional arousal. An adaptive filtering method is used to extract the respiratory cycle waveform from the respiratory signal, the respiratory frequency is calculated based on the peak detection algorithm, and the rate of change of respiratory depth per unit time is statistically analyzed to obtain respiratory characteristic parameters that reflect emotional state.
6. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 5, characterized in that, In step S3, after detecting the R wave using the Pan-Tompkins algorithm, the heart rate variability parameters are calculated using the SDNN and RMSSD of the RR interval sequence. SDNN is the standard deviation of the RR interval, and RMSSD is the root mean square of the adjacent differences. Both are Z-score normalized or min-max normalized. Skin conductance characteristics are extracted using sliding Fourier transform to obtain peak amplitude and response frequency. Respiratory features were obtained by adaptive filtering and peak detection to obtain respiratory cycle and depth variation rate.
7. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 1, characterized in that, Step S4 specifically includes: Keyframe extraction processing is performed on the collected infant behavior video data. Based on the calculation of inter-frame differences and adaptive threshold determination, video keyframes with significant action changes are extracted. Facial region detection and localization processing is performed on the extracted keyframe image sequence. A multi-task learning model based on cascaded convolutional neural networks is used to locate the facial region of infants and young children and detect key points to obtain the basic analysis region of facial expression changes. The facial image region after localization is processed for expression feature extraction. Based on a lightweight convolutional neural network model, local texture and global morphological features are extracted from the image to obtain facial expression feature vectors for infant emotion recognition. The keyframe image sequence is processed for limb movement region segmentation. A video action recognition model based on spatiotemporal convolutional network is used to perform pixel-level segmentation of the infant's limb movement region and extract the spatial distribution features of limb movements. The segmented limb movement regions are processed by motion feature encoding. Based on a three-dimensional convolutional neural network, the temporal continuity and spatial motion pattern features of the movements are extracted to generate limb movement feature vectors.
8. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 7, characterized in that, In step S4, key frames with motion changes are extracted by using the inter-frame difference histogram and adaptive threshold. A three-level cascaded convolutional neural network is used to locate the facial region. The confidence threshold is set to 0.85, the overlap rate threshold of the NMS algorithm is 0.3, and five facial key points are output. Pose normalization is performed by affine transformation. Then, a lightweight convolutional network is used to extract local texture and global morphological features. A weighted fusion method is used to generate expression and motion feature vectors.
9. The method for infant and toddler state recognition based on traditional Chinese medicine five-tone monitoring and analysis according to claim 1, characterized in that, Step S5 specifically includes: The extracted audio feature vectors are normalized to obtain a standardized audio feature matrix; Based on the predefined five-tone pitch benchmark values, the dynamic time warping algorithm is used to perform pattern matching on the fundamental frequency sequence in the standardized audio feature matrix to obtain the five-tone classification results. The five-tone classification results are matched using the traditional Chinese medicine five-tone-five-organ mapping rules. Based on the correspondence table of five-tone-five-organs in traditional Chinese medicine theory, the five organ systems corresponding to the current audio signal are determined, and a five organ-related label sequence is generated. The five organ association label sequences are assigned fuzzy membership values. Fuzzy set theory is introduced to model the non-unique mapping relationship between the five tones and the five organs, and the five tones-five organs fuzzy association matrix is output. The five-tone-five-organ fuzzy correlation matrix is fused with the infant's current physiological state information, and the matching probability between the five-tone features and the state of the five organs is updated based on the Bayesian inference mechanism.
10. An infant and toddler state recognition system based on the monitoring and analysis of the five tones in Traditional Chinese Medicine, characterized in that: The infant state identification method based on the five-tone monitoring and analysis of traditional Chinese medicine, as described in any one of claims 1-9, is used to identify the infant state.
Citation Information
Cited By
Children rhinitis atomization treatment control system and method based on behavior guidance and intelligent monitoring
CN122141079A
Sound recognition method based on fttr network, electronic device, medium and product
CN122417083A