An AI-based multi-modal emotion perception and adaptive correction-based intelligent nasal washing robot for children and a control method thereof

CN122681702APending Publication Date: 2026-09-04李佳琪
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610711639.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

[0003]现有洗鼻设备主要依赖单一的语音识别或简单的压力阈值判断来感知儿童情绪状态,缺乏多模态生理信号的融合分析能力

Benefits of technology

1.多模态协同感知的准确性提升

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122681702A_ABST
    Figure CN122681702A_ABST
Patent Text Reader

Abstract

The application belongs to the field of medical devices, and particularly relates to a child intelligent nasal washing robot based on AI multi-modal emotion perception and adaptive correction and a control method thereof, aiming to solve the problems of inaccurate emotion perception, difficult guarantee of operation standardization and difficult prevention of safety risks in the process of nasal washing of children. The nasal washing robot obtains multi-modal perception data through a sound collection component, a flexible pressure sensor array, a physiological signal sensor, a flow monitoring component and an environment perception component, adopts a three-stage progressive fusion architecture to identify the emotional state of children and evaluate the operation standardization, dynamically adjusts the flushing parameters according to the emotional state and the standardization level through an adaptive proportional integral differential control algorithm, realizes predictive protection of risks such as coughing and resistance escalation in combination with a trajectory prediction network, and continuously optimizes the personalized control strategy through an incremental learning controller. The system also integrates an emotional interaction guidance module, and outputs dynamic expressions, guidance dialogues and tactile feedback according to the real-time emotional state and the task stage. The application can effectively improve the accuracy and robustness of multi-modal emotion perception, realize predictive safety protection and personalized precise nursing, significantly reduce the risk of damage to the nasal mucosa and coughing of children, and improve the nasal washing compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical devices, specifically relating to a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction, and its control method. Background Technology

[0002] With the increasing incidence of respiratory diseases in children, nasal irrigation, as a non-pharmacological physical therapy, plays a crucial role in clearing nasal secretions and allergens. Home-based and intelligent medical care devices are gradually becoming an important support for improving children's health management, and the market demand for intelligent nasal irrigation products that can accurately sense children's conditions and provide personalized care services is increasingly urgent. The rapid development of AI multimodal collaborative sensing technology, adaptive predictive control algorithms, and emotion computing methods provides a solid technological foundation for the intelligent upgrading of children's intelligent nasal irrigation devices.

[0003] Existing nasal irrigation devices primarily rely on single-modal voice recognition or simple pressure threshold judgment to perceive children's emotional state, lacking the ability to fuse and analyze multimodal physiological signals. When children are in situations of silent resistance or background noise interference, the accuracy of single-modal recognition schemes drops significantly, failing to accurately assess the child's level of cooperation and psychological state. The safety protection mechanisms of existing devices mainly adopt a reactive approach, stopping the pump after a dangerous event occurs via threshold triggering. This passive protection strategy has a significant response delay and cannot provide preventative protection for scenarios requiring early intervention, such as coughing or tubing blockage, increasing the risk of nasal mucosal damage and coughing in children. Traditional nasal irrigation devices use fixed PID parameters or simple segmented control strategies, unable to adaptively adjust according to individual child differences and real-time status, making precise parameter optimization difficult and resulting in inconsistent irrigation comfort and compliance. The interactive feedback of existing devices mostly uses pre-recorded scripts and fixed animation sequences, unable to dynamically adjust the interaction strategy based on the child's real-time emotional state and task progress, lacking differentiated guidance mechanisms for different emotional states, making effective intervention difficult before a child's emotional state deteriorates. Some smart nasal irrigation devices rely excessively on cloud computing and network connectivity for safety judgment. In scenarios with network outages or network latency, the effectiveness of the safety protection mechanism cannot be guaranteed. The lack of a local safety fallback mechanism independent of the cloud affects the reliability of the device in various usage environments. Summary of the Invention

[0004] The purpose of this invention is to provide a children's intelligent nasal irrigation robot and its control method based on AI multimodal emotion perception and adaptive correction, which can effectively solve the problems in the background art.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A children's intelligent nasal irrigation robot and its control method based on AI multimodal emotion perception and adaptive correction include the following specific steps: Step 1: Multi-source perception data acquisition: Through the sound acquisition component, flexible pressure sensor array, physiological signal sensor, flow monitoring component, and environmental perception component installed on the nasal irrigation robot, the robot acquires in real time the child's voice audio signal, the temporal feature data of the handle gripping force, the physiological parameter data reflecting emotional arousal, the real-time flow rate data of the irrigation fluid, the device posture and motion data, and the external environment data, and transmits the acquired multimodal perception data to the embedded central processing unit (CPU); Step 2: Multimodal emotion recognition and classification: The embedded CPU calls the built-in emotion perception engine to preprocess the voice audio signal, extracts the Mel-frequency feature vector and combines it with a lightweight convolutional neural network for acoustic feature extraction, and combines the heart rate variability index and skin conductance response characteristics fed back by the physiological signal sensor to perform multimodal feature fusion operation based on a three-stage progressive fusion mechanism, identify and output the classification result of the child's current emotional state and the scores of valence, arousal, and dominance; Step 3: Operational standardization assessment: The embedded CPU calls the built-in operation correction... The engine compares real-time flow rate and grip pressure data with preset safety thresholds corresponding to specific age groups and nasal physiological structures. It monitors the spatial tilt angle of the nasal irrigation robot using a six-axis accelerometer and comprehensively evaluates instantaneous pressure peaks and flow rate stability, calculating and outputting the current nasal irrigation operation's compliance level. Step 4: Adaptive Rinsing Control. Based on the emotional state classification results and compliance level, the embedded central processing unit determines the target rinsing power using a preset adaptive proportional-integral-derivative control algorithm and dynamically adjusts the pulse width modulation drive signal parameters of the micro-adjustable rinsing pump to achieve closed-loop regulation of the rinsing fluid output pressure and flow rate. The gain parameter of the control algorithm is nonlinearly adaptively adjusted according to the real-time emotional state and compliance level. Step 5: Emotional Interactive Guidance. Based on the nasal irrigation task's phased progress and emotional state classification results, the embedded central processing unit schedules the LED array screen, audio output unit, and vibration feedback component in the interaction module to output corresponding dynamic facial expression sequences, guiding voice prompts, and tactile stimulation signals, guiding children to execute and complete the nasal irrigation task chain consisting of preparation, rinsing, cleaning, and completion stages.

[0006] Preferably, the multi-source sensing system in step 1 includes an audio sensing submodule, a physiological sensing submodule, a motion sensing submodule, a mechanical sensing submodule, a liquid sensing submodule, and an environmental sensing submodule. The audio sensing submodule uses a digital microphone, with the sampling frequency and quantization bit depth set to preset values. It performs real-time preprocessing of children's speech using an adaptive spectral subtraction noise reduction algorithm. The preprocessing process includes bandpass filtering to filter out noise below a preset low-frequency threshold and above a preset high-frequency threshold, and spectral subtraction to eliminate mechanical vibration noise generated by the nasal irrigation pump.

[0007] Preferably, the physiological sensing submodule integrates a photoplethysmography (PPG) sensor and a skin conductance sensor in the grip area. The PPG sensor emits infrared light of a specific wavelength and receives reflected light to extract real-time heart rate data of the child during gripping, and calculates time-domain indicators of heart rate variability, including the root mean square of the difference between adjacent normal heartbeats. The skin conductance sensor is used to monitor changes in sweat gland secretion caused by emotional tension, reflecting the degree of emotional arousal. The introduction of physiological signals solves the problem of single audio modality recognition failure in silent resistance scenarios.

[0008] Preferably, the motion sensing submodule employs a 6-axis inertial measurement unit, including a 3-axis accelerometer and a 3-axis gyroscope, for detecting the device's attitude angle, shaking amplitude, and nozzle direction. The sampling frequency is set to a preset value. By fusing the accelerometer and gyroscope data through Kalman filtering, a stable attitude angle estimate is output to determine the stability of the device held by a child, whether the nozzle contact angle is within a preset safe angle range, and whether abnormal motion patterns such as violent shaking occur.

[0009] Preferably, the mechanical sensing submodule arranges a flexible pressure sensor array in the handle area, with a sampling frequency of a preset value, distributed on the circumferential surface of the handle, to accurately capture muscle tremors and sudden changes in grip force caused by children's tension. The collected analog signals are converted into digital sequences by a multi-channel analog-to-digital converter, and the processor performs a moving average filter on the digital sequence to eliminate random interference caused by slight hand tremors and extract pressure envelope features reflecting emotional tension. The pipeline pressure sensor is used to monitor the instantaneous value of the flushing pressure, and triggers safety protection when the pressure exceeds a preset threshold.

[0010] Preferably, the liquid sensing submodule includes a miniature flow sensor, a liquid level sensor, and a temperature sensor. The flow sensor uses a miniature turbine flow meter or an ultrasonic flow meter to monitor the real-time flow rate of the flushing liquid with a preset response speed and a measurement accuracy that reaches a preset accuracy level. The liquid level sensor is used to detect the remaining amount of flushing liquid in the storage bottle. The temperature sensor uses an infrared temperature sensor to monitor the temperature of the flushing liquid in real time. When the temperature exceeds the preset temperature range, the flushing pump is prohibited from being turned on.

[0011] Preferably, the three-stage progressive fusion architecture in step 2 includes a first stage of modal independent encoding, a second stage of cross-modal interactive fusion, and a third stage of multi-task joint output. In the first stage, the audio encoding branch adopts the Mel spectrum feature extraction method combined with a lightweight convolutional neural network encoder, while the mechanics and motion encoding branches adopt an architecture combining a temporal convolutional network and a bidirectional long short-term memory network. The temporal convolutional network adopts a dilated causal convolution structure to support temporal dependencies covering a preset duration, and the bidirectional long short-term memory network is used to model bidirectional temporal context relationships. The physiological encoding branch extracts features from heart rate signals and electrodermal signals, including temporal and frequency domain features.

[0012] Preferably, the second stage employs a multi-level cross-modal attention mechanism to establish the association between different modalities. Sound-to-sensor attention is used to establish the association between sound events and device states, and sound-to-physiology attention is used to establish the association between sound events and physiological states. Audio features, sensor features, and physiological features are respectively used as queries, keys, and values ​​to be input into the cross-modal attention layer to calculate the interaction weights between modalities.

[0013] Preferably, the third-stage model output adopts a multi-task structure, simultaneously outputting three types of results: child emotional state classification including four categories: calm, tense, resistant, and crying; continuous scores of valence, arousal, and dominance; valence representing the degree of positive and negative emotions, with values ​​within a preset range; arousal representing the degree of physiological activation, with values ​​within a preset range; and dominance representing the ratio of willingness to cooperate to degree of resistance, with values ​​within a preset range; operation standardization level judgment including four levels: standard, slight deviation, significant deviation, and dangerous; and risk event type identification including multi-label classification such as choking, violent shaking, tubing blockage, and pressure exceeding limits; and simultaneously outputting the predicted probability and priority ranking of each type of risk.

[0014] Preferably, it also includes a predictive safety assessment step, which encodes a preset number of historical frames through a trajectory prediction network to predict the system state trajectory within a preset time period in the future. The input data includes pressure sequence, flow sequence, attitude angle sequence, heart rate sequence and grip pressure envelope. Based on the predicted trajectory, the probability of coughing risk, the probability of resistance escalation and the probability of pipeline blockage are calculated. When the predicted probability exceeds a preset threshold, a calming and deceleration action or a reverse suction action is triggered in advance to achieve preventive safety protection.

[0015] Preferably, in step 4, the adaptive proportional-integral-derivative controller adopts a nonlinear adaptive structure. The control gain coefficient is adaptively adjusted according to the child's emotional state, current stage, individual characteristics, and pipeline status. The nonlinear compensation term is dynamically calculated based on the real-time identification results. When the child's emotional state is tense, the proportional gain is reduced to improve system stability and extend system response time. When the emotional state is resistant, the integral gain is further reduced to avoid overshoot caused by the accumulation of integral terms. When the standardization level is slightly deviated, differential feedforward compensation is introduced to predict the pressure change trend. When there is a risk of pipeline blockage, the differential gain is reduced to avoid high-frequency oscillation.

[0016] Preferably, it also includes an incremental learning controller, which updates the personalized configuration of proportional, integral, and differential parameters based on the control sequence of the current task after each nasal irrigation task. The update algorithm adopts elastic weight consolidation technology to prevent the learning of new tasks from overwriting key parameters. The target flow rate is determined by the child's profile and current status. When the child is young, has a high number of historical resistances, or has a high number of recent abnormalities, the system automatically adopts a lower base flow rate. The adjustment of the base flow rate adopts a gradual strategy to avoid physical stimulation of the nasal mucosa caused by sudden pressure changes.

[0017] Preferably, the dynamic script generation system in step 5 adopts a template-based architecture, which includes three parts: stage templates, emotion modifiers, and personalized variables. The stage templates define the basic scripts for each task stage. The emotion modifiers adjust the tone and length of the scripts according to the real-time recognized emotional state. In a calm state, normal speaking speed and standard volume are used. In a tense state, the speaking speed is slowed down, the volume is lowered, and encouraging statements are added. In a resistant state, very short statements and a gentle tone are used. In a crying state, the system switches to a silent mode and uses an emoji screen and vibration feedback instead of voice. Personalized variables include the child's name, historical performance records, and reward badge status.

[0018] Preferably, it also includes a multi-dimensional feedback system, including four dimensions: visual feedback, auditory feedback, tactile feedback, and olfactory feedback. Visual feedback uses an LED expression screen to display dynamic expression sequences, with the expression content linked to the current task stage and emotional state. Auditory feedback uses preset voice prompts and sound effects. Tactile feedback uses a linear vibration motor to provide slight vibration stimulation as positive reinforcement when the child correctly completes key steps. Olfactory feedback is an optional module that releases a small amount of fragrance before rinsing begins.

[0019] Preferably, it also includes a local safety protection system, which runs multiple local safety protection rules independently of the cloud and network. The safety protection rules include two types: threshold-triggered and trend-predictive. Threshold-triggered rules include pressure limit protection, nozzle detachment protection, abnormal posture protection, cough trigger protection, severe shaking protection, abnormal temperature protection, insufficient liquid level protection, and abnormal flow protection. Trend-predictive rules include cough prediction protection, resistance escalation protection, pipeline blockage warning, and battery abnormality protection. The safety rules have a higher priority than the output of the finite state machine and the proportional-integral-derivative controller.

[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. Improved accuracy of multimodal collaborative sensing By introducing physiological signals, including photoplethysmography pulse wave heart rate and skin conductance response, and environmental perception, including background noise and ambient light, combined with a three-stage progressive fusion architecture, this invention can effectively solve the problem of recognition failure of a single audio modality in silent resistance and noise interference scenarios. The introduction of physiological signals fills the perceptual gap when children cannot express emotions through language. The environmental perception module can automatically identify the external noise environment and adjust the multimodal fusion weights. The multimodal information complements redundancy, significantly improving the accuracy and robustness of emotional state recognition and risk event detection.

[0021] 2. Reduced risks associated with predictive security measures The upgrade from passive response to proactive prevention in safety strategies enables the prediction and intervention of risk events such as choking and escalating resistance before they occur. The trajectory prediction network predicts future trajectory based on a preset amount of multimodal time-series data and calculates the predicted probability of various risks. When the predicted probability exceeds a threshold, protective actions are triggered in advance. Compared with the traditional post-event pump shutdown mechanism, the predictive safety module significantly reduces the risk of nasal mucosal damage and choking in children, providing a safer care environment for children.

[0022] 3. Enhanced personalization of adaptive control Through incremental learning controllers and a personalized parameter optimization mechanism, the device can continuously learn the stress response characteristics and psychological tolerance thresholds of specific children, automatically adjust proportional-integral-derivative control parameters and target flow settings. Incremental learning uses elastic weight consolidation technology to prevent new tasks from overriding key parameters. After a preset number of usage cycles, the controller parameters are fully adapted to the individual characteristics of specific children. The gradual adjustment strategy of basic flow avoids physical stimulation of the nasal mucosa by sudden pressure changes, achieving truly personalized and precise care.

[0023] 4. Enhanced adherence through emotional interaction The dynamic script generation system and multi-dimensional feedback mechanism can dynamically adjust the interaction strategy according to the child's real-time emotional state and task progress. The emotion modifier adjusts the tone, speed and volume of the script according to the emotional state. The gamified task chain design transforms the nasal rinsing process into a task story of taking care of a baby elephant. The virtual growth system establishes a long-term incentive mechanism through virtual badges and unlocking new skills, transforming short-term task compliance into long-term healthy behavior habits, significantly improving children's acceptance of nasal rinsing and parents' care efficiency.

[0024] 5. Reliability assurance of the local security system The local security protection system, which operates independently of the cloud and network, ensures that devices can perform reliable security protection actions in any usage scenario, including network outages, signal delays, and device offline events. Multiple local security rules cover both threshold-triggered and trend-predictive types. The hardware self-test logic eliminates potential sensor malfunctions before startup, and the priority setting of security rules ensures that predictive protection can cover the decision-making of the control layer, further improving the overall reliability of the system. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the overall technical solution architecture of a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction according to an embodiment of this application. Figure 2 This is a schematic diagram of the core principle framework of multimodal emotion recognition based on a three-stage progressive fusion architecture according to an embodiment of this application; Figure 3 This is a schematic diagram of the multi-source sensory data acquisition process framework of a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction according to an embodiment of this application; Figure 4 This is a schematic diagram of the adaptive proportional-integral-derivative control process framework based on emotional state and normativity level according to an embodiment of this application. Figure 5 This is a schematic diagram of the emotional interaction guidance process framework of a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction according to an embodiment of this application. Figure 6 This is a schematic diagram illustrating the predictive security assessment principle framework based on trajectory prediction networks according to embodiments of this application; Figure 7 This is a schematic diagram illustrating the local safety protection system and multi-level cloud interaction relationship and data flow of a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction, according to an embodiment of this application. Detailed Implementation Example 1

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0027] This invention provides a children's intelligent nasal irrigation robot and its control method based on AI multimodal emotion perception and adaptive correction. The method constructs a multi-source perception system using 12 types of sensors integrated into the robot body. This system collects audio, physiological, motion, mechanical, fluid parameter, and environmental parameter signals from the child during the nasal irrigation process and transmits them to an embedded central processing unit for comprehensive analysis and processing. Based on a three-stage progressive fusion architecture, it achieves accurate identification and classification of the child's emotional state. Combined with real-time evaluation of operational compliance and adaptive irrigation control, as well as emotional interactive guidance based on a dynamic speech generation system, it forms a complete perception-recognition-control-feedback closed-loop system.

[0028] The specific implementation process of acquiring multi-source sensing data in step 1 is as follows.

[0029] The audio perception submodule uses a Knowles series MEMS digital microphone as the sound acquisition component. This microphone has a sensitivity of -26dBFS, a signal-to-noise ratio of 65dB, a sampling frequency of 16kHz, and a quantization bit depth of 16 bits. The sound acquisition component has a built-in automatic gain control circuit that dynamically adjusts the gain range of the input signal to ensure effective audio data acquisition under different sound pressure levels. The preprocessing stage uses an adaptive spectral subtraction noise reduction algorithm. First, a bandpass filter is used to filter out low-frequency noise below 100Hz and high-frequency noise above 8000Hz. The transition band width of the bandpass filter is set to 20Hz, and the stopband attenuation reaches 40dB. Then, spectral subtraction is used to eliminate the mechanical vibration noise generated during the operation of the nasal irrigator (mainly concentrated between 200Hz and 500Hz). Specifically, this is achieved by estimating the noise power spectrum in the frequency domain and subtracting it from the noisy speech power spectrum to obtain the enhanced speech spectrum. The preprocessed audio signal is used for subsequent emotional acoustic feature extraction and voice command recognition.

[0030] The physiological sensing submodule integrates a PPG (photoplethysmography) pulse wave sensor and a skin conductance sensor in the grip area. The PPG sensor operates at a wavelength of 940nm infrared light, emitting infrared light onto the child's skin via a transmitting unit. After reflection, the receiving unit captures the light intensity change signal, with a sampling frequency of 100Hz and a sampling precision of 12 bits. The signal processing unit sequentially performs low-pass filtering (cutoff frequency 5Hz) and band-pass filtering (0.5Hz to 10Hz) on the raw PPG signal, extracts the pulse wave peak sequence, and calculates the heart rate value and heart rate variability index. The calculation of the heart rate variability time-domain index includes the root mean square (RMSSD) of the difference between adjacent normal heartbeats and the standard deviation (SDNN) of heart rate variability. The calculation window length for both RMSSD and SDNN is 5 minutes. The skin conductance sensor uses a dual-electrode configuration with an electrode spacing of 10mm, measuring changes in skin conductance at a sampling frequency of 50Hz. After low-pass filtering (cutoff frequency 2Hz), the mean EDA value, number of EDA peaks, and EDA response amplitude of the EDA signal are extracted. The threshold for EDA peak detection is set to the baseline value plus 0.05μS. The introduction of the physiological perception submodule enables the system to determine the level of emotional tension of children without them making a sound by observing changes in heart rate and EDA, effectively solving the problem of emotion recognition in silent resistance scenarios.

[0031] The motion sensing submodule employs a BMI270 6-axis inertial measurement unit, integrating a 3-axis accelerometer and a 3-axis gyroscope. The accelerometer range is set to ±16g with a resolution of 0.061mg / LSB, and the gyroscope range is set to ±2000dps with a resolution of 0.061dps / LSB. The sampling frequency is uniformly set to 50Hz. Data from the inertial measurement unit is transmitted to the embedded central processing unit via an I2C interface at a rate of 400kHz. The processor uses a Kalman filter algorithm to fuse the accelerometer and gyroscope data. The Kalman filter state vector includes attitude angles in quaternion form and gyroscope zero-bias estimates. Process noise parameters are adaptively adjusted based on static and dynamic states. The filtered attitude angle output includes pitch, roll, and yaw angles. The attitude estimation update frequency is 50Hz, and the static angle accuracy is better than 0.5 degrees. The motion sensing submodule simultaneously calculates the acceleration envelope area for sway detection. The calculation formula is the absolute value of the integral of the acceleration signal over time in the time domain. When the acceleration envelope area exceeds a preset threshold, it is determined to be a violent sway.

[0032] The mechanical sensing submodule consists of a flexible pressure sensor array and a pipeline pressure sensor. The flexible pressure sensor array contains 10 sensing points distributed on the circumferential surface of the handle in a uniform ring shape, with adjacent sensing points spaced 36 degrees apart. The flexible pressure sensor employs the piezoresistive principle, with a measurement range of 0 to 500g, linearity better than 0.5%FS, and a sampling frequency of 50Hz. The analog output signal of the sensor array is converted into a digital sequence by a multi-channel 12-bit analog-to-digital converter (ADC). The ADC has a sampling rate of 200kHz and a conversion accuracy of 0.1g. The processor performs a moving average filter on the digital sequences from the 10 sensing channels, with a sliding window length set to 5 sampling points (corresponding to a 100ms time window). The filtered data is used to extract pressure envelope features. The extracted pressure envelope features include the mean grip pressure, the standard deviation of grip pressure, and the rate of change of grip pressure. The rate of change of grip pressure is calculated by dividing the difference in the mean pressure at adjacent sampling times by the sampling interval. The pipeline pressure sensor is a MEMS type pressure sensor with a range of 0 to 200 kPa, an accuracy of ±1 kPa, and a response time of 10 ms. It is installed at the pipeline connection at the outlet of the peristaltic pump to monitor the instantaneous value of the flushing pressure in real time.

[0033] The liquid sensing submodule includes a miniature flow sensor, a level sensor, and a temperature sensor. The flow sensor uses a miniature turbine flow meter with a range of 0 to 500 mL / min, a measurement accuracy of ±2%FS, and a response time of 50 ms. It calculates the real-time flow velocity by detecting the pulse signal generated by the turbine rotation; the pulse signal frequency is proportional to the flow velocity, and the proportionality coefficient is determined by the flow meter calibration. The level sensor is a capacitive level sensor installed on the inner wall of the storage bottle. Its measurement principle is that changes in liquid level cause changes in capacitance; the liquid level is calculated by measuring the capacitance value, and the level measurement accuracy is ±2 mm. The temperature sensor uses an MLX90614 infrared temperature sensor with a range of -20 to 100 degrees Celsius, a measurement accuracy of ±0.5 degrees Celsius, a field of view of 90 degrees, and a target emissivity set to 0.95. It calculates the temperature value by receiving the infrared radiation energy from the surface of the flushing liquid. All sensor data from the liquid sensing submodule are transmitted to the central processing unit via an SPI interface with a clock frequency set to 10 MHz.

[0034] The environmental perception submodule includes an ambient light sensor and an ambient noise sensor. The ambient light sensor uses a photoresistor with a measurement range of 0.1 to 10000 lux and a response time of 100 ms. It provides input for adjusting the display brightness by detecting ambient light intensity. The ambient noise sensor reuses the digital microphone from the audio perception submodule. It calculates the ambient noise intensity through power spectral analysis. When the noise intensity exceeds a preset threshold (set to 60 dB SPL), the system automatically reduces the weighting coefficient of speech features in the multimodal fusion operation. The ambient noise intensity is calculated using an A-weighting algorithm.

[0035] The multimodal sensing data collected by the above six types of sensors are synchronously transmitted to the embedded central processing unit (CPU) through their respective communication interfaces. Audio data is transmitted in real time via the I2S interface at a sampling rate of 16kHz, with each frame containing 512 bytes of data. Physiological, motion, mechanical, and fluid data are transmitted via the I2C bus at a polling frequency of 50Hz to 100Hz. Environmental data is transmitted via the ADC channel at a sampling frequency of 10Hz. The embedded CPU uses an ESP32-S3 Ultra dual-core heterogeneous chip. Core 1, with a main frequency of 240MHz, is responsible for real-time sensor data acquisition and pump control PWM output, while Core 2, with a main frequency of 200MHz, is responsible for multimodal feature extraction and lightweight neural network inference. The processor maintains a circular data buffer to store the most recent 30 frames of time-series data. The buffer depth is set to 30 frames, and each frame contains feature vectors of various sensor data. Data updates use a timestamp synchronization mechanism to ensure the temporal consistency of each modality's data.

[0036] The specific implementation process of multimodal emotion recognition and classification in step 2 is as follows.

[0037] The embedded central processing unit calls the built-in emotion perception engine to perform multimodal emotion recognition and classification tasks. The software architecture of the emotion perception engine consists of an audio feature extraction module, a temporal feature extraction module, a physiological feature extraction module, and a three-stage fusion network module.

[0038] The audio feature extraction module employs the log-Mel spectral feature extraction method. The specific processing flow includes three stages: framing, windowing, FFT transformation, and Mel filter bank filtering. Framing divides the continuous audio signal into short frames of 25 milliseconds, with a frame shift of 10 milliseconds and a frame overlap rate of 60%. A Hanning window is used for windowing to reduce spectral leakage. Each windowed frame undergoes a 512-point FFT transformation to obtain the spectral amplitude, with an FFT output frequency resolution of approximately 31.25 Hz. The Mel filter bank contains 40 bandpass filters, with their center frequencies uniformly distributed according to the Mel scale. The conversion relationship between Mel frequency and linear frequency is Mel(f) = 2595 × log10(1 + f / 700). The filter bank outputs a 40-dimensional log-Mel spectral feature vector. The audio feature extraction module outputs data at 100 frames per second, with each frame corresponding to a 10-millisecond audio signal.

[0039] The temporal feature extraction module employs an architecture combining a temporal convolutional network and a bidirectional long short-term memory network. The input consists of temporal data collected by the mechanical perception and motion perception submodules, including grip pressure sequences, attitude angle sequences, and acceleration sequences. All sequences are uniformly resampled at 10Hz. The input data first undergoes channel fusion and preliminary feature extraction through a 1D convolutional layer with a kernel size of 3 and 64 output channels. Then, the data enters the TCN residual convolutional block. The TCN contains four residual blocks, each containing two layers of dilated causal convolutions with dilation coefficients of 1, 2, 4, and 8, respectively. The kernel size is uniformly 3, and the expanded effective receptive field covers approximately a 3-second historical window. The output of the TCN is fused with the input of a BiLSTM through residual connections. The BiLSTM contains two layers, each with 128 hidden units, and the bidirectional concatenation outputs a 256-dimensional temporal feature vector.

[0040] The physiological feature extraction module extracts time-domain and frequency-domain features from PPG heart rate and electrodermal (EDA) signals. Time-domain features include the mean heart rate (HR_mean, arithmetic mean of heart rate over the past 30 seconds), root mean square heart rate variability (RMSSD), and standard deviation of heart rate variability (SDNN). Frequency-domain features are obtained through power spectral density analysis of the heart rate sequence, including low-frequency power (LF, frequency range 0.04 to 0.15 Hz) and high-frequency power (HF, frequency range 0.15 to 0.4 Hz). The power ratio of LF to HF serves as an indicator of sympathetic-parasympathetic balance. EDA signal features include the mean EDA (arithmetic mean of EDA conductance over the past 10 seconds), the number of EDA peaks (the number of times the EDA signal exceeds a threshold per unit time), and the EDA response amplitude (the difference between the peak and the baseline). The physiological coding network employs a multilayer perceptron structure, containing two hidden layers with 32 and 64 neurons respectively. The ReLU activation function is used, and the output is a 64-dimensional physiological feature vector.

[0041] A three-stage fusion network achieves efficient fusion of audio, temporal, and physiological features. The first stage is modality-independent encoding. The audio encoder adopts a Mini-Conformer lightweight Transformer encoder structure, containing four convolutional downsampling layers (gradually downsampling the temporal dimension from 100 frames to 25 frames), six Conformer modules, and two fully connected layers. The Conformer modules combine the local feature extraction capabilities of convolutional neural networks with the long-range dependency modeling capabilities of self-attention mechanisms. The audio encoder outputs a 64-dimensional feature vector. The temporal encoder receives the 256-dimensional output of TCN+BiLSTM, maintaining the 256-dimensional output after passing through a linear projection layer. The physiological encoder receives the 64-dimensional output of a multilayer perceptron, maintaining the 64-dimensional output after passing through a linear projection layer.

[0042] The second stage is cross-modal interaction fusion, which employs a multi-level cross-modal attention mechanism to establish connections between different modalities. Audio-to-Sensor attention calculates attention weights using audio features as the query and temporal features as the key and value. The number of attention heads is set to 4, and the query, key, and value dimensions are all 64-dimensional. The interaction features between audio and temporal features are obtained through a scaled dot product attention calculation formula. Audio-to-Physio attention calculates attention weights using audio features as the query and physiological features as the key and value, using the same calculation method as Audio-to-Sensor attention. The output of cross-modal attention is concatenated and fused with the original audio features, temporal features, and physiological features to obtain a 384-dimensional joint feature vector. The mathematical expression of cross-modal fusion is as follows:

[0043] in, For audio feature vectors, For time series feature vectors, For physiological feature vectors, For Audio-to-Sensor attention output, This is for Audio-to-Physio attention output.

[0044] The third stage involves multi-task joint output. The joint feature vector is sequentially passed through a fully connected layer and a classification head for multi-task prediction. The emotion state classification head contains four output neurons, corresponding to four emotion states: calm, tense, resistant, and crying. After normalization by the Softmax function, it outputs the probability distribution of each state, along with three continuous scores: valence V (range -1, +1, negative values ​​indicate negative emotions, positive values ​​indicate positive emotions), arousal A (range 0, 1, representing the degree of physiological activation), and dominance D (range -1, +1, negative values ​​indicate resistance, positive values ​​indicate cooperation). The operational compliance level judgment head contains four output neurons, corresponding to four operational compliance levels: compliant, slightly deviated, significantly deviated, and dangerous. It also outputs deviation direction labels (such as angle shift, pressure exceeding limits, flow fluctuation, etc.). The risk event type identification head contains multiple binary output nodes, corresponding to risk event types such as choking, violent shaking, pipeline blockage, and pressure exceeding limits. It outputs the probability of occurrence of each event. If the probability exceeds 0.5, it is judged as a risk event of that type. At the same time, it outputs the priority ranking of each type of risk.

[0045] The specific implementation process of the operational standardization assessment in step 3 is as follows.

[0046] The embedded central processing unit invokes the built-in operation correction engine to perform an operation standardization assessment task. The operation correction engine accesses a built-in database of pediatric physiological parameters, which covers the average nasal cavity volume of children aged 3 to 12 years (approximately 20 mL for 3-year-olds, 30 mL for 6-year-olds, and 40 mL for 12-year-olds), mucosal tolerance pressure limits (set to 60 to 100 kPa according to age group), and recommended irrigation flow rate ranges (50 to 100 mL / min for children aged 3 to 6 years, and 100 to 200 mL / min for children aged 6 to 12 years). The operation standardization assessment comprehensively considers the following factors: deviation between real-time flow rate and recommended flow rate range, deviation between grip pressure and safety threshold, deviation between device attitude angle and safety angle range, and flow rate stability index.

[0047] The real-time flow velocity data is compared with the recommended flow velocity range to calculate the percentage deviation. The formula for calculating the percentage deviation is: Flow velocity deviation percentage = |(Measured flow velocity - Target flow velocity) / Target flow velocity| × 100%. A flow velocity deviation percentage of less than 15% is considered normal fluctuation; a deviation percentage between 15% and 30% is considered slight deviation; a deviation percentage between 30% and 50% is considered significant deviation; and a deviation percentage exceeding 50% is considered a dangerous state.

[0048] The grip pressure data is compared with a safety threshold, which is dynamically adjusted based on the child's grip strength baseline. The baseline is estimated by the sliding average of the grip pressure over the previous 5 seconds. The pressure deviation is calculated by comparing the absolute deviation of the instantaneous pressure from the baseline value with the standard deviation of the pressure. When the absolute deviation exceeds 200% of the baseline value or the standard deviation of the pressure exceeds the preset threshold, it is considered an abnormal grip state, which may reflect the child's emotional tension or improper operation.

[0049] The device's attitude angle data is compared with the safe angle range. The safe angle range is determined by querying a database of recommended values ​​for the corresponding age group using children's physiological parameters. The default safe pitch angle range is 45 to 60 degrees, and the safe roll angle range is ±30 degrees. If the pitch angle exceeds the safe range for more than 2 seconds, it is considered an abnormal attitude. Similarly, if the roll angle exceeds the safe range for more than 2 seconds, it is also considered an abnormal attitude.

[0050] Flow velocity stability is assessed by calculating the ratio of the standard deviation to the mean of the flow velocity signal (coefficient of variation), with a calculation window length of 10 seconds. A coefficient of variation less than 0.1 indicates stable flow velocity; a coefficient of variation between 0.1 and 0.3 indicates slight fluctuation; a coefficient of variation between 0.3 and 0.5 indicates significant fluctuation; and a coefficient of variation exceeding 0.5 indicates a dangerous state.

[0051] The overall assessment of operational compliance level employs a weighted scoring mechanism, with the following weight allocations for each factor: flow velocity deviation 0.3, grip pressure deviation 0.25, attitude angle deviation 0.25, and flow velocity stability 0.2. The final operational compliance level is determined based on the weighted scoring results: a comprehensive score greater than or equal to 0.8 is considered compliant; a comprehensive score between 0.6 and 0.8 is considered slightly deviated; a comprehensive score between 0.4 and 0.6 is considered significantly deviated; and a comprehensive score less than 0.4 is considered dangerous.

[0052] The specific implementation process of the adaptive flushing control in step 4 is as follows.

[0053] The embedded central processing unit determines the target flushing power based on the emotion state classification results and normativity level using an adaptive PID control algorithm. The control system adopts a composite control architecture combining a finite state machine and an adaptive PID controller. The finite state machine manages the 12 main states of the flushing process and their state transition logic, while the adaptive PID controller achieves precise closed-loop control of the flushing flow rate.

[0054] The state transition logic of the finite state machine is defined as follows: After the system is powered on, it first enters the standby state, waiting for the user's start command; upon receiving the start command, it enters the self-test state, sequentially checking the working status of the power (remaining power must be greater than 20%), liquid level (liquid level must be greater than 10%), temperature (temperature must be within the range of 35 to 42 degrees Celsius), pressure sensor (pressure sensor output must be within the normal range), flow sensor (flow sensor output must be within the normal range), and nozzle contact sensor (contact sensor must detect effective contact). If any sensor detects an abnormality, it enters the fault mode and alarms; after the self-test passes, it enters the preparation guidance state, displaying a preparation animation on the LED expression screen and playing the preparation guidance script through the audio output unit; after waiting for the IMU to detect that the device is being held (the holding pressure exceeds 30% of the baseline value) and the holding lasts for more than 1 second, it enters the nozzle contact detection state; the nozzle contact detection state confirms that the proximity sensor has been triggered, Once the IMU posture is within a safe angle range and the pressure distribution is even, it enters a low-speed pre-rinse state. The low-speed pre-rinse state operates at 30% of the standard flow rate for 5 seconds, and after the pressure stabilizes, it enters the normal rinsing state. In the normal rinsing state, it operates according to the standard flow rate set in the child's profile, while continuously monitoring emotional state and compliance level. If emotional tension, mild resistance, or slight posture deviation is detected, it enters a soothing and slow-down state, reducing the target flow rate to 50% of the standard flow rate and playing soothing words. After the child's emotions stabilize and the operation remains compliant for 5 seconds, it returns to the normal rinsing state. If the nozzle detaches, violent shaking occurs, or coughing triggers the system, it enters a pause-and-recovery state, pausing pump output and prompting the user to re-attach or wait for recovery. If pressure exceeds limits, temperature is abnormal, liquid level is insufficient, or a predicted risk event is detected, it enters an abnormal shutdown state, immediately stopping the pump and opening the electromagnetic pressure relief valve. After the rinsing time reaches the target and the cumulative flow rate meets the target, it enters a completion feedback state, playing a reward animation and issuing a virtual badge.

[0055] The adaptive PID controller aims to ensure the flushing flow rate tracks a setpoint, determined by the child's profile and current status. The baseline flow rate is determined as follows: for children under 4 years old with 3 or more previous resistance attempts, the baseline flow rate is 50% of the standard flow rate; for children aged 4 to 6 years old with 1 to 2 previous resistance attempts, the baseline flow rate is 75% of the standard flow rate; and for children over 6 years old with zero previous resistance attempts, the baseline flow rate is 100% of the standard flow rate. The baseline flow rate is adjusted gradually, starting at 30% of the initial flow rate and increasing by 20% every 5 seconds until the target flow rate is reached, avoiding sudden pressure changes that could physically irritate the nasal mucosa.

[0056] The output of the adaptive PID controller is the PWM duty cycle, which controls the speed of the peristaltic pump and thus adjusts the flushing flow rate. The standard form of the PID control law is:

[0057] in, The controller output (the PWM duty cycle at the current moment). This refers to the flow velocity deviation (the difference between the set value and the measured value). For proportional gain, For integral gain, This is the differential gain. The basic PID parameters are set as follows: , , .

[0058] The control gain parameter is nonlinearly adaptively adjusted based on real-time emotional state and normativity level. The calculation rule for the emotion compensation coefficient is: when the emotional state is calm, the compensation coefficient is 1.0. , , Maintain the baseline value; when the emotional state is tense, the compensation coefficient is 0.7. Reduced to 70% of the base value. Reduced to 70% of the base value. Maintain the baseline value to improve system stability and extend response time; when the emotional state is resistance, the compensation coefficient is 0.5. Reduced to 50% of the base value. Reduced to 50% of the base value. Maintain the baseline value to avoid overshoot caused by the accumulation of integral terms; when the emotional state is crying, the compensation coefficient is 0.3. Reduced to 30% of the base value. Reduced to 30% of the base value. Maintain the baseline value while reducing pump output to the minimum sustaining flow rate.

[0059] The rule governing the impact of normativity levels on control parameters is as follows: When the normativity level is a slight deviation, differential feedforward compensation is introduced, and the differential feedforward term is... ,in The rate of change of pressure, The feedforward coefficient (set to 0.5) is used to predict pressure change trends and adjust flow rates in advance; when the normativity level is significantly deviated, it is further reduced. The value is increased to 60% of the base value, and the differential feedforward coefficient is increased to 0.8; when the normative level is dangerous, the system is directly switched to the safe shutdown state.

[0060] The feedforward compensation term is calculated based on the pipeline pressure change trend. The pipeline pressure sensor monitors the outlet pressure in real time. When an upward pressure trend is detected and the flow deviation has not yet changed significantly, the feedforward compensation term increases the PWM duty cycle in advance to compensate for the impact of increased pipeline resistance on flow. When there is a risk of pipeline blockage (based on the output of the predictive safety assessment module), the flow rate is reduced. Reduce to 50% of the base value to avoid high-frequency oscillations.

[0061] After each nasal irrigation task, the incremental learning controller updates the personalized configuration of the PID parameters based on the control sequence of that task. The control sequence includes a PWM duty cycle sequence (sampled every 100ms), a flow feedback sequence, and a mood fluctuation trajectory. The parameter update algorithm employs elastic weight consolidation, with the objective function being the sum of squared control errors for the current task plus a regularization term. The regularization term penalizes the parameter offset relative to the historical optimal parameters. The loss function of elastic weight consolidation is as follows:

[0062] in, For task error loss, The elasticity coefficient (set to 5000). These are the diagonal elements of the Fisher information matrix. These are the historically optimal parameter values. After each task, the controller calculates the gradient based on the task data and updates the personalized parameters. After approximately 10 usage cycles, the controller parameters are fully adapted to the individual characteristics of the specific child, achieving precise care tailored to each individual.

[0063] The specific implementation process of the emotional interaction guidance in step 5 is as follows.

[0064] The embedded central processing unit schedules the interaction module to execute emotionally-oriented interactive guidance based on the phased progress of the nasal irrigation task and the emotional state classification results. The interaction module includes an LED expression screen, an audio output unit, a vibration feedback component, and an optional olfactory feedback component.

[0065] The dynamic dialogue generation system adopts a template-based architecture, comprising three components: a stage template library, an emotion modifier, and personalized variable population. The stage template library defines the basic dialogues for each task stage, containing approximately 200 pre-set dialogues covering eight stages: preparation, bonding, pre-rinsing, rinsing, soothing, pause, exception, and completion. Examples of basic dialogues for the preparation stage include: "The little elephant is ready, are you ready?"; the bonding stage includes: "Gently attach the little elephant's trunk"; the pre-rinsing stage includes: "We're about to begin, are you ready, kids?"; the rinsing stage includes: "Rush, rush, the little elephant's trunk is clean!"; the soothing stage includes: "It's okay, take your time, the little elephant is with you"; the pause stage includes: "Take a break, no rush"; the exception stage includes: "The little elephant isn't feeling well, let's check it together"; and the completion stage includes: "Great job! You're so brave! Thank you, little elephant!" The emotion modifier adjusts the tone, speed, volume, and content of the speech based on the real-time recognized emotional state. In a calm state, it uses a normal speaking speed (3-5 words per second), standard volume (set to 70dB SPL), and standard tone; in a tense state, it slows down the speaking speed (reduced to 0.8 times the normal speed), lowers the volume (reduced to 70% of the standard volume), and adds encouraging phrases (such as "You're great," "Keep it up!"); in a resistant state, it uses very short sentences (no more than 5 words per sentence), a gentle tone (speaking speed reduced to 0.6 times the normal speed, volume reduced to 50% of the standard volume), and repeated reassurances (such as "Don't be afraid," "Little elephant is here"); in a crying state, it switches to silent mode, does not play the spoken speech, and uses an emoji screen and vibration feedback instead of voice output.

[0066] Personalized variable population is performed using dialogues tailored to the child's profile information. Personalized variables include the child's name, performance history, and reward badge status. The child's name is retrieved from the child's profile and inserted into the dialogue, such as "Go, Xiaoming!"; the performance history is evaluated based on the completion rate of the last three nasal rinsing tasks, using encouraging phrases like "You did great last time, you can do it again this time!" and comforting phrases like "It's okay, we'll take it slow!"; the reward badge status is based on the virtual badges the child has earned, such as "You've earned the Little Warrior Badge!"

[0067] The multi-dimensional feedback system uses a 64×64 pixel LED dot matrix display screen at a frame rate of 30fps to show the dynamic facial expressions of the elephant character. The expression sequence is linked to the task stage and emotional state: the preparation stage displays a blinking animation, the fitting stage displays an expectant animation, the rinsing stage displays a smiling animation (when cooperating) or a worried animation (when nervous), and the completion stage displays a cheering animation and a growth animation showing the elephant's trunk getting clean. The facial animations use pre-rendered frame sequences stored in local flash memory, with each expression containing 8 to 16 frames.

[0068] Auditory feedback uses preset speech scripts and sound effects, with the speech content dynamically changing according to the progress of the nasal irrigation task. Sound effects include key tones, success prompts, reward sounds, and warning sounds, stored in local flash memory in WAV format. Speech synthesis uses pre-recorded and spliced ​​methods, selecting child-friendly voices, with a sampling rate of 16kHz and a quantization bit depth of 16 bits.

[0069] The tactile feedback uses a linear vibration motor to provide gentle vibrations as positive reinforcement when a child correctly completes key steps. The vibration motor's driving parameters are: vibration duration of 200 milliseconds, vibration frequency of 100 Hz, and vibration intensity of 0.5g. Key steps that trigger tactile feedback include: successfully fitting the nozzle, rinsing correctly for 5 seconds, calming down from tension, and task completion.

[0070] Olfactory feedback is an optional module that releases a small amount of fragrance before rinsing begins. The fragrance is either strawberry or milk-based and is dispersed to the device's air outlet via a miniature air pump. The purpose of olfactory feedback is to provide positive olfactory stimulation when children are near the device, reducing their aversion to the distinctive odor of medical devices.

[0071] The task chain design adopts a gamified framework of caring for a baby elephant, designing the nasal irrigation process as a complete task chain including a preparation phase, a rinsing phase, a cleaning phase, and an achievement phase. The preparation phase guides children to recognize the baby elephant and pick up the device; the rinsing phase guides children to cooperate in completing the nasal irrigation; the cleaning phase prompts children to maintain the correct posture until the task is completed; the achievement phase provides rewards and positive feedback. During the rinsing phase, if a child shows a relaxed state for more than 3 seconds and operates correctly, the system triggers a hidden reward mechanism, scheduling the voice module to play a specific reward message, "Great job! You did a great job! The baby elephant gives you a big thumbs up!", and sends a virtual badge to the mobile terminal application via the wireless communication module. A long-term incentive mechanism is implemented through a virtual growth system. After each correct completion of the nasal irrigation task, children not only receive a virtual badge but also unlock new skills or decorations for the baby elephant in the application. Badge types include Little Warrior Badge, Perseverance Star Badge, and Cooperation Master Badge.

[0072] The specific implementation process of the predictive security assessment steps is as follows.

[0073] The embedded central processing unit incorporates a predictive safety assessment module. This module encodes the most recent 30 frames of multimodal time-series data using a trajectory prediction network to predict the system's state trajectory within the next 3 seconds, enabling a safety strategy upgrade from passive response to proactive prevention. The input data for the predictive safety assessment module includes pressure sequences (from pipeline pressure sensors), flow sequences (from flow sensors), attitude angle sequences (from the motion sensing submodule), heart rate sequences (from the PPG processing results of the physiological sensing submodule), and grip pressure envelopes (from the mechanical sensing submodule). All sequences are uniformly resampled at 10Hz, and each frame contains 5 channels of time-series data.

[0074] The trajectory prediction network employs an encoder-decoder architecture. The encoder uses a TCN structure, containing four residual convolutional blocks. Each residual block contains two layers of dilated causal convolutions with dilation coefficients of 1, 2, 4, and 8, and a kernel size of 3. The encoder encodes the 30-frame input sequence into a fixed-dimensional hidden state vector. The decoder uses an autoregressive structure, outputting one frame of prediction data per decoding step. The decoding process executes 30 steps to obtain the prediction sequence for the next 30 frames. The decoder output includes predicted pressure values, predicted flow values, predicted attitude angles, predicted heart rate, and predicted grip pressure envelopes.

[0075] The probabilities of various risk events are calculated based on predicted trajectories. The probability of choking risk is calculated based on the following combination of features: sudden heart rate increase (heart rate increase exceeding 5 beats / minute in the next 3 seconds), cough precursor sound feature (predicted cough precursor sound pattern in the audio characteristics), and flow mutation feature (predicted flow change rate exceeding 30% / second). A choking risk event is defined as the presence of at least two of these three features with a probability exceeding 0.3. The probability of resistance to escalation is calculated based on the following combination of features: pressure fluctuation feature (predicted grip pressure standard deviation exceeding 200% of the baseline value), sway amplitude feature (predicted acceleration envelope area exceeding 60% of the safe threshold), and heart rate variability feature (predicted continuous decrease in RMSSD). A resistance to escalation risk event is defined as the presence of at least two of these three features with a probability exceeding 0.4. The probability of pipeline blockage is calculated based on the following combination of features: a decreasing flow rate (predicted flow rate is more than 5% lower than the current flow rate for 3 consecutive frames) and an increasing pressure rate (predicted pressure is more than 3% higher than the current pressure for 3 consecutive frames). When both features appear simultaneously and the probability exceeds 0.35, it is determined to be a pipeline blockage risk event.

[0076] When the predicted probability exceeds a preset threshold, the system triggers corresponding protective actions in advance. When the probability of choking exceeds 0.3, the system reduces the target flow rate by 20% and triggers reassuring dialogue; when the probability of resistance escalation exceeds 0.4, the system enters a reassuring and speed-reducing state in advance; when the probability of pipeline blockage exceeds 0.35, the system triggers a reverse suction action (the peristaltic pump reverses for 500 milliseconds at 50% of its normal speed) to clear any possible blockages.

[0077] The priority of safety rules is set as follows: predictive protection output > threshold-triggered rule output > FSM state machine / PID controller output. This means that the judgment result of the predictive module can override the decision of the control layer and directly trigger safety protection actions without waiting for state transition confirmation from the FSM state machine.

[0078] The specific implementation process of the local security protection system is as follows.

[0079] The local security protection system operates independently of the cloud and network, and includes 12 security protection rules to ensure that devices can perform reliable security protection actions in any usage scenario. The security protection rules are divided into two categories: threshold-triggered and trend-predictive.

[0080] The threshold-triggered rules include the following eight: Pressure upper limit protection rule: Sets a threshold of 80 kPa; if the duration exceeds 200 milliseconds, it triggers pump stop protection. Nozzle detachment protection rule: Detects the contact sensor status at 100-millisecond intervals; if three consecutive detections fail, it triggers a pause. Attitude anomaly protection rule: Sets the pitch angle safety range to ±30 degrees and the roll angle safety range to ±45 degrees; if the range is exceeded for more than 2 seconds, it triggers a pause. Cough trigger protection rule: Sets the audio recognition confidence threshold to 0.7; if the cough sound is recognized for more than 500 milliseconds and accompanied by a sudden change in flow, it triggers a pause. Severe shaking protection rule: Sets the acceleration envelope area threshold to 5000 (mg·ms); if the threshold is exceeded, it triggers a pause. Temperature anomaly protection rule: Sets the flushing fluid temperature safety range to 35 to 42 degrees Celsius; if the temperature exceeds this range, the flushing pump is prohibited from starting. Insufficient liquid level protection rule: Sets the lower limit liquid level threshold to 20%; if the level is below the threshold, it triggers an alert. Flow change anomaly protection rule: Sets the instantaneous flow rate change amplitude threshold to ±30%; if the change lasts for 500 milliseconds, it triggers flow adjustment.

[0081] The trend prediction rules include the following four: The cough prediction protection rule, based on the output of the predictive safety assessment module, triggers preventative speed reduction when the probability of cough risk exceeds 0.3; the resistance escalation protection rule, based on the output of the predictive safety assessment module, triggers preventative soothing speed reduction when the probability of resistance escalation exceeds 0.4; the pipeline blockage warning rule, based on the output of the predictive safety assessment module, triggers preventative reverse suction when the probability of pipeline blockage exceeds 0.35; and the battery anomaly protection rule monitors the battery voltage change trend, triggering a low battery warning when the voltage drop rate exceeds twice the normal discharge rate.

[0082] The safety protection actions, executed in descending order of priority, are: immediate pump stop, opening the solenoid pressure relief valve, audible and visual alarm, and status lock. Immediate pump stop is achieved by cutting off the peristaltic pump's PWM signal, with a response time of less than 10 milliseconds; opening the solenoid pressure relief valve is achieved by driving the solenoid valve coil. The solenoid valve is normally closed, with a rated voltage of 5V and a response time of less than 50 milliseconds; the audible and visual alarm is achieved by displaying warning expressions on the LED screen and playing warning sounds through the audio output unit; status lock locks the system in an abnormal shutdown state, requiring manual confirmation to unlock.

[0083] The hardware self-test logic executes before system startup. The self-test process includes applying minute excitation signals to each sensor component and detecting feedback current. During the audio sensing submodule self-test, a 1000Hz sine wave excitation signal is applied to the microphone to check if the ADC sampling value is within the expected range. During the physiological sensing submodule self-test, a calibration current is applied to the PPG sensor LED to detect the photodiode output signal strength. During the motion sensing submodule self-test, the sensor ID register is read and verified via the I2C interface. During the mechanical sensing submodule self-test, the zero-point output value of the pressure sensor is read and verified. During the liquid sensing submodule self-test, zero-point calibration of the flow sensor and internal reference voltage detection of the temperature sensor are performed. If a core sensor failure is detected, a safety lock state is entered, and a specific fault code is sent to the parent's mobile phone via the wireless communication module.

[0084] The embedded central processing unit maintains a watchdog timer with a timing period of 1 second. It automatically triggers a reset when the processor's main loop execution time exceeds 1 second, ensuring automatic system recovery in abnormal situations. The watchdog timer pauses its timing when entering an abnormal shutdown state or a completion feedback state. Example 2

[0085] To meet the special care needs of children under 3 years old, this embodiment provides a sensory strategy adjustment scheme and corresponding control parameter optimization configuration for young children.

[0086] Young children have limited language expression abilities and find it difficult to accurately express their emotional states through speech. Therefore, the system automatically adjusts the fusion weights of multimodal perception signals. The fusion weight of audio features is reduced to 0.3 (0.4 in the standard mode), the fusion weight of mechanical features is increased to 0.5 (0.35 in the standard mode), and the fusion weight of motion features is increased to 0.2 (0.25 in the standard mode). In the adjusted multimodal fusion formula, the weighting coefficients of each modality feature vector are redistributed to ensure that mechanical and motion signals play a dominant role in emotion recognition.

[0087] The 6-axis IMU, in child-friendly mode, monitors whether a child is struggling violently. If the acceleration envelope area output by the inertial measurement unit exceeds 60% of the preset safety range, even if the voice signal does not indicate any abnormality, the system determines that the operating environment is unstable and immediately suspends the flushing pump. When the acceleration envelope area exceeds 80% of the safety range, an emergency stop protection is triggered.

[0088] For young children, the "Ultra-Gentle Mode" has a pre-set parameter set, with the first pressure threshold set at 40% of the standard mode, or 32 kPa (the standard mode is 80 kPa). The flow rate rise slope is limited to a very small range, ensuring that pressure changes are almost imperceptible. The operation correction engine monitors the instantaneous fluctuations of the flow sensor every 5 milliseconds (the standard mode monitors every 20 milliseconds). Once a tiny change in breathing resistance is detected (pressure rise rate exceeding 0.5 kPa / second), the machine immediately shuts down.

[0089] In early childhood mode, the emoji screen uses high-contrast color block animations with 30% increased color saturation, simplifying visual complexity to suit the visual cognitive characteristics of young children. Voice content primarily imitates animal sounds and simple reduplicated words, such as monosyllabic phrases like "ee-ee," "ya-ya," and "woo-woo," with each sentence lasting less than 2 seconds. In early childhood mode, the olfactory module prioritizes releasing a strawberry scent, with the fragrance concentration increased by 20%.

[0090] In terms of hardware structure, the nozzle uses a spherical protective cover made of food-grade silicone with a hardness of Shore A 30. The nozzle body and the protective cover are connected by snap-fit, and the protective cover is removable for cleaning. The purpose of the spherical protective cover is to prevent mechanical injury to children caused by the nozzle entering the nasal cavity too deeply during shaking. The reservoir bottle has an anti-backflow one-way valve structure, and the one-way valve opening pressure is set at 5kPa to ensure that nasal secretions do not flow back into the pump body and reservoir bottle. Generality Statement

[0091] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the specific embodiments described above, which are merely illustrative and not restrictive. The scope of protection of the present invention is defined by the appended claims. Various equivalent modifications or substitutions made by those skilled in the art to the present invention without departing from its principles fall within the scope of protection of the present invention.

Claims

1. A control method for a children's intelligent nasal irrigation robot based on AI multimodal emotion perception and adaptive correction, characterized in that, Includes the following steps: Step 1: Multi-source sensor data acquisition. This involves acquiring the child's voice audio signals via a sound acquisition component, collecting temporal characteristic data of the handle grip force via a flexible pressure sensor array, acquiring physiological parameter data reflecting emotional arousal via a physiological signal sensor, acquiring real-time flow rate data of the rinsing fluid via a flow monitoring component, acquiring device posture and motion data via a motion sensing component, and acquiring external environmental data via an environmental sensing component. All multi-modal sensor data are then synchronously transmitted to the embedded central processing unit (CPU). Step 2: Multi-modal emotion recognition and classification. The embedded CPU uses its built-in emotion perception engine to preprocess the voice audio signals, extracting Mel-frequency feature vectors and combining them with a lightweight convolutional neural network for acoustic feature extraction. Simultaneously, it combines heart rate variability and skin conductance response characteristics fed back from the physiological signal sensor, performing a multi-modal feature fusion operation based on a three-stage progressive fusion mechanism to identify and output the child's current emotional state classification result, along with valence, arousal, and dominance scores. Step 3: Operational compliance assessment. The embedded CPU uses its built-in operational correction engine to combine the real-time acquired flow rate data with the grip pressure data. The data is compared in real time with preset safety thresholds corresponding to specific age groups and nasal physiological structures. The spatial tilt angle of the nasal irrigation robot is monitored by a six-axis accelerometer, and the instantaneous peak pressure and flow rate stability are comprehensively evaluated to calculate and output the standardization level of the current nasal irrigation operation. Step 4: Adaptive irrigation control. The embedded central processing unit determines the target irrigation power based on the emotional state classification results and standardization level through a preset adaptive proportional-integral-derivative control algorithm, and dynamically adjusts the pulse width modulation drive signal parameters of the micro adjustable irrigation pump to achieve closed-loop regulation of the output pressure and output flow rate of the irrigation fluid. The gain parameter of the control algorithm is nonlinearly adaptively adjusted according to the real-time emotional state and standardization level. Step 5: Emotional interactive guidance. Based on the phased progress of the nasal irrigation task and the emotional state classification results, the embedded central processing unit schedules the display screen, audio output unit, and vibration feedback component in the interaction module to output corresponding dynamic facial expression sequences, guiding voice prompts, and tactile stimulation signals to guide children to execute and complete the nasal irrigation task chain consisting of preparation, irrigation, cleaning, and completion stages.

2. The control method according to claim 1, characterized in that, The specific implementation of multi-source sensing data acquisition in step 1 is as follows: The audio sensing submodule uses a digital microphone to collect sound and performs real-time preprocessing of children's speech through an adaptive spectral subtraction noise reduction algorithm. The preprocessing process includes bandpass filtering to filter out noise below a preset low-frequency threshold and above a preset high-frequency threshold, and spectral subtraction to eliminate mechanical vibration noise generated by the nasal irrigator. The physiological sensing submodule integrates a photoplethysmography (PPG) sensor and a skin conductance sensor in the grip handle area. The PPG sensor extracts real-time heart rate data by emitting infrared light of a specific wavelength and receiving reflected light, and calculates heart rate variability indicators including the root mean square of the difference between adjacent normal heartbeat intervals. The skin conductance sensor is used to monitor changes in sweat gland secretion caused by emotional stress; the motion sensing submodule uses a 6-axis inertial measurement unit, which uses Kalman filtering to fuse accelerometer and gyroscope data to output stable attitude angle estimates; the force sensing submodule arranges a flexible pressure sensor array in the handle area, and the collected analog signals are converted into digital sequences by a multi-channel analog-to-digital converter. The processor performs moving average filtering on the digital sequences to eliminate random interference and extracts pressure envelope features that reflect emotional stress; the liquid sensing submodule includes a miniature flow sensor, a liquid level sensor, and a temperature sensor to monitor the real-time flow rate of the flushing fluid, the remaining flushing fluid in the storage bottle, and the flushing fluid temperature.

3. The control method according to claim 1, characterized in that, Step 2's three-stage progressive fusion architecture includes: The first stage is modal-independent encoding. The audio encoding branch uses Mel-spectrum feature extraction combined with a lightweight convolutional neural network encoder. The mechanics and motion encoding branches employ an architecture combining a temporal convolutional network and a bidirectional long short-term memory network. The temporal convolutional network uses a dilated causal convolutional structure. The physiological encoding branch extracts temporal and frequency domain features from heart rate and electrodermal signals. The second stage is cross-modal interactive fusion, employing a multi-level cross-modal attention mechanism to establish connections between different modalities. Sound-to-sensor attention is used to establish connections between sound events and device states, and sound-to-physiology attention is used to establish connections between sound events and physiological states. The third stage is multi-task joint output, simultaneously outputting the child's emotional state classification results, valence arousal dominance continuous scores, operational compliance level judgment results, and risk event type identification results.

4. The control method according to claim 1, characterized in that, The specific implementation of the operational standardization assessment in step 3 is as follows: The embedded central processing unit calls the children's physiological parameter database, which covers the average nasal cavity volume, mucosal tolerance pressure limit, and recommended irrigation flow rate range for children of different ages; the operational standardization assessment comprehensively considers the deviation between the real-time flow rate and the recommended flow rate range, the deviation between the grip pressure and the safety threshold, the deviation between the device posture angle and the safety angle range, and the flow rate stability index; the comprehensive determination of the standardization level adopts a weighted scoring mechanism, with the weights of flow rate deviation, grip pressure deviation, posture angle deviation, and flow rate stability allocated according to a preset ratio, and finally the standardization level is determined as standard, slight deviation, significant deviation, or dangerous based on the weighted scoring results.

5. The control method according to claim 1, characterized in that, In step 4, the control gain coefficient of the adaptive proportional-integral-derivative controller is adaptively adjusted according to the child's emotional state, current stage, individual characteristics, and pipeline status; the nonlinear compensation term is dynamically calculated based on the real-time identification results: when the child's emotional state is tense, the proportional gain is reduced to improve system stability and extend system response time; when the emotional state is resistant, the integral gain is further reduced to avoid overshoot caused by the accumulation of integral terms; when the normativity level is slightly deviated, differential feedforward compensation is introduced to predict the pressure change trend. When there is a risk of blockage in the tubing, the differential gain is reduced to avoid high-frequency oscillations; the target flow rate is determined by the child's profile and current status, and the basic flow rate is adjusted gradually to avoid physical irritation to the nasal mucosa caused by sudden pressure changes.

6. The control method according to claim 1, characterized in that, It also includes an incremental learning controller, which updates the personalized configuration of proportional, integral, and derivative parameters based on the control sequence of the current task after each nasal irrigation task. The update algorithm uses elastic weight consolidation technology to prevent the learning of key parameters from being overwritten by new tasks. When the child is young, has a history of resistance, or has a lot of recent abnormalities, the system automatically uses a lower base flow rate; the base flow rate is adjusted using a gradual strategy, starting from 30% of the initial flow rate and gradually increasing to the target flow rate.

7. The control method according to claim 1, characterized in that, In step 5, the dynamic script generation system adopts a template-based architecture, which includes three parts: stage templates, emotion modifiers, and personalized variables. The stage templates define the basic scripts for each task stage. The emotion modifiers adjust the tone and length of the scripts according to the real-time recognized emotional state: the calm state uses normal speaking speed and standard volume, the tense state uses slower speaking speed, lower volume and more encouraging sentences, the resistant state uses very short sentences and a gentle tone, and the crying state switches to silent mode and uses an emoji screen and vibration feedback instead of voice. Personalized variables include the child's name, historical performance records and reward badge status.

8. The control method according to claim 1, characterized in that, It also includes a multi-dimensional feedback system, comprising four dimensions: visual feedback, auditory feedback, tactile feedback, and olfactory feedback; visual feedback uses a display screen to show dynamic facial expression sequences, with the content of the expressions linked to the current task stage and emotional state; auditory feedback uses preset voice prompts and sound effects; The tactile feedback uses a linear vibration motor to provide gentle vibrations as positive reinforcement when the child correctly completes key steps; the olfactory feedback releases a small amount of fragrance before rinsing begins.

9. The control method according to claim 1, characterized in that, It also includes a predictive safety assessment step, which encodes a preset number of historical frames through a trajectory prediction network to predict the system state trajectory within a preset time period in the future; the input data includes pressure sequence, flow sequence, attitude angle sequence, heart rate sequence and grip pressure envelope; based on the predicted trajectory, the probability of coughing risk, the probability of resistance escalation and the probability of pipeline blockage are calculated; when the predicted probability exceeds a preset threshold, a calming deceleration or reverse suction action is triggered in advance to achieve preventive safety protection; the safety rules have higher priority than the output of the finite state machine and the proportional-integral-derivative controller.

10. The control method according to claim 1, characterized in that, It also includes a local safety protection system, which operates independently of the cloud and network, with multiple local safety protection rules. The safety protection rules are divided into two categories: threshold-triggered and trend-predictive. Threshold-triggered rules include pressure limit protection, nozzle detachment protection, abnormal posture protection, cough trigger protection, severe shaking protection, abnormal temperature protection, insufficient liquid level protection, and abnormal flow protection. Trend-predictive rules include cough prediction protection, resistance escalation protection, pipeline blockage warning, and battery abnormality protection. The execution of safety protection actions includes immediate pump stoppage, opening of the electromagnetic pressure relief valve, audible and visual alarm prompts, and status lock. The hardware self-test logic is executed before system startup, applying excitation signals to each sensor component and detecting feedback current.