A method, apparatus, device and medium for monitoring mouth breathing of children

By collecting multimodal physiological data through a non-contact multimodal sensor array and combining it with thermal imaging, radar, and audio analysis, accurate monitoring of children's mouth breathing in the home environment is achieved. This solves the problems of high invasiveness and high cost in existing technologies, provides long-term and dynamic monitoring data support, and improves the accuracy and timeliness of diagnosis.

CN121176862BActive Publication Date: 2026-02-24AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511736730.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-24
Estimated Expiration
2045-11-25

AI Technical Summary

Technical Problem

Existing polysomnography methods, such as PSG, are highly invasive, costly, and difficult to monitor for long periods in diagnosing mouth breathing in children. They cannot accurately reflect daily sleep patterns and cannot continuously monitor intermittent or posture-dependent mouth breathing events.

Method used

A non-contact sensor array, including thermal imaging, radar, and audio sensors, is used to collect multimodal physiological data. Through the fusion analysis of thermodynamic change characteristics, respiratory effort waveforms, and respiratory audio spectrum characteristics, a pre-trained open-mouth breathing recognition model is used to make accurate judgments and generate monitoring reports.

Benefits of technology

It enables non-intrusive, long-term, and comfortable monitoring of children's mouth breathing in a home environment, provides accurate assessment of functional mouth breathing, supports the accuracy and timeliness of clinical diagnosis, overcomes the limitations of a single sensor, and provides long-term dynamic data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121176862B_ABST
    Figure CN121176862B_ABST
Patent Text Reader

Abstract

The application discloses a kind of child mouth breathing monitoring method, device, equipment and medium, comprising: using non-contact sensor group synchronous acquisition target child during sleep Multimodal physiological data, Multimodal physiological data at least include face thermal imaging video stream, chest and abdomen micro-motion signal and environmental audio signal;Multimodal physiological data is handled, respectively extract and breathe related Multimodal features, Multimodal features include based on the thermodynamic change feature of extracting oral-nasal region in face thermal imaging video stream, based on the respiratory effort waveform of extracting in chest and abdomen micro-motion signal, and based on the sound source spatial position and spectral feature of extracting respiratory sound in environmental audio signal;Multimodal features are input to pre-trained mouth breathing identification model, and the judgment result that target child occurs functional mouth breathing event in specific time period is output;Based on the judgment result, the mouth breathing monitoring report of target child is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of respiratory monitoring technology, and in particular to a method, device, equipment and medium for monitoring children's open-mouth breathing. Background Technology

[0002] Mouth breathing in children is a common but serious health problem that should not be ignored. It may indicate upper airway obstruction (such as adenoid hypertrophy), and if left untreated, it can easily lead to serious consequences such as maxillofacial developmental abnormalities (forming "adenoid facies"), tooth decay, and sleep apnea. Therefore, early and accurate monitoring and assessment of children's breathing patterns during sleep has significant clinical value.

[0003] Currently, the standard for clinical diagnosis of mouth breathing is polysomnography (PSG) performed in a hospital sleep laboratory. This method indirectly determines the breathing pathway by attaching thermal / pressure airflow sensors to the child's face and placing electromyography sensors in the jaw. However, PSG has significant limitations: First, it is highly invasive; the sensors and cables cause significant discomfort and a foreign body sensation for children, severely interfering with their natural sleep state and leading to data distortion, the so-called "first night effect," making the monitoring results unable to accurately reflect daily sleep patterns. Second, PSG must be performed in a specific laboratory environment, usually only recording data for a single night, making it costly and difficult to popularize, and unable to achieve long-term, continuous home monitoring of children's breathing patterns, thus potentially missing intermittent or posture-dependent mouth breathing events. Summary of the Invention

[0004] This specification provides one or more embodiments of a method, device, equipment, and medium for monitoring children's open-mouth breathing, in order to solve the technical problems mentioned in the background art.

[0005] One or more embodiments of this specification employ the following technical solutions:

[0006] This specification provides one or more embodiments of a method for monitoring mouth breathing in children, the method comprising:

[0007] Multimodal physiological data of the target child during sleep is collected synchronously using a non-contact sensor array. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0008] The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0009] The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period.

[0010] Based on the judgment results, an open-mouth breathing monitoring report for the target child is generated.

[0011] It should be noted that this invention, by employing non-contact multimodal sensing technologies (such as thermal imaging, radar, and audio) to replace the invasive method of traditional PSG that requires attaching sensors to the face, fundamentally eliminates the interference of the device on children's sleep, making the monitoring process imperceptible. This allows for the first-ever acquisition of children's real and natural sleep breathing data in a home environment. Furthermore, by fusing and analyzing the thermodynamic characteristics of the nasal and oral regions (quantifying airflow pathways), chest and abdominal respiratory effort (reflecting respiratory resistance), and the sound source and spectral characteristics of respiratory sounds (distinguishing sound sources and characteristics), this method achieves accurate judgment of "functional mouth breathing" (i.e., airflow actually passing through the oral cavity) rather than just the act of "opening the mouth." This overcomes the limitations of single sensors, which are susceptible to environmental interference and cannot distinguish the nature of airflow. Ultimately, this ability to operate comfortably and continuously at home and provide accurate results makes home monitoring for weeks or even months possible. This provides doctors with long-term, dynamic data support that traditional single PSG tests cannot match, enabling the capture of intermittent or posture-dependent respiratory events and the assessment of disease progression, greatly improving the accuracy of clinical diagnosis and the timeliness of intervention.

[0012] Furthermore, the thermodynamic change feature is the airflow ratio between the mouth and nose, and the extraction of the thermodynamic change features of the mouth and nose region based on the facial thermal imaging video stream includes:

[0013] The center point of the nostril and the center point of the lip crease in the thermal imaging video stream are located based on the key point detection network.

[0014] Using the center point of the nostril as a reference, the region of interest in the nasal cavity is dynamically defined;

[0015] Using the center point of the labial suture as a reference, the region of interest in the oral cavity is dynamically defined;

[0016] Determine the first average temperature change rate per unit time in the region of interest of the nasal cavity;

[0017] Determine the second average temperature change rate per unit time in the oral cavity region of interest;

[0018] The airflow ratio between the mouth and nose is determined based on the ratio of the first average temperature change rate to the second average temperature change rate.

[0019] It should be noted that this invention utilizes thermal imaging video streams to non-contactly monitor temperature changes in the nasal and oral regions, and innovatively proposes the quantitative indicator of "nasal-oral airflow ratio." This fundamentally avoids the discomfort and interference caused by attaching sensors to children's faces, achieving imperceptible physiological signal acquisition. Furthermore, by dynamically locating the center of the nostrils and lip creases and defining the region of interest through a key point detection network, this method can adapt to the slight head movements during children's sleep, ensuring the continuity and stability of monitoring. Most importantly, by calculating the average temperature change rate per unit time in the nasal and oral regions and using their ratio as the basis for judgment, this method transforms traditional qualitative and subjective observations (such as whether the mouth is open) into quantitative and objective measurements of airflow pathways. This effectively distinguishes between habitual mouth breathing and true functional mouth breathing caused by nasal congestion. This key advancement makes it possible to obtain accurate judgments with clear clinical value comparable to polysomnography in natural sleep scenarios such as at home.

[0020] Furthermore, the extraction of respiratory effort waveform based on the chest and abdominal micromotion signals includes:

[0021] By using frequency-modulated continuous wave radar sensors deployed in the monitoring area, electromagnetic waves are continuously emitted toward the child's chest and abdomen and their echoes are received to obtain the raw baseband signal.

[0022] The original baseband signal is processed by Fast Fourier Transform to select the specific distance cell with the highest signal-to-noise ratio;

[0023] Phase demodulation is performed on the signal sequence within the specific distance unit to obtain a phase change signal that is proportional to the displacement of the chest and abdominal surface;

[0024] A bandpass filter with a specified passband frequency is applied to the phase change signal to separate the breathing effort waveform caused by breathing.

[0025] It should be noted that this invention employs a frequency-modulated continuous wave radar sensor to non-contactly detect subtle movements of the chest and abdomen. This firstly achieves precise perception of body surface displacement during respiration, completely avoiding the restrictive feeling associated with wearable sensors. Secondly, by performing a fast Fourier transform on the radar echo signal and intelligently selecting the specific distance unit with the highest signal-to-noise ratio, this method effectively eliminates environmental clutter interference, stably locking the monitoring target onto the child's chest and abdomen, ensuring the accuracy and reliability of the signal source. Furthermore, by performing phase demodulation on the signal from the specific distance unit, minute displacement information is extracted from the echo signal with high precision. Then, a bandpass filter removes large body movements and high-frequency noise unrelated to respiration, ultimately separating a clean, continuous periodic waveform representing respiratory effort. This complete technical chain from radar signal to respiratory effort waveform makes it possible to continuously and stably acquire key physiological indicators reflecting the patency of the upper airway without contact with the child's body. This provides indispensable objective quantitative evidence for accurately determining whether functional mouth breathing, requiring additional respiratory effort due to nasal obstruction, exists.

[0026] Furthermore, the extraction of the sound source spatial location and spectral features of breathing sounds based on the environmental audio signal includes:

[0027] Multi-channel audio signals from the environment are simultaneously acquired by deploying a multi-microphone array in the monitoring area;

[0028] Determine the direction angle of arrival of the multi-channel audio signal to locate the spatial position of the sound source of the breathing sound;

[0029] Determine whether the spatial location of the sound source is within a spatial range centered on the child's mouth and nose area;

[0030] If so, the multi-channel audio signal is determined to be a candidate breath sound signal;

[0031] The candidate breath sound signals are processed to convert them from the time domain to the frequency domain, resulting in a spectrum.

[0032] The energy distribution features within a preset frequency band are extracted from the spectrum as spectral features. The preset frequency band includes a feature band used to distinguish between nasal breathing and mouth breathing.

[0033] It should be noted that this invention, by deploying a multi-microphone array to collect multi-channel audio signals, firstly achieves spatial sound field perception capability at the hardware level, laying the foundation for subsequent sound source localization. Then, by using the generalized cross-correlation function method or beamforming algorithm to calculate the direction angle of sound arrival, the system can intelligently determine the spatial source of breathing sounds, effectively distinguishing breathing sounds from the child's mouth and nose area from other irrelevant noise in the environment, significantly improving the anti-interference capability of target signal acquisition. Based on this, by setting a spatial range centered on the mouth and nose area as a judgment condition, this method achieves automatic and accurate screening of candidate breathing sound signals, ensuring the reliability of subsequent analysis objects. Finally, the filtered signals undergo frequency domain transformation and the energy distribution of characteristic frequency bands is extracted, transforming the physical characteristics of sound into quantifiable spectral features. This process enables the system to objectively distinguish breathing patterns based on the inherent differences between mouth breathing and nasal breathing in the acoustic spectrum, ultimately providing reliable dual evidence regarding the source and nature of breathing sounds—evidence lacking in traditional single audio analysis methods—for accurately judging functional open-mouth breathing.

[0034] Furthermore, the step of inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event occurring in the target child within a specific time period includes:

[0035] The multimodal features are input into a pre-trained mouth breathing recognition model. The weights of the multimodal features are dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model, and the judgment result of functional mouth breathing events of the target child within a specific time period is output.

[0036] It should be noted that this invention inputs multimodal features such as thermal imaging, radar micro-motion, and audio into a pre-trained recognition model, and dynamically allocates the weights of each feature using its internal attention fusion mechanism. This allows the model to simulate the comprehensive decision-making process of a clinician, meaning it does not mechanically rely on a single indicator, but intelligently adjusts the level of trust in different features based on the specific context of each breathing event (such as whether the thermal imaging signal is weakened due to occlusion by a blanket, or whether the audio signal is interfered with by environmental noise). This dynamic weighted fusion strategy greatly enhances the robustness of the system in complex home environments, enabling it to effectively cope with the challenge of single sensors being susceptible to interference and failure. Thus, it comprehensively and with high confidence judges whether the nasal airflow path, respiratory effort, and breath sound characteristics all point to functional mouth breathing caused by upper airway obstruction, ultimately significantly improving the accuracy and reliability of automatic recognition in real home scenarios.

[0037] Furthermore, the open-mouth breathing recognition model is a fusion model constructed based on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism; wherein, the spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result.

[0038] It should be noted that this invention employs a spatiotemporal graph convolutional network to model the spatiotemporal relationships between thermal imaging feature points. This allows the system to move beyond isolated analysis of local temperature changes and instead grasp the dynamic thermodynamic patterns of the nasal and oral regions during the respiratory cycle as a whole. This enables it to more robustly handle data fluctuations caused by transient interference or local occlusion from the external environment. Simultaneously, by combining a multi-head self-attention mechanism to dynamically weight and fuse different modal features from thermal imaging, radar micro-motion, and audio, the model can mimic the comprehensive decision-making thinking of clinical experts. For each specific respiratory event, it adaptively focuses on the feature cues that contribute the most (e.g., relying more on thermal imaging and radar features when audio signals are interfered with, and assigning higher weight to spectral features when it is necessary to distinguish the nature of breathing). This deep understanding of spatiotemporal context and intelligent balancing of the contribution of multimodal information greatly enhances the recognition model's contextual awareness and reasoning ability in real and changing home sleep environments. Ultimately, this achieves a more clinically relevant, high-precision, and robust automated judgment of functional open-mouth breathing events.

[0039] Furthermore, based on the judgment result, a mouth breathing monitoring report for the target child is generated, including:

[0040] Based on the judgment results, the monitoring indicators of mouth breathing events during the entire monitoring period are statistically analyzed. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of occurrence, and the correlation between mouth breathing events and sleep stages.

[0041] Based on the monitoring indicators of the mouth breathing events, a mouth breathing monitoring report for the target child is generated.

[0042] It should be noted that this invention, by systematically statistically analyzing the model's judgment results on individual respiratory events throughout the entire nightly monitoring period, elevates discrete event-level judgments into macroscopic indicators with clear clinical significance, including frequency of occurrence and average duration. This transforms the output from the identification of a single phenomenon to a holistic assessment of the child's overnight breathing pattern. Furthermore, by analyzing the correlation between mouth breathing events and different sleep stages, the report can reveal whether breathing problems are concentrated in specific sleep periods (such as deep sleep or REM sleep). This crucial information helps assess the severity and potential patterns of the problem. Ultimately, this structured monitoring report, encompassing macroscopic statistics and in-depth correlation analysis, provides doctors with high-value information far exceeding simple event counting, including time-based and pattern-based analysis. This greatly assists in the transformation of clinical diagnosis from qualitative judgment to data-driven decision-making, enhancing the professionalism and clinical usability of home monitoring results.

[0043] This specification provides one or more embodiments of a child mouth breathing monitoring device, comprising:

[0044] The acquisition unit uses a non-contact sensor array to simultaneously acquire multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0045] The extraction unit processes the multimodal physiological data and extracts multimodal features related to breathing. The multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0046] The output unit inputs the multimodal features into a pre-trained mouth breathing recognition model and outputs the judgment result of the target child's functional mouth breathing event within a specific time period.

[0047] The report generation unit generates an open-mouth breathing monitoring report for the target child based on the judgment result.

[0048] This specification provides one or more embodiments of a child mouth breathing monitoring device, comprising:

[0049] At least one processor and bus; and,

[0050] A memory communicatively connected to the at least one processor; wherein,

[0051] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0052] Multimodal physiological data of the target child during sleep is collected synchronously using a non-contact sensor array. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0053] The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0054] The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period.

[0055] Based on the judgment results, an open-mouth breathing monitoring report for the target child is generated.

[0056] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0057] Multimodal physiological data of the target child during sleep is collected synchronously using a non-contact sensor array. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0058] The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0059] The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period.

[0060] Based on the judgment results, an open-mouth breathing monitoring report for the target child is generated.

[0061] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0062] This invention replaces the invasive method of traditional PSG, which requires attaching sensors to the face, with non-contact multimodal sensing technologies (such as thermal imaging, radar, and audio). This fundamentally eliminates the interference of the device with children's sleep, making the monitoring process imperceptible. For the first time, it can obtain real and natural sleep breathing data of children in a home environment. On this basis, by fusing and analyzing the thermodynamic characteristics of the oral and nasal regions (quantifying airflow pathways), chest and abdominal breathing effort (reflecting respiratory resistance), and the sound source and spectral characteristics of breathing sounds (distinguishing sound sources and characteristics), this method achieves accurate judgment of "functional mouth breathing" (i.e., airflow actually passes through the oral cavity) rather than just the action of "opening the mouth." This overcomes the limitations of single sensors, which are susceptible to environmental interference and cannot distinguish the nature of airflow. Finally, this ability to operate comfortably at home for a long time and provide accurate judgment results makes home monitoring for weeks or even months possible. This provides doctors with long-term, dynamic data support that traditional single PSG tests cannot match in capturing intermittent or posture-dependent respiratory events and assessing the trend of disease progression, greatly improving the accuracy of clinical diagnosis and the timeliness of intervention. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0064] Figure 1 A flowchart illustrating a method for monitoring mouth breathing in children, provided for one or more embodiments of this specification;

[0065] Figure 2 A schematic diagram of a child mouth breathing monitoring device provided for one or more embodiments of this specification;

[0066] Figure 3 This is a schematic diagram of the structure of a child mouth breathing monitoring device provided for one or more embodiments of this specification. Detailed Implementation

[0067] This specification provides a method, device, equipment, and medium for monitoring children's open-mouth breathing through its embodiments.

[0068] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0069] Figure 1 This diagram illustrates a flowchart of a method for monitoring mouth breathing in children, provided for one or more embodiments of this specification. This process can be executed by a child mouth breathing monitoring system. Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0070] The method flow steps of the embodiments in this specification are as follows:

[0071] S101, using a non-contact sensor array to synchronously collect multimodal physiological data of the target child during sleep, the multimodal physiological data including at least facial thermal imaging video stream, chest and abdominal micro-motion signals and environmental audio signals.

[0072] In the embodiments described in this specification, this step is performed by a non-contact sensor array deployed in the child's bedroom. The sensor array comprises three core units: a thermal imaging camera, a millimeter-wave radar sensor, and a microphone array. To ensure strict time alignment of the data, all sensors are synchronously triggered by a unified master clock. At the start of data acquisition, the thermal imaging camera outputs a facial thermal imaging video stream at a fixed frame rate (e.g., several frames per second), covering the frontal facial area of ​​the child while sleeping; the millimeter-wave radar sensor continuously emits electromagnetic waves towards the child's chest and abdomen and receives the echoes, generating a raw baseband signal containing distance and phase information, which corresponds to the micro-movement signal of the chest and abdomen; the microphone array synchronously records ambient audio signals. All raw data are timestamped and packaged into a multimodal physiological dataset for one monitoring session, temporarily stored in the local device memory, in preparation for subsequent feature extraction.

[0073] S102, the multimodal physiological data is processed to extract multimodal features related to breathing. The multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdomen micro-motion signals, and sound source spatial location and spectral features of breathing sounds extracted from the environmental audio signals.

[0074] In the embodiments described in this specification, this step is performed on a local computing unit (such as an edge computing device) and is intended to extract effective respiratory-related features from raw data.

[0075] For thermodynamic feature extraction, facial thermal imaging video streams are processed. First, a pre-trained keypoint detection network automatically identifies and tracks the coordinates of the nostril center and labial fold center in each frame. Next, two fixed rectangular regions are dynamically defined centered on these coordinates, serving as the nasal cavity region of interest (ROI) and the oral cavity region of interest (ROI), respectively. Then, the average rate of temperature change within each region per unit time (e.g., one respiratory cycle) is calculated. Finally, the temperature change rate of the oral cavity region is divided by the temperature change rate of the nasal cavity region to obtain the quantified oral-nasal airflow ratio, which serves as the thermodynamic change feature.

[0076] For extracting the respiratory effort waveform, the micro-motion signals of the chest and abdomen (i.e., the raw radar baseband signal) are processed. First, the signal is processed using range-dimensional FFT, and specific distance cells corresponding to the distance between the chest and abdomen of the child are identified by energy peak detection. Then, the signal sequence within this cell is demodulated to obtain a phase change signal proportional to the body surface displacement. Finally, a bandpass filter covering the normal respiratory frequency range is applied to this phase signal to filter out high-frequency noise and low-frequency drift. The resulting clean, periodic waveform is the respiratory effort waveform.

[0077] For audio feature extraction, environmental audio signals are processed. First, the time difference of sound arrival at different microphones is calculated using the generalized cross-correlation function method, thereby estimating the direction of arrival angle of the sound and locating the spatial position of the sound source. The system presets a spatial range centered on the child's mouth and nose; only sound signals located within this range are identified as candidate breath sound signals. Subsequently, the candidate signal is framed, windowed, and subjected to short-time Fourier transform to generate a spectrogram. Finally, the energy distribution of preset characteristic frequency bands that can distinguish between nasal breathing (predominantly low frequencies) and mouth breathing (enhanced mid-to-high frequencies) is calculated from the spectrogram to obtain spectral features.

[0078] S103, the multimodal features are input into the pre-trained mouth breathing recognition model, and the judgment result of the target child having functional mouth breathing events within a specific time period is output.

[0079] In this embodiment of the specification, this step aligns all the features extracted in S102 (nasal-oral airflow ratio, breathing effort waveform, sound source location, and audio spectrum features) by time and combines them into a single feature vector, which is then input into a pre-trained mouth breathing recognition model. This model employs an attention fusion mechanism to dynamically evaluate the reliability of different features at the current moment and assign corresponding weights, thereby outputting a judgment result for each short time period (e.g., 30 seconds). This result clearly determines whether a functional mouth breathing event occurred within that time period and includes a confidence level.

[0080] S104, Based on the judgment result, generate an open-mouth breathing monitoring report for the target child.

[0081] In the embodiments described in this specification, this step integrates and analyzes the continuous judgment results output by S103. The system iterates through the judgment results of the entire monitoring period (e.g., throughout the night), and statistically analyzes key monitoring indicators of mouth breathing events, including total frequency, average duration, and correlation with different sleep stages. Finally, the system automatically fills these indicators into a structured template to generate a concise mouth breathing monitoring report, which allows parents or doctors to intuitively understand the child's breathing pattern throughout the night.

[0082] It should be noted that this invention, by employing non-contact multimodal sensing technologies (such as thermal imaging, radar, and audio) to replace the invasive method of traditional PSG that requires attaching sensors to the face, fundamentally eliminates the interference of the device on children's sleep, making the monitoring process imperceptible. This allows for the first-ever acquisition of children's real and natural sleep breathing data in a home environment. Furthermore, by fusing and analyzing the thermodynamic characteristics of the nasal and oral regions (quantifying airflow pathways), chest and abdominal respiratory effort (reflecting respiratory resistance), and the sound source and spectral characteristics of respiratory sounds (distinguishing sound sources and characteristics), this method achieves accurate judgment of "functional mouth breathing" (i.e., airflow actually passing through the oral cavity) rather than just the act of "opening the mouth." This overcomes the limitations of single sensors, which are susceptible to environmental interference and cannot distinguish the nature of airflow. Ultimately, this ability to operate comfortably and continuously at home and provide accurate results makes home monitoring for weeks or even months possible. This provides doctors with long-term, dynamic data support that traditional single PSG tests cannot match, enabling the capture of intermittent or posture-dependent respiratory events and the assessment of disease progression, greatly improving the accuracy of clinical diagnosis and the timeliness of intervention.

[0083] Furthermore, the thermodynamic change feature is the nasal-oral airflow ratio. When extracting the thermodynamic change features of the nasal and oral regions based on the facial thermal imaging video stream, the center points of the nostrils and the labial folds in the thermal imaging video stream can be located based on a key point detection network; the nasal cavity region of interest is dynamically delineated based on the center point of the nostrils; the oral cavity region of interest is dynamically delineated based on the center point of the labial folds; a first average temperature change rate per unit time of the nasal cavity region of interest is determined; a second average temperature change rate per unit time of the oral cavity region of interest is determined; and the nasal-oral airflow ratio is determined based on the ratio of the first average temperature change rate to the second average temperature change rate.

[0084] It should be noted that this invention primarily relies on a keypoint detection network pre-trained on a large amount of thermal imaging facial data. This network is deployed on a local computing unit to process the acquired thermal imaging video stream frame by frame. For each frame of thermal imaging image, the network automatically outputs the pixel coordinates of the center points of the nostrils and the lip creases.

[0085] After successfully locating key points, the system dynamically defines the analysis area based on these coordinates. Specifically, the system uses the center point of the nostrils in each frame as a reference to define a fixed-size rectangular or circular area around it as the nasal cavity region of interest. Similarly, a region of interest for the oral cavity is defined based on the center point of the labial folds in each frame. This dynamic definition method based on key point coordinates ensures that even if the child makes slight head movements during sleep, these two regions of interest will always accurately cover the nostrils and lip areas, guaranteeing the continuity of data acquisition.

[0086] Next, the system will perform time-series analysis on the temperature data in these two regions. It will calculate the average rate of temperature change for all pixels within the nasal cavity region of interest over a unit of time (e.g., one complete respiratory cycle), which is defined as the first average rate of temperature change. Similarly, the second average rate of temperature change will be calculated within the oral cavity region of interest. Since exhaled airflow is warmer and inhaled airflow is cooler, this rate of temperature change directly reflects the strength of the respiratory airflow.

[0087] Finally, the system divides the second average temperature change rate (reflecting oral airflow) by the first average temperature change rate (reflecting nasal airflow) to calculate the nasal airflow ratio. This ratio is a normalized quantitative indicator: when the ratio is close to 0, it indicates that almost all airflow passes through the nasal cavity; when the ratio is significantly greater than 1, it indicates that the airflow mainly passes through the oral cavity, thus achieving an objective and quantitative judgment of the breathing pathway.

[0088] It should be noted that this invention utilizes thermal imaging video streams to non-contactly monitor temperature changes in the nasal and oral regions, and innovatively proposes the quantitative indicator of "nasal-oral airflow ratio." This fundamentally avoids the discomfort and interference caused by attaching sensors to children's faces, achieving imperceptible physiological signal acquisition. Furthermore, by dynamically locating the center of the nostrils and lip creases and defining the region of interest through a key point detection network, this method can adapt to the slight head movements during children's sleep, ensuring the continuity and stability of monitoring. Most importantly, by calculating the average temperature change rate per unit time in the nasal and oral regions and using their ratio as the basis for judgment, this method transforms traditional qualitative and subjective observations (such as whether the mouth is open) into quantitative and objective measurements of airflow pathways. This effectively distinguishes between habitual mouth breathing and true functional mouth breathing caused by nasal congestion. This key advancement makes it possible to obtain accurate judgments with clear clinical value comparable to polysomnography in natural sleep scenarios such as at home.

[0089] Furthermore, when extracting the respiratory effort waveform based on the chest and abdomen micro-motion signals, a frequency-modulated continuous wave radar sensor deployed in the monitoring area can continuously emit electromagnetic waves to the child's chest and abdomen and receive their echoes to obtain the original baseband signal; the original baseband signal is processed by fast Fourier transform to filter out the specific range cell with the highest signal-to-noise ratio; the signal sequence in the specific range cell is phase-demodulated to obtain a phase change signal proportional to the displacement of the chest and abdomen surface; a bandpass filter with a specified passband frequency is applied to the phase change signal to separate the respiratory effort waveform caused by breathing.

[0090] It should be noted that the implementation of this scheme begins with the deployment of a frequency-modulated continuous wave radar sensor. This sensor is placed within the monitoring area, directly facing the child's chest and abdomen. During its operation, it continuously emits electromagnetic waves with linearly varying frequencies and receives the echoes reflected from the child's body surface. By mixing the transmitted and received signals, the raw baseband signal containing range and phase information is obtained.

[0091] Next, the system performs a Fast Fourier Transform (FFT) on the original baseband signal. This process converts the signal from the time-amplitude domain to the range-amplitude domain (i.e., range-dimensional FFT), separating the reflected signals from targets at different distances in the spectrum. The system then scans the entire spectrum, automatically identifying and selecting the specific range cell with the strongest signal energy and highest signal-to-noise ratio. This cell corresponds to the precise distance between the child's chest / abdomen and the radar sensor, thus achieving range-dimensional focusing of the monitored target and effectively eliminating interference from other stationary objects and clutter indoors.

[0092] After locking onto a specific distance cell, the system performs phase demodulation processing on the signal sequence within that cell. Since minute changes in the propagation path of electromagnetic waves are sensitively reflected in the echo phase, and respiratory movements of the chest and abdomen cause periodic changes in the propagation distance, a phase change signal proportional to the displacement of the chest and abdomen surface can be obtained by calculating the phase information.

[0093] Finally, the system applies a bandpass filter with a specified passband frequency to the demodulated phase change signal. The passband frequency range of this filter is set to cover the typical frequencies of normal human breathing. Through this filtering process, the low-frequency periodic displacement components generated by breathing in the signal are preserved and enhanced, while higher-frequency heartbeat micro-movements, random body swaying, and low-frequency DC components and noise are effectively filtered out. The clean, smooth periodic waveform output after filtering is the respiratory effort waveform, which characterizes the degree of respiratory exertion.

[0094] It should be noted that this invention employs a frequency-modulated continuous wave radar sensor to non-contactly detect subtle movements of the chest and abdomen. This firstly achieves precise perception of body surface displacement during respiration, completely avoiding the restrictive feeling associated with wearable sensors. Secondly, by performing a fast Fourier transform on the radar echo signal and intelligently selecting the specific distance unit with the highest signal-to-noise ratio, this method effectively eliminates environmental clutter interference, stably locking the monitoring target onto the child's chest and abdomen, ensuring the accuracy and reliability of the signal source. Furthermore, by performing phase demodulation on the signal from the specific distance unit, minute displacement information is extracted from the echo signal with high precision. Then, a bandpass filter removes large body movements and high-frequency noise unrelated to respiration, ultimately separating a clean, continuous periodic waveform representing respiratory effort. This complete technical chain from radar signal to respiratory effort waveform makes it possible to continuously and stably acquire key physiological indicators reflecting the patency of the upper airway without contact with the child's body. This provides indispensable objective quantitative evidence for accurately determining whether functional mouth breathing, requiring additional respiratory effort due to nasal obstruction, exists.

[0095] Furthermore, when extracting the spatial location and spectral features of the respiratory sound source based on the environmental audio signal, a multi-microphone array deployed in the monitoring area can be used to simultaneously collect multi-channel audio signals from the environment; the direction of arrival angle of the multi-channel audio signal can be determined by the generalized cross-correlation function method or beamforming algorithm, thereby locating the spatial location of the respiratory sound source; it is determined whether the spatial location of the sound source is within the spatial range centered on the child's mouth and nose area; if so, the multi-channel audio signal is determined to be a candidate respiratory sound signal; the candidate respiratory sound signal is processed (frame segmentation, windowing, and short-time Fourier transform) to convert the candidate respiratory sound signal from the time domain to the frequency domain to obtain a spectrum; the energy distribution features within a preset frequency band are extracted from the spectrum as spectral features, and the preset frequency band includes the feature band used to distinguish between nasal breathing and mouth breathing.

[0096] It should be noted that this invention primarily relies on a multi-microphone array correctly deployed in the monitoring area. Each microphone in this array synchronously acquires ambient sound, generating a synchronized multi-channel audio signal.

[0097] Subsequently, the system uses either the generalized cross-correlation function method or beamforming algorithm to process these synchronization signals. The core purpose of both algorithms is to estimate the direction of arrival angle of the sound by calculating the time difference between the arrival of the sound at different microphones, thereby locating the spatial position of the sound source of the breathing sound.

[0098] Next, the system compares the calculated sound source location with a preset spatial range centered on the child's mouth and nose area. This preset range is typically pre-set based on the sensor's installation location and viewing angle. Only when the sound source is determined to be within this specific spatial range will the corresponding audio segment be recognized by the system as a candidate breath sound signal. This step effectively filters out interference from the environment or other sound sources.

[0099] For successfully identified candidate breath sound signals, the system performs standard audio signal processing. First, it performs frame segmentation, dividing the continuous signal into short time segments; then, it performs windowing on each frame to reduce spectral leakage; finally, it performs a short-time Fourier transform on the windowed signal, converting it from the time domain to the frequency domain, to obtain a spectrum that displays the frequency components changing over time.

[0100] Ultimately, from the spectrogram, the system focuses on analyzing the characteristic frequency bands that are pre-defined based on acoustic knowledge and can effectively distinguish between nasal and mouth breathing. By calculating the energy distribution within these specific frequency bands (e.g., the energy ratio or concentration of different frequency bands), spectral features for pattern recognition are extracted.

[0101] It should be noted that this invention, by deploying a multi-microphone array to collect multi-channel audio signals, firstly achieves spatial sound field perception capability at the hardware level, laying the foundation for subsequent sound source localization. Then, by using the generalized cross-correlation function method or beamforming algorithm to calculate the direction angle of sound arrival, the system can intelligently determine the spatial source of breathing sounds, effectively distinguishing breathing sounds from the child's mouth and nose area from other irrelevant noise in the environment, significantly improving the anti-interference capability of target signal acquisition. Based on this, by setting a spatial range centered on the mouth and nose area as a judgment condition, this method achieves automatic and accurate screening of candidate breathing sound signals, ensuring the reliability of subsequent analysis objects. Finally, the filtered signals undergo frequency domain transformation and the energy distribution of characteristic frequency bands is extracted, transforming the physical characteristics of sound into quantifiable spectral features. This process enables the system to objectively distinguish breathing patterns based on the inherent differences between mouth breathing and nasal breathing in the acoustic spectrum, ultimately providing reliable dual evidence regarding the source and nature of breathing sounds—evidence lacking in traditional single audio analysis methods—for accurately judging functional open-mouth breathing.

[0102] Furthermore, when inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event of the target child within a specific time period, the multimodal features can be input into the pre-trained mouth breathing recognition model, and the weights of the multimodal features can be dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model to output a judgment result on the functional mouth breathing event of the target child within a specific time period.

[0103] It should be noted that this invention is based on a pre-trained open-mouth breathing recognition model that has already been trained using a large amount of labeled multimodal data. During operation, the system aligns and combines multiple features from previous steps—namely, the mouth-nose airflow ratio extracted from thermal imaging data, the breathing effort waveform extracted from radar signals, and the sound source location and spectral features extracted from audio signals—according to a unified timestamp, forming a complete multimodal feature vector, which is then input into the model.

[0104] The core processing element within the model is the attention fusion mechanism. This mechanism automatically analyzes the current input feature vector, evaluating the reliability and importance of each feature dimension at the current moment. For example, when an audio signal is briefly disturbed by noise, the mechanism reduces the weight of audio-related features; when a thermal imaging signal is clear and stable, its weight is increased accordingly. Through this dynamic allocation, the model can simulate the thought process of an expert making a comprehensive judgment, rather than mechanically treating all features equally.

[0105] Finally, based on these weighted feature information, the model performs comprehensive calculations and inferences to output a judgment result for a specific time period (such as a time window). This result is usually a classification conclusion (e.g., "a functional mouth breathing event occurred" or "no event occurred"), and may be accompanied by a confidence score indicating the degree of certainty in the judgment.

[0106] It should be noted that this invention inputs multimodal features such as thermal imaging, radar micro-motion, and audio into a pre-trained recognition model, and dynamically allocates the weights of each feature using its internal attention fusion mechanism. This allows the model to simulate the comprehensive decision-making process of a clinician, meaning it does not mechanically rely on a single indicator, but intelligently adjusts the level of trust in different features based on the specific context of each breathing event (such as whether the thermal imaging signal is weakened due to occlusion by a blanket, or whether the audio signal is interfered with by environmental noise). This dynamic weighted fusion strategy greatly enhances the robustness of the system in complex home environments, enabling it to effectively cope with the challenge of single sensors being susceptible to interference and failure. Thus, it comprehensively and with high confidence judges whether the nasal airflow path, respiratory effort, and breath sound characteristics all point to functional mouth breathing caused by upper airway obstruction, ultimately significantly improving the accuracy and reliability of automatic recognition in real home scenarios.

[0107] Furthermore, the open-mouth breathing recognition model is a fusion model constructed based on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism; wherein, the spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result.

[0108] It should be noted that, firstly, the spatiotemporal graph convolutional network component is used to process thermal imaging feature point data. This network models the sequence of key points (such as the center points of the nostrils and the lip creases) extracted from the thermal imaging video stream as a graph structure, where nodes represent feature points and edges represent the spatial connections between points. The network captures the spatial dependencies between feature points through graph convolution operations, and simultaneously uses temporal convolution to analyze the change patterns of feature points across different time frames, thereby learning the spatiotemporal characteristics of the thermodynamic changes in the nasal and oral regions during respiration.

[0109] Subsequently, a multi-head self-attention mechanism component was used to fuse multimodal features. This mechanism takes as input the thermodynamic features output from the spatiotemporal graph convolutional network, the breathing effort waveform features extracted directly from radar signals, and the sound source location and spectral features extracted from audio signals. Through multiple parallel attention heads, the mechanism calculates the importance weight of each feature to the final judgment, dynamically adjusting the contribution of different features (e.g., reducing the weight of audio signals when they are interfered with, and increasing the weight of thermal imaging signals when they are clear), thus achieving adaptive feature fusion across modalities.

[0110] Finally, the fused features are fed into a fully connected layer or classifier, which outputs a judgment on whether the target child experienced functional mouth breathing events within a specific time period. The entire model learns parameters using labeled data during the training phase and is directly used for inference during the deployment phase, without the need for real-time adjustments.

[0111] It should be noted that this invention employs a spatiotemporal graph convolutional network to model the spatiotemporal relationships between thermal imaging feature points. This allows the system to move beyond isolated analysis of local temperature changes and instead grasp the dynamic thermodynamic patterns of the nasal and oral regions during the respiratory cycle as a whole. This enables it to more robustly handle data fluctuations caused by transient interference or local occlusion from the external environment. Simultaneously, by combining a multi-head self-attention mechanism to dynamically weight and fuse different modal features from thermal imaging, radar micro-motion, and audio, the model can mimic the comprehensive decision-making thinking of clinical experts. For each specific respiratory event, it adaptively focuses on the feature cues that contribute the most (e.g., relying more on thermal imaging and radar features when audio signals are interfered with, and assigning higher weight to spectral features when it is necessary to distinguish the nature of breathing). This deep understanding of spatiotemporal context and intelligent balancing of the contribution of multimodal information greatly enhances the recognition model's contextual awareness and reasoning ability in real and changing home sleep environments. Ultimately, this achieves a more clinically relevant, high-precision, and robust automated judgment of functional open-mouth breathing events.

[0112] Furthermore, when generating the mouth breathing monitoring report for the target child based on the judgment result, the monitoring indicators of mouth breathing events throughout the entire monitoring period can be statistically analyzed based on the judgment result. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of mouth breathing events, and the correlation between mouth breathing events and sleep stages. Based on the monitoring indicators of mouth breathing events, the mouth breathing monitoring report for the target child is generated.

[0113] It should be noted that this step begins with the integration and analysis of the chronological sequence of judgment results output from the preceding step (S103). The system first iterates through all results throughout the entire monitoring period (e.g., overnight) to identify all time periods marked as "occurring functional mouth breathing events".

[0114] The system calculates the total number of mouth breathing events. Typically, for standardization, the average number of events per unit time (e.g., per hour) is further calculated, i.e., the frequency of occurrence. For each identified mouth breathing event, the system calculates the duration of the single event based on its start and end timestamps. Subsequently, the average duration of all events and the percentage of total mouth breathing duration to total recorded time throughout the night are calculated. The system temporally matches the occurrence time of mouth breathing events with sleep stages (e.g., light sleep, deep sleep, REM sleep) estimated through other analytical methods (e.g., sleep staging algorithms based on radar micro-motion signals or body movement signals). By analyzing the frequency or duration percentage of events in different sleep stages, the correlation between mouth breathing events and sleep stages is determined.

[0115] It should be noted that this invention, by systematically statistically analyzing the model's judgment results on individual respiratory events throughout the entire nightly monitoring period, elevates discrete event-level judgments into macroscopic indicators with clear clinical significance, including frequency of occurrence and average duration. This transforms the output from the identification of a single phenomenon to a holistic assessment of the child's overnight breathing pattern. Furthermore, by analyzing the correlation between mouth breathing events and different sleep stages, the report can reveal whether breathing problems are concentrated in specific sleep periods (such as deep sleep or REM sleep). This crucial information helps assess the severity and potential patterns of the problem. Ultimately, this structured monitoring report, encompassing macroscopic statistics and in-depth correlation analysis, provides doctors with high-value information far exceeding simple event counting, including time-based and pattern-based analysis. This greatly assists in the transformation of clinical diagnosis from qualitative judgment to data-driven decision-making, enhancing the professionalism and clinical usability of home monitoring results.

[0116] Figure 2 This specification provides a schematic diagram of the structure of a child mouth breathing monitoring device according to one or more embodiments, including: a data acquisition unit 201, an extraction unit 202, an output unit 203, and a report generation unit 204.

[0117] The acquisition unit 201 uses a non-contact sensor array to synchronously acquire multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0118] Extraction unit 202 processes the multimodal physiological data and extracts multimodal features related to breathing. The multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdomen micro-motion signals, and sound source spatial location and spectral features of breathing sounds extracted from the environmental audio signals.

[0119] The output unit 203 inputs the multimodal features into a pre-trained mouth breathing recognition model and outputs the judgment result of the target child's functional mouth breathing event within a specific time period.

[0120] The report generation unit 204 generates an open-mouth breathing monitoring report for the target child based on the judgment result.

[0121] Figure 3 A schematic diagram of a child mouth breathing monitoring device provided for one or more embodiments of this specification includes:

[0122] At least one processor and bus; and,

[0123] A memory communicatively connected to the at least one processor; wherein,

[0124] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0125] Multimodal physiological data of the target child during sleep is collected synchronously using a non-contact sensor array. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0126] The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0127] The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period.

[0128] Based on the judgment results, an open-mouth breathing monitoring report for the target child is generated.

[0129] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following:

[0130] Multimodal physiological data of the target child during sleep is collected synchronously using a non-contact sensor array. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals.

[0131] The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals.

[0132] The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period.

[0133] Based on the judgment results, an open-mouth breathing monitoring report for the target child is generated.

[0134] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0135] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0136] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0137] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The aforementioned units can be implemented in hardware or software.

[0140] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0141] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for monitoring mouth breathing in children, characterized in that, The method includes: The non-contact sensor array synchronously collects multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals. The sensor array includes a thermal imaging camera, a millimeter-wave radar sensor, and a microphone array. All sensors are synchronously triggered by a unified master control clock. The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals. The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period. Based on the judgment results, a mouth breathing monitoring report for the target child is generated; The thermodynamic change feature is the airflow ratio between the mouth and nose. The extraction of the thermodynamic change features of the mouth and nose region based on the facial thermal imaging video stream includes: The center point of the nostril and the center point of the lip crease in the thermal imaging video stream are located based on the key point detection network. Using the center point of the nostril as a reference, the region of interest in the nasal cavity is dynamically defined; Using the center point of the labial suture as a reference, the region of interest in the oral cavity is dynamically defined; Determine the first average temperature change rate per unit time in the region of interest of the nasal cavity; Determine the second average temperature change rate per unit time in the oral cavity region of interest; The airflow ratio between the mouth and nose is determined based on the ratio of the first average temperature change rate to the second average temperature change rate. When the ratio is close to 0, it indicates that the airflow passes through the nasal cavity; when the ratio is greater than 1, it indicates that the airflow passes through the oral cavity. The extraction of respiratory effort waveform based on the chest and abdominal micromotion signals includes: By using frequency-modulated continuous wave radar sensors deployed in the monitoring area, electromagnetic waves are continuously emitted toward the child's chest and abdomen and their echoes are received to obtain the raw baseband signal. The original baseband signal is processed by Fast Fourier Transform to select the specific distance cell with the highest signal-to-noise ratio; Phase demodulation is performed on the signal sequence within the specific distance unit to obtain a phase change signal that is proportional to the displacement of the chest and abdominal surface; A bandpass filter with a specified passband frequency is applied to the phase change signal to separate the breathing effort waveform caused by breathing, the breathing effort waveform representing the degree of breathing effort; The extraction of the sound source spatial location and spectral features of breathing sounds based on the environmental audio signal includes: Multi-channel audio signals from the environment are simultaneously acquired by deploying a multi-microphone array in the monitoring area; Determine the direction angle of arrival of the multi-channel audio signal to locate the spatial position of the sound source of the breathing sound; Determine whether the spatial location of the sound source is within a spatial range centered on the child's mouth and nose area; If so, the multi-channel audio signal is determined to be a candidate breath sound signal; The candidate breath sound signals are processed to convert them from the time domain to the frequency domain, resulting in a spectrum. The energy distribution features within a preset frequency band are extracted from the spectrum as spectral features, and the preset frequency band includes a feature band used to distinguish between nasal breathing and mouth breathing. The step of inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event of the target child within a specific time period includes: The multimodal features are input into a pre-trained mouth breathing recognition model. The weights of the multimodal features are dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model, and the judgment result of functional mouth breathing events of the target child within a specific time period is output. The open-mouth breathing recognition model is a fusion model built on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism. The spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result. Generate an open-mouth breathing monitoring report for the target child, including: Based on the judgment results, the monitoring indicators of mouth breathing events are statistically analyzed throughout the monitoring period. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of occurrence, and the correlation between mouth breathing events and sleep stages. The occurrence time of mouth breathing events is matched temporally with the sleep stages estimated by the sleep staging algorithm through radar micro-motion signals or body movement signals analysis. The correlation between mouth breathing events and sleep stages is determined by analyzing the proportion of occurrence frequency or duration of mouth breathing events in different sleep stages. Based on the monitoring indicators of the mouth breathing events, a mouth breathing monitoring report for the target child is generated.

2. A device for monitoring mouth breathing in children, characterized in that, include: The acquisition unit uses a non-contact sensor array to synchronously acquire multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals. The sensor array includes a thermal imaging camera, a millimeter-wave radar sensor, and a microphone array. All sensors are synchronously triggered by a unified master control clock. The extraction unit processes the multimodal physiological data and extracts multimodal features related to breathing. The multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals. The output unit inputs the multimodal features into a pre-trained mouth breathing recognition model and outputs the judgment result of the target child's functional mouth breathing event within a specific time period. The report generation unit generates a mouth breathing monitoring report for the target child based on the judgment result; The thermodynamic change feature is the airflow ratio between the mouth and nose. The extraction of the thermodynamic change features of the mouth and nose region based on the facial thermal imaging video stream includes: The center point of the nostril and the center point of the lip crease in the thermal imaging video stream are located based on the key point detection network. Using the center point of the nostril as a reference, the region of interest in the nasal cavity is dynamically defined; Using the center point of the labial suture as a reference, the region of interest in the oral cavity is dynamically defined; Determine the first average temperature change rate per unit time in the region of interest of the nasal cavity; Determine the second average temperature change rate per unit time in the oral cavity region of interest; The airflow ratio between the mouth and nose is determined based on the ratio of the first average temperature change rate to the second average temperature change rate. When the ratio is close to 0, it indicates that the airflow passes through the nasal cavity; when the ratio is greater than 1, it indicates that the airflow passes through the oral cavity. The extraction of respiratory effort waveform based on the chest and abdominal micromotion signals includes: By using frequency-modulated continuous wave radar sensors deployed in the monitoring area, electromagnetic waves are continuously emitted toward the child's chest and abdomen and their echoes are received to obtain the raw baseband signal. The original baseband signal is processed by Fast Fourier Transform to select the specific distance cell with the highest signal-to-noise ratio; Phase demodulation is performed on the signal sequence within the specific distance unit to obtain a phase change signal that is proportional to the displacement of the chest and abdominal surface; A bandpass filter with a specified passband frequency is applied to the phase change signal to separate the breathing effort waveform caused by breathing, the breathing effort waveform representing the degree of breathing effort; The extraction of the sound source spatial location and spectral features of breathing sounds based on the environmental audio signal includes: Multi-channel audio signals from the environment are simultaneously acquired by deploying a multi-microphone array in the monitoring area; Determine the direction angle of arrival of the multi-channel audio signal to locate the spatial position of the sound source of the breathing sound; Determine whether the spatial location of the sound source is within a spatial range centered on the child's mouth and nose area; If so, the multi-channel audio signal is determined to be a candidate breath sound signal; The candidate breath sound signals are processed to convert them from the time domain to the frequency domain, resulting in a spectrum. The energy distribution features within a preset frequency band are extracted from the spectrum as spectral features, and the preset frequency band includes a feature band used to distinguish between nasal breathing and mouth breathing. The step of inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event of the target child within a specific time period includes: The multimodal features are input into a pre-trained mouth breathing recognition model. The weights of the multimodal features are dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model, and the judgment result of functional mouth breathing events of the target child within a specific time period is output. The open-mouth breathing recognition model is a fusion model built on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism. The spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result. Generate an open-mouth breathing monitoring report for the target child, including: Based on the judgment results, the monitoring indicators of mouth breathing events are statistically analyzed throughout the monitoring period. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of occurrence, and the correlation between mouth breathing events and sleep stages. The occurrence time of mouth breathing events is matched temporally with the sleep stages estimated by the sleep staging algorithm through radar micro-motion signals or body movement signals analysis. The correlation between mouth breathing events and sleep stages is determined by analyzing the proportion of occurrence frequency or duration of mouth breathing events in different sleep stages. Based on the monitoring indicators of the mouth breathing events, a mouth breathing monitoring report for the target child is generated.

3. A device for monitoring mouth breathing in children, characterized in that, include: At least one processor and bus; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The non-contact sensor array synchronously collects multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals. The sensor array includes a thermal imaging camera, a millimeter-wave radar sensor, and a microphone array. All sensors are synchronously triggered by a unified master control clock. The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals. The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period. Based on the judgment results, a mouth breathing monitoring report for the target child is generated; The thermodynamic change feature is the airflow ratio between the mouth and nose. The extraction of the thermodynamic change features of the mouth and nose region based on the facial thermal imaging video stream includes: The center point of the nostril and the center point of the lip crease in the thermal imaging video stream are located based on the key point detection network. Using the center point of the nostril as a reference, the region of interest in the nasal cavity is dynamically defined; Using the center point of the labial suture as a reference, the region of interest in the oral cavity is dynamically defined; Determine the first average temperature change rate per unit time in the region of interest of the nasal cavity; Determine the second average temperature change rate per unit time in the oral cavity region of interest; The airflow ratio between the mouth and nose is determined based on the ratio of the first average temperature change rate to the second average temperature change rate. When the ratio is close to 0, it indicates that the airflow passes through the nasal cavity; when the ratio is greater than 1, it indicates that the airflow passes through the oral cavity. The extraction of respiratory effort waveform based on the chest and abdominal micromotion signals includes: By using frequency-modulated continuous wave radar sensors deployed in the monitoring area, electromagnetic waves are continuously emitted toward the child's chest and abdomen and their echoes are received to obtain the raw baseband signal. The original baseband signal is processed by Fast Fourier Transform to select the specific distance cell with the highest signal-to-noise ratio; Phase demodulation is performed on the signal sequence within the specific distance unit to obtain a phase change signal that is proportional to the displacement of the chest and abdominal surface; A bandpass filter with a specified passband frequency is applied to the phase change signal to separate the breathing effort waveform caused by breathing, the breathing effort waveform representing the degree of breathing effort; The extraction of the sound source spatial location and spectral features of breathing sounds based on the environmental audio signal includes: Multi-channel audio signals from the environment are simultaneously acquired by deploying a multi-microphone array in the monitoring area; Determine the direction angle of arrival of the multi-channel audio signal to locate the spatial position of the sound source of the breathing sound; Determine whether the spatial location of the sound source is within a spatial range centered on the child's mouth and nose area; If so, the multi-channel audio signal is determined to be a candidate breath sound signal; The candidate breath sound signals are processed to convert them from the time domain to the frequency domain, resulting in a spectrum. The energy distribution features within a preset frequency band are extracted from the spectrum as spectral features, and the preset frequency band includes a feature band used to distinguish between nasal breathing and mouth breathing. The step of inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event of the target child within a specific time period includes: The multimodal features are input into a pre-trained mouth breathing recognition model. The weights of the multimodal features are dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model, and the judgment result of functional mouth breathing events of the target child within a specific time period is output. The open-mouth breathing recognition model is a fusion model built on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism. The spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result. Generate an open-mouth breathing monitoring report for the target child, including: Based on the judgment results, the monitoring indicators of mouth breathing events are statistically analyzed throughout the monitoring period. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of occurrence, and the correlation between mouth breathing events and sleep stages. The occurrence time of mouth breathing events is matched temporally with the sleep stages estimated by the sleep staging algorithm through radar micro-motion signals or body movement signals analysis. The correlation between mouth breathing events and sleep stages is determined by analyzing the proportion of occurrence frequency or duration of mouth breathing events in different sleep stages. Based on the monitoring indicators of the mouth breathing events, a mouth breathing monitoring report for the target child is generated.

4. A non-volatile computer storage medium, characterized in that, It stores computer-executable instructions, which, when executed by a computer, can achieve the following: The non-contact sensor array synchronously collects multimodal physiological data of the target child during sleep. The multimodal physiological data includes at least facial thermal imaging video stream, chest and abdominal micro-motion signals, and environmental audio signals. The sensor array includes a thermal imaging camera, a millimeter-wave radar sensor, and a microphone array. All sensors are synchronously triggered by a unified master control clock. The multimodal physiological data is processed to extract multimodal features related to breathing. These multimodal features include thermodynamic change features of the mouth and nose region extracted from the facial thermal imaging video stream, respiratory effort waveform extracted from the chest and abdominal micro-motion signals, and sound source spatial location and spectral features of respiratory sounds extracted from the environmental audio signals. The multimodal features are input into a pre-trained mouth breathing recognition model, which outputs a judgment result on the functional mouth breathing event that occurred in the target child within a specific time period. Based on the judgment results, a mouth breathing monitoring report for the target child is generated; The thermodynamic change feature is the airflow ratio between the mouth and nose. The extraction of the thermodynamic change features of the mouth and nose region based on the facial thermal imaging video stream includes: The center point of the nostril and the center point of the lip crease in the thermal imaging video stream are located based on the key point detection network. Using the center point of the nostril as a reference, the region of interest in the nasal cavity is dynamically defined; Using the center point of the labial suture as a reference, the region of interest in the oral cavity is dynamically defined; Determine the first average temperature change rate per unit time in the region of interest of the nasal cavity; Determine the second average temperature change rate per unit time in the oral cavity region of interest; The airflow ratio between the mouth and nose is determined based on the ratio of the first average temperature change rate to the second average temperature change rate. When the ratio is close to 0, it indicates that the airflow passes through the nasal cavity; when the ratio is greater than 1, it indicates that the airflow passes through the oral cavity. The extraction of respiratory effort waveform based on the chest and abdominal micromotion signals includes: By using frequency-modulated continuous wave radar sensors deployed in the monitoring area, electromagnetic waves are continuously emitted toward the child's chest and abdomen and their echoes are received to obtain the raw baseband signal. The original baseband signal is processed by Fast Fourier Transform to select the specific distance cell with the highest signal-to-noise ratio; Phase demodulation is performed on the signal sequence within the specific distance unit to obtain a phase change signal that is proportional to the displacement of the chest and abdominal surface; A bandpass filter with a specified passband frequency is applied to the phase change signal to separate the breathing effort waveform caused by breathing, the breathing effort waveform representing the degree of breathing effort; The extraction of the sound source spatial location and spectral features of breathing sounds based on the environmental audio signal includes: Multi-channel audio signals from the environment are simultaneously acquired by deploying a multi-microphone array in the monitoring area; Determine the direction angle of arrival of the multi-channel audio signal to locate the spatial position of the sound source of the breathing sound; Determine whether the spatial location of the sound source is within a spatial range centered on the child's mouth and nose area; If so, the multi-channel audio signal is determined to be a candidate breath sound signal; The candidate breath sound signals are processed to convert them from the time domain to the frequency domain, resulting in a spectrum. The energy distribution features within a preset frequency band are extracted from the spectrum as spectral features, and the preset frequency band includes a feature band used to distinguish between nasal breathing and mouth breathing. The step of inputting the multimodal features into a pre-trained mouth breathing recognition model and outputting a judgment result on the functional mouth breathing event of the target child within a specific time period includes: The multimodal features are input into a pre-trained mouth breathing recognition model. The weights of the multimodal features are dynamically allocated through the attention fusion mechanism in the mouth breathing recognition model, and the judgment result of functional mouth breathing events of the target child within a specific time period is output. The open-mouth breathing recognition model is a fusion model built on a spatiotemporal graph convolutional network and a multi-head self-attention mechanism. The spatiotemporal graph convolutional network is used to model the spatiotemporal relationship between thermal imaging feature points, and the multi-head self-attention mechanism is used to dynamically evaluate and fuse the contribution weights of thermodynamic features, breathing effort waveform features and audio features to the final judgment result. Generate an open-mouth breathing monitoring report for the target child, including: Based on the judgment results, the monitoring indicators of mouth breathing events are statistically analyzed throughout the monitoring period. The monitoring indicators of mouth breathing events include the frequency of occurrence of mouth breathing events, the duration of occurrence, and the correlation between mouth breathing events and sleep stages. The occurrence time of mouth breathing events is matched temporally with the sleep stages estimated by the sleep staging algorithm through radar micro-motion signals or body movement signals analysis. The correlation between mouth breathing events and sleep stages is determined by analyzing the proportion of occurrence frequency or duration of mouth breathing events in different sleep stages. Based on the monitoring indicators of the mouth breathing events, a mouth breathing monitoring report for the target child is generated.

Citation Information

Patent Citations

  • Non-contact oral respiration detection device and method and storage medium

    CN111568388A

  • Sleep respiration detection method and system for apnea syndrome, and cloud platform

    CN120419911A

  • Respiration diagnostic device and respiration diagnostic program

    JP2018023453A

  • Systems and methods for contactless sleep monitoring

    US20210177343A1