Non-intrusive non-interference sleep breathing state monitoring method and non-intrusive non-interference sleep breathing state monitoring system
By using non-invasive video acquisition and artificial intelligence algorithms, convenient, comfortable, and low-interference sleep breathing monitoring is achieved in a natural sleep environment, solving the problems of cumbersome operation and high interference of existing equipment, and providing accurate physiological parameter analysis and evaluation results.
Patent Information
- Application Number
- CN202511779515.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-01-23
AI Technical Summary
Existing sleep breathing monitoring devices are expensive, cumbersome to operate, and highly interfering. They are difficult to reflect the true physiological situation in a natural sleep environment and lack privacy protection and effective monitoring solutions for sleep breathing physiological activities.
Using a non-invasive video acquisition module, high-definition video segmentation and feature extraction, combined with artificial intelligence algorithms, it can monitor physiological activities during sleep, construct a multi-dimensional sleep state assessment model, and conduct non-intrusive and privacy-protected sleep breathing state monitoring.
It enables convenient, comfortable, and low-interference monitoring of sleep breathing in a natural sleep environment, providing accurate physiological parameter analysis and evaluation results, and is suitable for large-scale screening of sleep breathing-related physiological states.
Smart Images

Figure CN121370071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a non-invasive and interference-free method and system for monitoring sleep breathing. Background Technology
[0002] Sleep-disordered breathing refers to abnormal sleep states and breathing during sleep, which seriously endanger human health. In recent years, the incidence of sleep-disordered breathing patients, represented by OSAHS (obstructive sleep apnea-hypopnea syndrome), has been gradually increasing. During sleep, patients experience repeated partial or complete obstruction of the upper airway, resulting in apnea or hypopnea, with snoring and daytime sleepiness as the main clinical manifestations.
[0003] Existing polysomnography (PSG) devices for monitoring sleep breathing physiological parameters have significant technical drawbacks: high equipment cost, cumbersome operation procedures, and the need for professional personnel to perform monitoring and data interpretation; during monitoring, electrodes and sensors need to be placed on multiple parts of the subject's body, which can easily cause discomfort.
[0004] Limited by the testing environment, interference factors such as the "first night effect" and the requirement for fixed body positions make it difficult for monitoring data to reflect the true physiological situation under natural sleep conditions. Furthermore, there are individual and observer differences in parameter measurement results. Existing video monitoring technologies are only applied to security and identity authentication scenarios and have not developed effective monitoring solutions for the special characteristics of sleep respiratory physiology and privacy protection needs.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] According to one aspect of this application, a non-invasive and interference-free method for monitoring sleep breathing is provided, comprising: acquiring visual data and associated modal data of multiple parts of the human body during sleep; preprocessing the acquired sleep-related data, blurring private parts of the human body through a privacy protection unit, using a segmentation unit to achieve accurate segmentation and time alignment of 10 specific part videos, and combining PSG tags to complete data quality control and time synchronization; performing detailed reasoning on the preprocessed sleep data, analyzing pixel grayscale changes and motion amplitude features of each part video, and identifying sleep-related events including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement. The system generates multi-dimensional sleep event details; it comprehensively analyzes these details, integrates multimodal data to construct a sleep state assessment model, and combines statistical features for weighted integration to achieve two to five categories of sleep stages, generating core sleep events; it optimizes the presentation of the comprehensive analysis results, displaying video capture content, sleep event details, and sleep stage results in real time on the user interface, generating sleep breathing state assessment information; and it integrates sleep breathing state assessment information, multimodal raw data, and inference analysis process data to generate sleep breathing state early warning information that includes sleep quality level, quantitative indicators of breathing abnormalities, and sleep structure characteristics.
[0007] Another aspect of this application discloses a non-invasive and interference-free sleep breathing monitoring device, comprising: an acquisition module for acquiring visual data and associated modal data of multiple parts of the human body during sleep; a processing module for preprocessing the acquired sleep-related data, blurring private parts of the human body through a privacy protection unit, accurately segmenting and aligning 10 specific body part videos using a segmentation unit, and performing data quality control and time synchronization using PSG tags; and performing detailed reasoning on the preprocessed sleep data, analyzing pixel grayscale changes and motion amplitude features of each body part video, and identifying features including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement. The system collects sleep-related events and generates multi-dimensional sleep event details. It then performs comprehensive analysis of these details, integrates multimodal data to construct a sleep state assessment model, and combines statistical features for weighted integration to achieve two to five categories of sleep stages, generating core sleep events. The system optimizes the presentation of the comprehensive analysis results, displaying video capture content, sleep event details, and sleep stage results in real-time on the user interface, generating sleep breathing state assessment information. Finally, it integrates the sleep breathing state assessment information, multimodal raw data, and inference analysis process data to generate sleep breathing state early warning information that includes sleep quality levels, quantitative indicators of respiratory abnormalities, and sleep structure characteristics.
[0008] According to another aspect of this application, an electronic device includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute, by executing the executable instructions, a non-invasive and disturbance-free sleep breathing monitoring method as described above.
[0009] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described non-invasive and interference-free method for monitoring sleep breathing.
[0010] This application provides a non-invasive and interference-free method and system for monitoring sleep breathing. Based on previously accumulated high-definition video data related to sleep breathing and synchronized PSG data, a standardized data annotation system is constructed. Using PSG monitoring parameters as a reference, video segmentation processing with physiological monitoring significance is performed on the high-definition video of the sleep process. A non-invasive video acquisition module is used to capture physiological activity characteristics such as chest and abdominal movements during sleep without contact with the human body or disturbing sleep. Feature extraction and analysis of the acquired video data are performed using computer vision technology. Combined with artificial intelligence algorithms trained on medical big data, the system achieves automated identification, quantitative analysis, and state assessment of sleep breathing-related physiological parameters. This application overcomes the technical limitations of traditional contact-based monitoring, eliminating the need for invasive sensors and enabling monitoring in a natural sleep environment. It offers advantages such as ease of operation, high comfort, and low interference. Through core technologies like video segmentation, feature extraction, and AI analysis, it achieves precise capture and data-driven assessment of sleep-related respiratory physiological activities. The output monitoring results can serve as reference data for analyzing sleep-related respiratory physiological states, providing a new technical pathway for convenient and accurate monitoring of sleep-related physiological parameters. It is suitable for large-scale screening scenarios of sleep-related respiratory physiological states and possesses significant technical practicality and application value.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] Figure 1 A flowchart illustrating a non-invasive and disturbance-free sleep breathing monitoring method provided in an embodiment of this application is shown. Figure 2 A schematic diagram of a non-invasive and disturbance-free sleep breathing monitoring device provided in an embodiment of this application is shown. Detailed Implementation
[0013] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0014] In one embodiment, this application also proposes a non-invasive and undisturbed method and system for monitoring sleep breathing. Figure 1 A schematic flowchart of a non-invasive and undisturbed sleep breathing monitoring method according to an embodiment of this application is shown.
[0015] S101 acquires visual data and related modal data of multiple parts of the human body during sleep. The visual data of multiple parts of the human body includes video data of 10 specific parts such as the eyes, mouth, lips, Adam's apple, chest, and abdomen collected by visual sensors.
[0016] In one embodiment, the visual sensor is a commercially available Hikvision 4-megapixel ultra-high-definition camera with a resolution of no less than 2560. 1440. Ensure the collected video data clearly captures subtle physiological activities during sleep. Deploy 1-5 sensors, covering the headboard, footboard, sides, and above the bed (e.g., ceiling), with at least one sensor located at the headboard, footboard, or above the bed. All sensors should be within three meters of the bed's center to ensure comprehensive coverage of the sleep area. In a bedroom scenario, install one sensor on the wall at the head of the bed, one on the ceiling at the foot of the bed, and one on each side of the bed, for a total of four sensors working together to comprehensively capture the dynamics of the eyes, chest, abdomen, and other parts of the body during sleep, avoiding data loss due to obstruction from a single angle.
[0017] The sensor acquires real-time RGB video and optical flow data of the entire human body. Through image segmentation technology, it accurately extracts 10 independent video streams from specific body parts. Each video stream generates a unique timestamp, ensuring complete temporal alignment of data across all parts, laying the foundation for subsequent synchronous analysis. The 10 specific body parts include the eyes, mouth, lips, Adam's apple, chest, abdomen, entire head, both lower limbs, both upper limbs, and the entire torso. Each video stream focuses on the physiological activity characteristics of its corresponding part. After acquiring full-body video of a sleeping person, the sensor automatically segments the video stream from the eyes, specifically capturing subtle fluctuations in the eyelid skin. Simultaneously, it segments independent video streams from the chest and abdomen, accurately recording the rise and fall of the ribcage and abdomen. The 10 video streams generate timestamps synchronously, ensuring that eyelid fluctuations and chest / abdominal movements at any given moment can be analyzed in a corresponding manner.
[0018] In addition to 10 channels of visual data from specific body parts, the system also supports the access of multimodal correlated data, including bioradar data, snoring data, PSG data, and optional infrared thermal imaging data. Infrared thermal imaging data uses the face as the monitoring and positioning area, supplementing the monitoring dimensions of visual data by analyzing the thermal changes in respiratory airflow and the microcirculation perfusion under the skin. Other modal data are acquired synchronously through dedicated acquisition equipment, complementing the visual data. While acquiring 10 channels of visual data, the system obtains radar signals of human respiratory frequency through a bioradar device deployed at the bedside, and records snoring intensity and frequency data through a bedside sound acquisition device. If the ambient light is dim, the infrared thermal imaging function can be activated to capture the thermal trajectory of facial respiratory airflow. All correlated modal data are synchronized with the timestamps of the visual data, enabling multi-dimensional collaborative data analysis.
[0019] S102 preprocesses the collected sleep-related data, blurs the private parts of the human body through the privacy protection unit, uses the segmentation unit to achieve accurate segmentation and time alignment of 10 specific part videos, and completes data quality control and time synchronization by combining PSG tags.
[0020] In one implementation, the collected sleep-related data is categorized and processed to generate an adaptation correlation result between data types and preprocessing mechanisms. The sleep-related data includes full-body RGB video data, optical flow data, and optional infrared thermal imaging data collected by visual sensors. The associated data is PSG-tagged data. All collected data is categorized to clarify the adaptation relationship between data types and corresponding preprocessing mechanisms, ensuring that each type of data has a targeted processing path. The core sleep-related data is divided into two categories: one is the basic data directly collected by the visual sensors, including full-body RGB video data, optical flow data, and optional infrared thermal imaging data. This type of data focuses on the visual characteristics and thermal changes during human sleep. The other category is the associated reference data, i.e., PSG-tagged data, which serves as the clinical gold standard data for subsequent quality control and time synchronization. Based on data characteristics, adaptation and association rules are established. RGB video data and optical flow data are adapted to "privacy protection + part segmentation" preprocessing, infrared thermal imaging data are adapted to "thermal feature extraction + privacy re-checking" preprocessing, and PSG tag data are adapted to "expert verification + format standardization" preprocessing. Finally, a one-to-one correspondence and association result between data types and preprocessing mechanisms is generated.
[0021] A full-body RGB video and optical flow data of a subject during nighttime sleep were collected, along with synchronously recorded PSG tag data. Simultaneously, infrared thermal imaging was used to acquire facial thermal data. After classification, the RGB video and optical flow data were associated with a "privacy protection + body part segmentation" process, the infrared thermal imaging data with a "thermal feature extraction + privacy review" process, and the PSG tag data with an "expert verification + format standardization" process, forming clear and compatible association results that lay the foundation for subsequent targeted processing.
[0022] The adaptation and association results are analyzed in a targeted manner to generate a list of core preprocessing variables, including a privacy-preserving regional positioning accuracy variable, a segmentation integrity variable for 10 video segments, a PSG tag quality control and verification variable, and a data time alignment synchronization error variable. The adaptation and association results are then broken down to extract key control variables for each type of data preprocessing, constructing a core preprocessing variable list to ensure the preprocessing process is quantifiable and controllable. The core variables are set around four core aspects: privacy protection, segmentation, tag quality control, and time alignment. The privacy-preserving regional positioning accuracy variable controls the accuracy of privacy-preserving part positioning, with a quantification standard of positioning error ≤ 1cm; the segmentation integrity variable for 10 video segments ensures no omissions in segmenting 10 specific human body parts, with a quantification standard of segmentation coverage ≥ 99%; the PSG tag quality control and verification variable ensures tag accuracy, with a quantification standard of three-expert verification consistency rate ≥ 95%; and the data time alignment synchronization error variable controls the time deviation between visual data and PSG tags, with a quantification standard of synchronization error ≤ 50ms.
[0023] After analyzing the above-mentioned adaptation and association results, specific values for core preprocessing variables are generated: the privacy protection area positioning accuracy variable is set to "positioning error ≤ 0.8cm", the 10-channel part video segmentation integrity variable is set to "segmentation coverage ≥ 99.5%", the PSG tag quality control verification variable is set to "consistency rate of three experts ≥ 96%", and the data time alignment synchronization error variable is set to "synchronization error ≤ 30ms". Preprocessing quality is ensured by clearly defining quantitative indicators.
[0024] The adaptation and association results, the core preprocessing variable list, and the original sleep-related data and PSG tag data are simultaneously verified and screened for compliance. Combining the tag verification mechanism and cross-validation rules, privacy-protected processing results, time-aligned 10-channel body part video data, and standardized PSG tag association data are generated. The simultaneous verification process, using the tag verification mechanism and cross-validation rules, comprehensively verifies the adaptation and association results, the core preprocessing variable list, and the original sleep-related data and PSG tag data, while simultaneously conducting compliance screening to remove invalid data and non-compliant content. The simultaneous verification phase focuses on checking data integrity (e.g., no missing video frames), variable fit (e.g., positioning accuracy meets set standards), and temporal consistency (e.g., visual data matches PSG tag timestamps). The tag verification mechanism requires three trained sleep apnea physicians to independently verify the PSG tags; disagreements are resolved through collective discussion and voting, and weekly cross-validation is conducted to eliminate subjective bias. The compliance screening phase removes video data with excessive ambiguity and PSG tags that fail verification, ensuring the output data is compliant and valid. The final result is three types of target data: privacy protection processing results (privacy parts are blurred and frame-to-frame stable), time-aligned 10-channel part video data (timestamps of each video are synchronized and segmented completely), and standardized PSG tag associated data (uniform format and verified).
[0025] The raw data from the aforementioned subjects were simultaneously verified. A missing frame in the RGB video was discovered and removed. Three PSG labels showed disagreement among experts; these were corrected through collective discussion and voting and passed verification. After privacy protection processing, Kalman filtering was used to ensure stable and unwavering blurring of the 3cm region above and below the line connecting the nipples across consecutive frames. The final output included a complete video with blurred private areas, videos of 10 body parts (including the eyes and mouth) with perfectly aligned timestamps, and standardized PSG label-related data with a 97% expert verification consistency rate, providing high-quality data support for subsequent analysis and inference.
[0026] S103 performs detailed reasoning on the preprocessed sleep data, analyzes the pixel grayscale changes and motion amplitude features of the video of each part, identifies sleep-related events including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement, and generates multi-dimensional sleep event details.
[0027] In one implementation, ISP image signal processing technology is used to extract features from 10 preprocessed video feeds of different body parts, generating pixel grayscale change sequences, motion amplitude quantification data, and a visual semantic feature set for each body part. ISP image signal processing technology is used to extract features from each of the 10 time-aligned preprocessed video feeds of different body parts, accurately capturing core features related to sleep physiological events in each video feed and generating three types of standardized feature data. For each body part video, grayscale value changes are extracted in chronological order to form a pixel grayscale change sequence, reflecting the subtle dynamics of the body part. Parameters such as displacement and velocity of the body part's motion are calculated using inter-frame difference and optical flow methods to generate motion amplitude quantification data, quantifying the intensity of physiological activity. Image semantic analysis technology is used to extract key semantic information that can characterize the physiological state of the body part, forming a visual semantic feature set, providing high-dimensional support for event recognition.
[0028] When extracting features from videos of the "chest" area, the ISP technology records the grayscale value of each pixel at a frequency of 30 frames per second, generating a grayscale change sequence of the chest undulation process; by calculating the displacement distance of the chest area between adjacent frames, motion amplitude quantification data (such as undulation amplitude of 0.5-3cm and undulation frequency of 12-20 times / min) is generated; at the same time, visual semantic features such as "chest expansion" and "chest contraction" are extracted, which together with the feature data of other 9 areas (such as abdomen, eyes, etc.) form a complete feature set.
[0029] This study integrates feature sets with sleep-related event recognition rules, physiological activity association standards for different body parts, and medical annotation thresholds to establish a precise mapping between features and events, generating an event recognition fusion dataset. The extracted feature sets are deeply integrated with pre-defined sleep-related event recognition rules, physiological activity association standards for different body parts, and medical annotation thresholds to clarify the correspondence between different feature combinations and specific sleep events, establishing a precise mapping and generating an event recognition fusion dataset. Sleep-related event recognition rules cover the judgment logic for various events such as eye opening, respiratory rhythm, and blood oxygen changes. Physiological activity association standards for different body parts clarify the synergistic relationship of features in different body parts (e.g., the synchronicity of chest and abdominal movement). Medical annotation thresholds are set based on the PSG gold standard (e.g., a decrease in blood oxygen saturation ≥4% is considered a hypoxic event). Through multi-dimensional information fusion, the accuracy of the feature-event mapping is ensured.
[0030] The pixel grayscale change sequence of the "lips" area (such as a continuous decrease in the grayscale value of lip color), the quantitative data of movement amplitude (such as mouth opening amplitude ≥1cm), and the medical annotation threshold (blood oxygen saturation decrease ≥4%) are fused to establish a mapping relationship between "decline in lip color grayscale + mouth opening amplitude meets the standard + blood oxygen threshold trigger" and "blood oxygen saturation decrease event". At the same time, combined with the respiratory rhythm characteristics of "chest" and "abdomen", a mapping relationship between "sudden drop in chest and abdominal fluctuation amplitude + duration ≥10 seconds" and "apnea event" is established. All mapping relationships and corresponding feature data are integrated into an event recognition fusion dataset.
[0031] A multi-site sleep event classification and reasoning model is constructed based on a fusion dataset. Using 10 video feeds from different body sites as input, visual semantic features as core parameters, and medical annotation standards as the judgment criteria, the model classifies and identifies specific physiological events, generating initial sleep event details. Based on an event recognition fusion dataset, a multi-site sleep event classification and reasoning model is built. The model uses feature data corresponding to 10 video feeds as input, visual semantic features as core parameters, and medical annotation standards (PSG examination result calibration) as the judgment criteria. Deep learning algorithms are used to train the model's ability to classify and identify various sleep events. The model incorporates feature matching logic for different sleep events, automatically analyzing the input multi-site feature data to determine whether it matches the feature patterns of a specific sleep event, and then outputting initial sleep event details, clarifying the event type and preliminary association information.
[0032] After receiving feature data from 10 different areas, the model identifies "rapid eye movement events" by analyzing the grayscale change sequence (minor fluctuations in eyelid skin) and visual semantic features ("eyelid tremors") of the "both eyes" area, combined with medical annotation standards (eyelid tremor frequency ≥ 5 times / second during REM sleep). It also identifies "swallowing events" by analyzing the quantitative data of the movement amplitude of the "Adam's apple" area (vertical displacement of the Adam's apple ≥ 0.3cm). All identified events are arranged in chronological order to generate detailed initial sleep event data, including basic information such as event type, initial occurrence time, and associated location.
[0033] By combining the specificity and correlation of physiological activities in different body parts, the initial sleep event details data are calibrated with timestamps and validated with event types. This clarifies the occurrence time, duration, associated body parts, and confidence level of various events, generating multi-dimensional sleep event details including eye opening, respiratory rhythm, and blood oxygen changes. The initial sleep event details data undergoes dual optimization processing to ensure accuracy and completeness, taking into account the specificity and correlation of physiological activities in different body parts. In the timestamp calibration stage, the unified timestamps of 10 body part videos are used as a benchmark to correct deviations in the event occurrence times in the initial data, ensuring consistency in timestamps for related events at the same time point (such as apnea and decreased blood oxygen). In the event type validation stage, based on the synergistic relationship of physiological activities in different body parts (such as respiratory events requiring verification from both chest and abdominal features), contradictory events are eliminated (e.g., a single "sudden drop in abdominal movement" without chest feature support is deemed invalid), and missed events are supplemented (e.g., identifying hypoxic events not initially identified through the synergistic identification of lip color changes and abnormal respiratory rhythm). Ultimately, the precise occurrence time, duration, associated location, and identification confidence level (e.g., confidence level 92%-98%) of various events are determined, generating detailed sleep event information that includes multiple dimensions such as eye opening, eye movement, respiratory rhythm, blood oxygenation changes, and body movement.
[0034] The initially identified "apnea events" (marked as occurring between 02:15:30 and 02:15:45) were calibrated and verified. Correction was performed using unified timestamps from 10 different body parts, confirming the actual occurrence time as 02:15:28-02:15:43 (lasting 15 seconds). Combining the blood oxygen saturation drop characteristics of the "lips" area (95% confidence) with the rise and fall stagnation characteristics of the "chest" and "abdomen" areas (96% confidence), the event type was verified as valid. Simultaneously, associated body parts "lips," "chest," and "abdomen" were added. The final event details were: "Event Type: Apnea Event, Occurrence Time: 02:15:28, Duration: 15 seconds, Associated Body Parts: Chest, Abdomen, Lips, Confidence: 95%." All event details were integrated to form a complete multi-dimensional sleep event detail.
[0035] S104 comprehensively analyzes detailed information on multi-dimensional sleep events, integrates multimodal data to construct a sleep state assessment model, and combines statistical features for weighted integration to achieve two to five categories of sleep stages and generate core sleep events.
[0036] In one implementation, multi-dimensional sleep event details, multimodal correlation data, and medical annotation standard data are classified, extracted, and feature-aligned to generate 10-way site-specific event temporal feature information, multimodal complementary feature information, sleep stage determination threshold information, and core event association rule information. The multi-dimensional sleep event details, multimodal correlation data (bioradar, snoring, PSG, etc.), and medical annotation standard data are systematically classified and extracted, and feature alignment processing is completed to generate four types of standardized information. From the sleep event details of 10 sleep regions, data such as the occurrence time and characteristic intensity of events in each region are extracted in chronological order to form the temporal characteristic information of events in 10 sleep regions; from the multimodal association data, complementary features unique to each modality (such as respiratory depth data of bioradar, decibel and frequency data of snoring) are extracted to generate multimodal complementary feature information; based on the PSG gold standard and clinical diagnostic guidelines, the critical values for sleep stage determination are defined (such as deep sleep body movement frequency ≤ 5 times / hour) to form sleep stage determination threshold information; the triggering conditions and association logic of various core sleep events (apnea, hypoxia, etc.) are sorted out (such as apnea must be accompanied by cessation of chest and abdominal movement + decrease in blood oxygen) to generate core event association rule information.
[0037] After processing the multi-source data of the subjects, the temporal feature information of the 10 site events includes temporal data such as "rapid eye movement events (continuous occurrence from 01:30 to 01:45)" and "chest movement events (frequency 18 times / min)"; the multimodal complementary feature information includes data such as "breathing depth 3cm" detected by bioradar and "maximum decibel 65dB" recorded by snoring device; the sleep stage determination threshold information clearly defines "body movement frequency of 5-15 times / hour during light sleep" and "eyelid pulsation frequency ≥5 times / second during rapid eye movement"; the core event association rule information clearly defines "apnea event: chest and abdominal movement cessation ≥10 seconds + blood oxygen saturation decrease ≥4%".
[0038] Based on the goal of accurate multimodal fusion assessment, this method correlates respiratory rhythm fluctuation features and body movement frequency features from 10 site-specific event time-series features with physiological signal features from the bioradar / snoring modality. A weighted average integration algorithm is then used to quantify the contribution and synergistic information of each modality in sleep assessment, generating information to support multimodal data fusion. Specifically, it performs cross-modal correlation between key physiological features (respiratory rhythm fluctuation features and body movement frequency features) from 10 site-specific event time-series features and physiological signal features (respiratory depth features from bioradar and frequency features of snoring) from multimodal correlated data. A weighted average integration algorithm is then used to quantify the contribution and synergistic information of each modality in sleep assessment, generating information to support multimodal data fusion. The algorithm calculates the weights of different modal features for sleep event recognition (e.g., visual modality respiratory rhythm feature weight 0.6, bio-radar respiratory depth feature weight 0.3, snoring frequency feature weight 0.1), clarifies the complementary relationship and synergistic effect of various modalities (e.g., visual modality captures chest and abdominal movements, bio-radar supplements respiratory depth, and together improves the accuracy of sleep apnea recognition), and provides a quantitative basis for subsequent fusion inference.
[0039] The respiratory rhythm fluctuation characteristics of "chest + abdomen" (weight 0.55), the respiratory depth characteristics of bio-radar (weight 0.3), and the frequency characteristics of snoring (weight 0.15) are correlated. The weighted average integration algorithm calculates that when "chest and abdomen fluctuation amplitude ≤ 0.5cm + respiratory depth ≤ 1cm + snoring frequency drops to 0", the confidence of the synergistic pointing to the sleep apnea event increases to 92%. This quantitative relationship and weight allocation together constitute the basis information for multimodal data fusion.
[0040] Based on the goals of detailed sleep staging and accurate event identification, sleep staging thresholds and core event association rules are mapped to staging level adaptation models and event type identification optimization models, respectively, generating dynamic threshold and rule adaptation information. The staging level adaptation model dynamically adjusts the sleep staging thresholds based on individual differences such as age and sleep habits (e.g., the deep sleep body movement frequency threshold for elderly subjects is adjusted to ≤8 times / hour). The event type identification optimization model dynamically optimizes core event association rules based on real-time monitoring data (e.g., the threshold for decreased blood oxygen saturation in low-temperature nighttime environments is adjusted to ≥3%), ensuring that the assessment results are adapted to individual and environmental differences.
[0041] For elderly subjects, the staging and grading adaptation model dynamically adjusted the standard of "deep sleep body movement frequency ≤ 5 times / hour" to "≤ 8 times / hour"; the event type identification and optimization model, combined with environmental data of nighttime room temperature of 18℃, dynamically adjusted the blood oxygen saturation decrease threshold of hypoxia events from ≥4% to ≥3%. The generated dynamic threshold and rule adaptation information clearly stated that "deep sleep determination for elderly subjects: body movement frequency ≤ 8 times / hour + eyelid pulsation frequency ≤ 2 times / second" and "hypoxia events in low temperature environment: blood oxygen saturation decrease ≥ 3% + decrease in lip color grayscale ≥ 20%".
[0042] Based on multimodal data fusion information and dynamic threshold and rule adaptation information, core features in multidimensional sleep event details are labeled and weighted to generate preliminary sleep state staging results and candidate core sleep events. The core features (such as respiratory rhythm, body movement frequency, and blood oxygenation changes) in multidimensional sleep event details are labeled and weighted according to quantified weights to generate preliminary sleep state staging results and candidate core sleep events. By calculating the matching score for each sleep stage (e.g., 85 points for deep sleep and 60 points for light sleep), the preliminary sleep stage is determined; simultaneously, potential core events that conform to dynamic rules are screened and labeled as candidate core sleep events, providing a basis for further verification.
[0043] After weighting the subject data, the following parameters were calculated: body movement frequency 6 times / hour (weight 0.3, score 2.4), eyelid pulsation frequency 1 time / second (weight 0.2, score 0.2), and respiratory rhythm stability (weight 0.5, score 4.5). The combined score was 7.1, which matched the criteria for deep sleep and generated a preliminary sleep stage result of "02:00-03:00 is the deep sleep period". At the same time, the event of "pause in chest and abdominal movement for 12 seconds (weight 0.6) + decrease in blood oxygen saturation of 3.5% (weight 0.4)" was selected and marked as a candidate core sleep event "suspected hypoxia accompanied by apnea".
[0044] Key physiological signal features in multimodal correlated data are labeled and substituted into a sleep event recognition model to calculate the frequency, duration, and impact of events, identifying core sleep event features across multiple dimensions. A dynamic weighted fusion strategy is used to integrate multi-source features, and preliminary identification results of core sleep events are generated through fusion confidence calculation and outlier filtering. Similarly, key physiological signal features in multimodal correlated data (such as a sudden drop in respiratory depth or a sudden cessation of snoring in bioradar) are labeled and substituted into a sleep event recognition model to calculate the frequency, duration, and impact of events (such as a 15-second apnea affecting a 4% decrease in blood oxygen saturation), identifying core sleep event features across multiple dimensions. A dynamic weighted fusion strategy is used to integrate multi-source features, and preliminary identification results of core sleep events are generated through fusion confidence calculation (such as a 93% collaborative confidence level for multimodal features) and outlier filtering (removing transient interference data from sensors).
[0045] Substituting features such as "breathing depth suddenly drops from 3cm to 0.8cm" from bio-radar, "snoring device decibels suddenly drop from 60dB to 0dB" from snoring device, "chest and abdominal movements stop for 15 seconds" from visual modality, and "blood oxygen saturation drops from 98% to 94%" into the model, it was calculated that the event lasted for 15 seconds, affecting blood oxygen by 4%, with a multimodal fusion confidence of 93%. After filtering out one frame of sensor interference data, the preliminary identification result of the core sleep event was generated as "02:15-02:15:15 sleep apnea event (accompanied by hypoxia)".
[0046] The process integrates and validates the preliminary sleep stage results with the preliminary identification results of core sleep events to achieve a two- to five-category sleep stage, generating a comprehensive assessment result that includes core sleep events such as sleep apnea and hypoxia. This integration and validation process eliminates contradictory data (e.g., if the preliminary stage indicates deep sleep but high-frequency body movement is present, feature weights need to be rechecked) and supplements missing information (e.g., if core events are not associated with corresponding blood oxygenation data, backtesting is required). Ultimately, the process achieves a two- to five-category sleep stage (e.g., five categories: N1, N2, N3, REM sleep, and wakefulness), generating a comprehensive assessment result that includes core sleep events such as sleep apnea, hypoxia, and forced breathing, clearly defining the occurrence time, duration, correlation characteristics, and confidence level of each type of event.
[0047] During the integration and verification process, it was found that the initial stage "02:00-03:00 deep sleep period" and the candidate event "02:15 apnea" were consistent, and the multimodal feature co-confidence was 93%. After further verification of the feature corresponding to this event, "25% decrease in lip color grayscale", the validity of the event was confirmed. The final comprehensive assessment result was generated as follows: sleep stage (five categories) "02:00-02:05 N2 stage, 02:05-03:00 N3 stage (deep sleep)"; core sleep event "02:15-02:15:15 apnea event (lasting 15 seconds, confidence level 93%), accompanied by hypoxia (blood oxygen saturation 94%)".
[0048] S105 optimizes the presentation of comprehensive analysis results, displaying video capture content, sleep event details, and sleep stage results in real time on the user interface, and generating sleep breathing status assessment information.
[0049] In one implementation, the comprehensive analysis results are categorized and core information is extracted to generate sleep staging results, detailed information on core sleep events, quantitative information on physiological indicators, and multimodal raw video clips. The sleep-related data obtained from the comprehensive analysis are systematically categorized, and various types of core information are precisely extracted to generate four types of standardized display data. Data such as the staging time period and proportion under different classification standards are extracted from the sleep staging results to form sleep staging results information; key information such as the occurrence time, duration, and confidence level of core sleep events such as sleep apnea and hypoxia are analyzed to generate detailed information on core sleep events; the specific values and trends of physiological indicators such as respiratory rate, number of body movements, and blood oxygen saturation are quantitatively extracted to generate quantitative information on physiological indicators; and multimodal raw video clips corresponding to core events and key physiological states (such as chest + abdominal videos during sleep apnea) are extracted to generate multimodal raw video clip information, providing complete data support for subsequent display.
[0050] After processing, the sleep staging results are as follows: "Five categories: N1 stage (00:30-01:10, 12%), N2 stage (01:10-03:40, 45%), N3 stage (03:40-05:00, 20%), REM sleep (05:00-06:20, 18%), and wakefulness (06:20-06:30, 5%)"; the core sleep event details are "02:15-02:15:15 apnea event (...)". The data includes: a 15-second video clip (93% confidence level) showing the duration of apnea; a hypoxic event from 03:00 to 03:00:08 (oxygen saturation 94%, 91% confidence level); quantitative information on physiological indicators including: an average thoracic respiratory rate of 16 breaths / min, a total number of body movements of 32, and a minimum oxygen saturation of 92%; and original multimodal video clips including: a 20-second video clip of the chest and abdomen at 02:15 when apnea occurred, and a 15-second video clip of the lips corresponding to the hypoxic event at 03:00.
[0051] Based on the user's intuitive viewing objective, the display interface adaptation system analyzes user scenarios and key information needs, quantifies the differences in information presentation requirements among different users, and generates scenario-based adaptation parameters. Based on the core objective of intuitive user viewing, the display interface adaptation system deeply analyzes user scenarios (clinical diagnosis, home monitoring) and key information needs, quantifies the differences in information presentation requirements among different users, and generates scenario-based adaptation parameters. In clinical diagnosis scenarios, users (doctors) require detailed data and original videos, emphasizing data accuracy and completeness; in home monitoring scenarios, users (general users) require concise conclusions and key warnings, emphasizing intuitiveness and ease of understanding. By quantifying the differences in requirements (e.g., doctors require 90% data detail, while general users require 85% conciseness), the system clarifies the information presentation format, level of detail, and key content to be displayed in different scenarios.
[0052] The scenario-based adaptation parameters for clinical diagnosis are: "Presentation format: table + original video + detailed curve; Detail level: display all original physiological indicators and complete event details; Key content: sleep stage percentage curve, core event video playback, and blood oxygen saturation change trend"; The scenario-based adaptation parameters for home monitoring are: "Presentation format: chart + key warnings + simplified conclusions; Detail level: only display core indicators (sleep quality level, number of abnormal events); Key content: sleep quality score (82 points), sleep apnea event warning (2 times), and suggested improvement directions."
[0053] Based on the constraints of real-time display and historical backtracking, the information update frequency, data loading rate, and storage usage are detected and optimized to generate display performance guarantee parameters, including real-time data refresh standards, historical data caching rules, and large file chunking loading schemes. The real-time data refresh standards clearly define the update frequency of different data (e.g., physiological indicators refresh once per second, sleep staging updates once every 5 minutes); historical data caching rules determine caching priorities (core event videos > physiological indicator data > sleep staging results) to ensure fast retrieval during backtracking; the large file chunking loading scheme uses a "load keyframes first + load complete content later" approach for large files such as video clips to avoid loading stutters and ensure smooth display.
[0054] The performance guarantee parameters are as follows: Real-time data refresh standard: physiological indicators (blood oxygen, respiratory rate) 1 time / second, sleep stage status 5 minutes / time, core events are pushed in real time; historical data caching rules: core event videos are cached for 30 days, physiological indicator data is cached for 90 days, and sleep stage results are cached for a long time. The caching priority is: event video > physiological data > stage results; large file segment loading scheme: video segments are divided into 10-second units. The 3-second keyframe when the event occurs is loaded first, and the remaining part is loaded asynchronously in the background with a loading delay of ≤1 second.
[0055] To ensure privacy and security, the video clips to be displayed undergo a re-examination to blur private areas. A tiered information access permission mechanism is implemented, generating privacy protection parameters for display. Specifically, the original multimodal video clips to be displayed undergo a re-examination to ensure that the blurring effect in private areas such as the 3cm area above and below the line connecting the breasts and the inner sides of the groin meets standards and poses no risk of privacy leakage. A tiered information access permission mechanism is also implemented to clearly define the access permissions for different users (doctors, regular users, and administrators). (For example, doctors can access all videos and data, while regular users can only access their own simplified data and videos without privacy concerns). This process generates privacy protection parameters that balance privacy protection with display needs.
[0056] The privacy protection parameters are: "Privacy blur re-check: The area within 3cm above and below the line connecting the breasts and the inner side of the groin area of all displayed video clips are blurred twice to ensure that the blur intensity is ≥80% and the inter-frame stability error is ≤2%; Access permission level: Doctor permission (account and password verification, can access all data + original video), ordinary user permission (identity verification, can access simplified data + privacy-blurred video), administrator permission (multi-factor verification, can manage user permissions + data backup)".
[0057] By combining the characteristics of the user display interface, scenario-based adaptation parameters, display performance assurance parameters, display privacy protection parameters, and interface display capabilities are correlated and matched to dynamically adjust the information presentation format and display priority. Specifically, based on the characteristics of the user display interface (desktop, mobile, tablet), scenario-based adaptation parameters, display performance assurance parameters, display privacy protection parameters, and interface display capabilities (resolution, screen size, ease of use) are precisely correlated and matched to dynamically adjust the information presentation format and display priority. Desktop devices, with their large screens and ease of use, can display multi-window data and high-definition video; mobile devices, with their smaller screens and emphasis on portability, prioritize the display of core conclusions and warning information. By matching the performance limits and usage scenarios of different interfaces, the display effect is ensured to adapt to the characteristics of the terminal.
[0058] After matching on the computer, the display format is "three-window layout (left: sleep stage curve + physiological indicator table; middle: core event details + video playback area; right: original video after privacy blurring); display priority: core event details > sleep stage curve > physiological indicator table > video playback"; the display format on the mobile device is "single-window scrolling layout (top: sleep quality score + abnormal event warning; middle: key physiological indicator card; bottom: simplified sleep stage pie chart + video thumbnail); display priority: sleep quality score > abnormal event warning > key indicator card > stage pie chart".
[0059] The above processing results are integrated and verified. Through display effect testing, privacy compliance review, and user experience simulation evaluation, a sleep breathing state assessment information display solution adapted to multiple scenarios and terminals is generated, presenting video capture content, sleep event details, and sleep stage results in real time. The above-mentioned scenario-based adaptation parameters, display performance assurance parameters, display privacy protection parameters, and dynamically adjusted presentation format are comprehensively integrated and verified. Three types of core tests ensure the solution's compliance, smoothness, and adaptability, generating a sleep breathing state assessment information display solution adapted to multiple scenarios and terminals. Display effect testing verifies the clarity and layout rationality of information presentation on different terminals; privacy compliance review verifies the privacy blurring effect and the effectiveness of permission hierarchy; user experience simulation evaluation verifies user satisfaction with the display format in different scenarios. The final solution clarifies the information presentation format, output sequence, display priority, and interaction rules for each scenario and terminal, achieving the goal of real-time presentation of video capture content, sleep event details, and sleep stage results.
[0060] After integration and verification, the computer-based display solution for clinical diagnosis scenarios is as follows: "Upon startup, the sleep stage curve and core event details are loaded first (time ≤ 2 seconds), followed by the asynchronous loading of physiological indicator tables and video clips; video clips can be dragged and played, physiological indicator curves can be zoomed in for viewing, and events and videos can be precisely linked (clicking on an event automatically jumps to the corresponding video time point)." The mobile-based display solution for home monitoring scenarios is as follows: "Upon startup, the sleep quality score (82 points) and abnormal event warning (2 episodes of apnea) are directly displayed. Swiping down allows viewing key indicator cards (average respiratory rate 16 breaths / min, minimum blood oxygen 92%) and simplified stage pie charts. Clicking on a warning event allows viewing a privacy-blurred video clip (duration 15 seconds); real-time data is refreshed every 1 second, and abnormal events are alerted via pop-up windows in real time."
[0061] S106 integrates sleep breathing state assessment information, multimodal raw data, and inference analysis process data to generate sleep breathing state early warning information that includes sleep quality level, quantitative indicators of breathing abnormalities, and sleep structure characteristics.
[0062] In one implementation, the sleep staging results and core sleep event data from the sleep breathing status assessment information are compared with the OSAHS screening thresholds and sleep quality assessment standards in medical diagnostic criteria to generate a medical fit quantitative factor for the assessment results. This factor characterizes the degree of fit between the monitoring results and clinical diagnostic needs. Core data (sleep staging results and core sleep event data) from the sleep breathing status assessment information are extracted and compared with the OSAHS screening thresholds and sleep quality assessment standards in medical diagnostic criteria to generate a comprehensive fit quantitative factor. This factor directly represents the degree of fit between the monitoring results and clinical diagnostic needs (values range from 0 to 1, with a higher fit closer to 1). OSAHS screening thresholds include standards such as an apnea-hypopnea index (AHI) ≥5 times / hour and a minimum blood oxygen saturation ≤90%. Sleep quality assessment standards cover indicators such as the percentage of sleep stages (e.g., deep sleep ≥20% is considered acceptable) and body movement frequency. By quantifying the degree of conformity between the monitoring data and the standards, the medical fit quantitative factor is determined.
[0063] The subjects' sleep stage results were "deep sleep 20% and REM sleep 18%", and the core sleep event data were "AHI 12 times / hour and lowest blood oxygen saturation 92%". After fitting analysis with medical standards, the percentage of deep sleep met the standard (fit 1.0), the percentage of REM sleep met the standard (fit 1.0), the AHI exceeded the standard (fit 0.6, because 12 times / hour is higher than 5 times / hour but not reaching the severe OSAHS standard of 30 times / hour), and the lowest blood oxygen saturation was close to the threshold (fit 0.8, because 92% is slightly higher than 90%). The comprehensive calculation yielded a medical fit quantitative factor of 0.85 for the evaluation results.
[0064] Visual feature fluctuation patterns and modality fusion deviation information in the multimodal raw data are compared and correlated with historical case feature databases and parameter optimization records in the model training database to generate a data quality verification quantification factor. In-depth analysis of key information (visual feature fluctuation patterns and modality fusion deviation information) in the multimodal raw data is performed, and trend comparison and correlation analysis are conducted with historical case feature databases and parameter optimization records in the model training database to generate a data quality verification quantification factor (values range from 0 to 1, with values closer to 1 indicating higher data quality). Visual feature fluctuation patterns must conform to physiological activity patterns (e.g., respiratory rhythm fluctuations should be periodic), and modality fusion deviation information must be controlled within allowable ranges (e.g., respiratory rate deviation between visual modality and bioradar modality ≤ 2 breaths / min). By comparing with historical high-quality data features, the degree of deviation in the current data is quantified.
[0065] In the visual characteristic fluctuation pattern of the subjects, the chest respiratory rhythm showed a stable periodicity (95% consistency with high-quality characteristics of historical cases), but there were 3 physiologically insignificant abrupt changes in lip color grayscale (deviation rate of 5%). In modal fusion deviation information, the average deviation of respiratory rate between visual modality and bioradar modality was 1 time / min (≤ allowable deviation of 2 times / min), and there were no obvious abnormalities in the fused data. After trend comparison and correlation analysis, the generated data quality verification quantification factor was 0.93.
[0066] Based on the assessment results, medically adapted quantitative factors and data quality verification quantitative factors, combined with the dynamic weight allocation logic and abnormal event grading criteria in multi-dimensional sleep event association rules, sleep quality level, respiratory abnormality degree, and sleep structure integrity are fused to generate a quality-abnormality-structure correlation quantitative feature. This feature integrates the quantitative values of three core information categories: sleep quality (e.g., excellent / good / average / poor), respiratory abnormality grading (e.g., none / mild / moderate / severe), and sleep structure integrity (e.g., complete / mostly complete / incomplete), providing unified feature support for the final early warning information generation.
[0067] Example: Combining the assessment results with a medical adaptation quantitative factor of 0.85 (weight 0.6) and a data quality verification quantitative factor of 0.93 (weight 0.4), the basic fusion value is calculated by dynamic weighting as 0.85×0.6+0.93×0.4=0.878. Then, combining the abnormal event grading criteria (AHI 12 times / hour is judged as mild respiratory abnormality) and the sleep structure integrity assessment (all stages are complete and the proportion is reasonable, judged as complete), the final quality-abnormality-structure correlation quantitative feature is generated as "sleep quality quantitative value 0.88, respiratory abnormality grading quantitative value 0.3 (mild corresponds to 0.3), sleep structure integrity quantitative value 1.0 (complete corresponds to 1.0)".
[0068] This system utilizes a fusion architecture combining computer vision and AI to analyze and process the quantitative features related to sleep quality, abnormalities, and structure, generating sleep breathing state warning information that includes sleep quality ratings, quantitative indicators of respiratory abnormalities, and sleep structure characteristic parameters. The system performs in-depth analysis of the quantitative features related to sleep quality, abnormalities, and structure, and combines this with judgment logic formed through model training to generate sleep breathing state warning information containing three core types of information. Sleep quality ratings are based on quantitative sleep quality values (e.g., ≥0.8 for excellent, 0.6-0.8 for good, 0.4-0.6 for average, and <0.4 for poor). Quantitative indicators of respiratory abnormalities clearly define the type, grade, and key parameters of the abnormality (e.g., mild OSAHS, AHI 12 breaths / hour, minimum blood oxygen saturation 92%). Sleep structure characteristic parameters include quantitative data such as the proportion of each sleep stage and body movement frequency, ensuring that the warning information is accurate, comprehensive, and has clinical reference value.
[0069] Through analysis using a fusion architecture of computer vision and AI, a sleep quality quantification value of 0.88 corresponds to a grade of "excellent"; a breathing abnormality grading quantification value of 0.3 corresponds to "mild OSAHS," and a breathing abnormality quantification index is generated by combining key parameters; a sleep structure integrity quantification value of 1.0 corresponds to structural integrity, and parameters such as the proportion of each stage are output. The final generated sleep breathing state warning information is: "Sleep quality grade: excellent; breathing abnormality quantification index: mild OSAHS (AHI 12 times / hour, lowest blood oxygen saturation 92%); sleep structure characteristic parameters: N1 stage 12%, N2 stage 45%, N3 stage 20%, REM sleep 18%, wakefulness 5%, total number of body movements 32."
[0070] In one implementation, such as Figure 2 As shown, this application also provides a non-invasive and disturbance-free sleep breathing monitoring device, comprising: The acquisition module 201 is used to acquire visual data and related modal data of multiple parts of the human body during sleep. The processing module 202 is used to preprocess the collected sleep-related data. It blurs private parts of the body using a privacy protection unit, and uses a segmentation unit to achieve precise segmentation and time alignment of 10 specific body part videos. PSG tags are used for data quality control and time synchronization. The preprocessed sleep data is then subjected to detailed reasoning, analyzing pixel grayscale changes and motion amplitude characteristics of each body part video to identify sleep-related events including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement, generating multi-dimensional sleep event details. This multi-dimensional sleep event details are then comprehensively analyzed, and a sleep state assessment model is constructed by integrating multimodal data. Statistical features are used for weighted integration to achieve two to five categories of sleep stages, generating core sleep events. The comprehensive analysis results are then optimized and presented, displaying the video capture content, sleep event details, and sleep stage results in real time on the user interface, generating sleep breathing state assessment information. Finally, the sleep breathing state assessment information, multimodal raw data, and reasoning analysis process data are integrated to generate sleep breathing state early warning information including sleep quality level, quantitative indicators of respiratory abnormalities, and sleep structure characteristics.
[0071] The computer-readable storage medium provided in the above embodiments of this application and the non-invasive and interference-free sleep breathing monitoring method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application stored therein.
[0072] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating a non-invasive, undisturbed sleep breathing monitoring method, electronic device, electronic device, and readable storage medium are basically similar to the embodiments of the non-invasive, undisturbed sleep breathing monitoring method described above, and are therefore described simply. Relevant parts can be referred to in the description of the embodiments of the non-invasive, undisturbed sleep breathing monitoring method described above.
Claims
1. A non-invasive and disturbance-free method for monitoring sleep breathing, characterized in that, include: Acquire visual data and related modal data of multiple parts of the human body during sleep. The visual data of multiple parts of the human body includes 10 specific parts of video data collected by visual sensors, including the eyes, mouth, lips, Adam's apple, chest, and abdomen. The collected sleep-related data is preprocessed, and the privacy protection unit blurs the private parts of the human body. The segmentation unit is used to achieve accurate segmentation and time alignment of 10 specific parts of the video. Combined with PSG tags, data quality control and time synchronization are completed. Detailed reasoning is performed on the preprocessed sleep data to analyze the pixel grayscale changes and motion amplitude features of each part of the video, and to identify sleep-related events including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement, generating multi-dimensional sleep event details; By comprehensively analyzing detailed information on multi-dimensional sleep events and integrating multimodal data to construct a sleep state assessment model, and by combining statistical features with weighted integration, a two- to five-category sleep stage is achieved, generating core sleep events. The comprehensive analysis results are optimized and presented in real time on the user interface, displaying video capture content, sleep event details, sleep stage results, and generating sleep breathing status assessment information. By integrating sleep breathing state assessment information, multimodal raw data, and inference analysis process data, early warning information on sleep breathing state, including sleep quality level, quantitative indicators of breathing abnormalities, and sleep structure characteristics, is generated.
2. The method as described in claim 1, characterized in that, The collected sleep-related data is preprocessed, and privacy-preserving units blur private parts of the body. A segmentation unit is used to achieve precise segmentation and time alignment of 10 video feeds targeting specific body parts. PSG tags are combined to complete data quality control and time synchronization, including: The collected sleep-related data is classified and processed to generate an adaptation association result between data types and preprocessing mechanisms. The sleep-related data includes full-body RGB video data, optical flow data and optional infrared thermal imaging data collected by visual sensors, and the associated data is PSG tag data. The adaptation and association results are analyzed in a targeted manner to generate a list of core preprocessing variables, including the privacy-preserving regional positioning accuracy variable, the segmentation integrity variable of 10 video segments, the quality control and verification variable of PSG tags, and the synchronization error variable of data time alignment. The adaptation and association results, the core preprocessing variable list, and the original sleep-related data and PSG tag data are synchronously verified and screened for compliance. Combining the tag verification mechanism and cross-validation rules, privacy-protected processing results, time-aligned 10-channel body part video data, and standardized PSG tag association data are generated.
3. The method as described in claim 1, characterized in that, Detailed reasoning is performed on the preprocessed sleep data, analyzing pixel grayscale changes and motion amplitude features of different body parts in the video. Sleep-related events, including eye opening, eye movement, respiratory rhythm, changes in blood oxygen saturation, and body movement, are identified, generating multi-dimensional sleep event details, including: ISP image signal processing technology is used to extract features from the preprocessed 10-channel video data to generate pixel grayscale change sequences, motion amplitude quantization data and visual semantic feature sets for each part. Data fusion processing is performed on the feature set, sleep-related event recognition rules, physiological activity association standards of various parts, and medical annotation threshold information to establish a precise mapping relationship between features and events, and generate an event recognition fusion dataset; A multi-site sleep event classification and reasoning model is constructed based on a fusion dataset. The model uses 10 video feeds of different body parts as input dimensions, visual semantic features as core parameters, and medical annotation standards as the judgment criteria. This enables the model to classify and identify specific physiological events and generate initial sleep event details. By combining the specificity and correlation of physiological activities in various parts of the body, the initial sleep event details are calibrated with timestamps and the event types are verified to clarify the occurrence time, duration, associated parts and confidence level of various events, and generate detailed sleep event information including eye opening, respiratory rhythm and blood oxygenation changes.
4. The method as described in claim 3, characterized in that, A comprehensive analysis of detailed information on multi-dimensional sleep events is conducted, and a sleep state assessment model is constructed by integrating multimodal data. This model is then weighted and integrated using statistical features to achieve two to five categories of sleep stages, generating core sleep events, including: The system performs classification, extraction, and feature alignment processing on multidimensional sleep event details, multimodal correlation data, and medical annotation standard data to generate 10-way site event temporal feature information, multimodal complementary feature information, sleep stage determination threshold information, and core event association rule information. Based on the goal of accurate assessment through multimodal fusion, the respiratory rhythm fluctuation features and body movement frequency features in the time sequence features of 10 site events are correlated with the physiological signal features of the bio-radar / snoring modality. The weighted average integration algorithm is combined to quantify the contribution and synergistic information of various modalities in sleep assessment and generate information on the basis for multimodal data fusion. Based on the goals of sleep stage segmentation and accurate event identification, the sleep stage determination threshold and core event association rules are respectively matched with the stage level adaptation model and the event type identification optimization model to generate dynamic threshold and rule adaptation information. Based on multimodal data fusion information and dynamic threshold and rule adaptation information, the core features in multidimensional sleep event details are labeled and weighted to generate preliminary sleep state staging results and candidate core sleep events. Key physiological signal features in multimodal correlation data are labeled and substituted into a sleep event recognition model to calculate the frequency, duration and impact of events, identify core sleep event features in multimodal dimensions, integrate multi-source features by combining dynamic weighted fusion strategy, and generate preliminary identification results of core sleep events by fusion confidence calculation and abnormal data filtering. The preliminary sleep stage results and the preliminary identification results of core sleep events are integrated and processed with rules to achieve two to five categories of sleep stages and generate a comprehensive assessment result that includes core sleep events such as sleep apnea and hypoxia.
5. The method as described in claim 1, characterized in that, The comprehensive analysis results are optimized and presented, displaying video capture content, sleep event details, and sleep staging results in real time on the user interface, generating sleep breathing status assessment information, including: The comprehensive analysis results are categorized and core information is extracted to generate sleep staging results, detailed information on core sleep events, quantitative information on physiological indicators, and information on multimodal raw video clips. Based on the user's intuitive viewing target, the display interface adaptation system analyzes the user's usage scenarios and key information needs, quantifies the differences in the information presentation needs of different users, and generates scenario-based adaptation parameters. Based on real-time display and historical backtracking constraints, the information update frequency, data loading rate, and storage usage are detected and optimized to generate display performance assurance parameters, including real-time data refresh standards, historical data caching rules, and large file chunking loading schemes. In accordance with privacy and security protection requirements, the displayed video clips undergo a blurring re-examination of private areas, an information access permission grading mechanism is set up, and display privacy protection parameters are generated; By combining the characteristics of the user display interface, the scenario adaptation parameters, display performance guarantee parameters, display privacy protection parameters and interface display capabilities are associated and matched to dynamically adjust the information presentation format and display priority. The above processing results are integrated and verified. Through display effect testing, privacy compliance review and user experience simulation evaluation, a sleep breathing state assessment information display solution adapted to multiple scenarios and terminals is generated, which presents video collection content, sleep event details and sleep stage results in real time.
6. The method as described in claim 5, characterized in that, The sleep breathing state assessment information, multimodal raw data, and inference analysis process data are integrated to generate sleep breathing state early warning information that includes sleep quality level, quantitative indicators of breathing abnormalities, and sleep structure characteristics, including: The sleep staging results and core sleep event data in the sleep breathing status assessment information are compared with the OSAHS screening threshold and sleep quality assessment criteria in the medical diagnostic standards to generate a medical fit quantitative factor for the assessment results. The medical fit quantitative factor for the assessment results represents the degree of fit between the monitoring results and the clinical diagnostic needs. The visual feature fluctuation patterns and modality fusion deviation information in the multimodal raw data are compared and correlated with the historical case feature library and parameter optimization records in the model training database to generate a data quality verification quantification factor. Based on the evaluation results, medical adaptation quantitative factors and data quality verification quantitative factors, combined with the dynamic weight allocation logic and abnormal event classification judgment criteria in the multi-dimensional sleep event association rules, the sleep quality level, respiratory abnormality degree and sleep structure integrity are fused to generate quality-abnormality-structure association quantitative features. Based on a computer vision and AI fusion architecture, the quantitative features of the quality-abnormality-structure correlation are analyzed and processed to generate sleep breathing state early warning information that includes sleep quality level assessment results, quantitative indicators of breathing abnormalities, and sleep structure characteristic parameters.
7. A non-invasive and disturbance-free sleep breathing monitoring device, characterized in that, The device includes: The acquisition module is used to acquire visual data and related modal data of multiple parts of the human body during sleep. The processing module preprocesses the collected sleep-related data. It blurs private body parts using a privacy protection unit, and uses a segmentation unit to accurately segment and time-align 10 specific body part videos. PSG tags are used for data quality control and time synchronization. The module then performs detailed reasoning on the preprocessed sleep data, analyzing pixel grayscale changes and motion amplitude characteristics of each body part video to identify sleep-related events including eye opening, eye movement, respiratory rhythm, blood oxygen saturation changes, and body movement, generating multi-dimensional sleep event details. This multi-dimensional sleep event details are then comprehensively analyzed, fusing multimodal data to construct a sleep state assessment model. Statistical features are used for weighted integration to achieve two to five categories of sleep stages, generating core sleep events. The comprehensive analysis results are then optimized and presented, displaying the video capture content, sleep event details, and sleep stage results in real-time on the user interface, generating sleep breathing state assessment information. Finally, the sleep breathing state assessment information, multimodal raw data, and reasoning analysis process data are integrated to generate sleep breathing state early warning information including sleep quality level, quantitative indicators of respiratory abnormalities, and sleep structure characteristics.
8. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the non-invasive, non-disruptive sleep breathing monitoring method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements a non-invasive and interference-free method for monitoring sleep breathing as described in any one of claims 1 to 6.