Method, apparatus and computer program for generating a sleep analysis model for predicting sleep states based on acoustic information

JP2025526211A5Pending Publication Date: 2026-07-30ASLEEP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ASLEEP
Filing Date
2023-07-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Conventional sleep analysis methods using wearable devices are inconvenient, require periodic management, and struggle with accuracy when not properly attached or in multi-user environments, necessitating a technology that can accurately analyze sleep states using acoustic information from a user's environment without special equipment.

Method used

An artificial neural network model that processes sleep acoustic information through preprocessing, conversion to frequency domains, and application of deep learning techniques, including unsupervised and semi-supervised learning, to infer sleep states using a user's terminal.

Benefits of technology

Enables accurate, real-time sleep state analysis in unrestricted environments, improving sleep health management by leveraging acoustic information without the need for wearable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention provides an artificial neural network model for determining a user's sleep state based on acoustic information detected in the user's sleep environment. A method for generating the sleep analysis model may include acquiring sleep acoustic information of a user, preprocessing the sleep acoustic information, and analyzing the preprocessed sleep acoustic information to obtain sleep state information. In accordance with the present invention, acoustic information related to the sleep environment can be easily acquired through a user terminal (e.g., a mobile terminal) carried by the user, and the user's sleep stage can be analyzed based on the acquired acoustic information to determine the sleep state.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is directed to analyzing a user's sleep state, and more particularly, to analyzing a user's sleep state based on acoustic information acquired in the user's sleep environment. [Background technology]

[0002] There are various ways to maintain and improve health, such as exercise and dietary therapy, but the most important thing is to manage sleep well, which accounts for more than 30% of our daily time. However, despite the simple substitution of labor by machines and the abundance of money in our lives, modern people are unable to get a good night's sleep due to irregular eating habits, lifestyle habits, and stress, and suffer from sleep disorders such as insomnia, excessive sleep, sleep apnea syndrome, nightmares, night terrors, and sleepwalking.

[0003] According to the National Health Insurance Service, the number of people with sleep disorders in Korea increased by an average of 8% per year from 2014 to 2018, and the number of people who received treatment for sleep disorders in Korea in 2018 reached approximately 570,000.

[0004] Furthermore, according to a 2019 sleep-related survey, 62% of adults worldwide are unable to get as much sleep as they would like, and 67% of adults experience at least one sleep disorder each night. Furthermore, while 8 in 10 adults worldwide wish to improve their sleep, 60% are unable to seek help from a medical professional, and 44% of adults worldwide have experienced a decline in the quality of their sleep over the past five years.

[0005] Deep sleep is recognized as an important factor affecting physical and mental health, and interest in deep sleep is increasing. However, in order to improve sleep disorders, patients must visit specialized medical institutions in person, which requires additional testing fees, and continuous management is difficult, so users are not making enough effort to seek treatment.

[0006] As sleep problems become increasingly serious, the need for sleep health management is increasing, and the sleep tech market, which aims to solve sleep problems with technology, is also growing rapidly.

[0007] In addition, when analyzing and inferring information about sleep for sleep health management, more accurate inference is required through learning various types of data in a multimodal manner rather than using only one type of data.

[0008] Korean Patent Publication No. 10-2003-0032529 discloses a sleep induction device and sleep induction method that receives input of a user's physical information and outputs vibrations and / or ultrasound waves in a frequency band detected through repeated learning according to the user's physical condition during sleep, thereby enabling optimal sleep induction.

[0009] However, conventional technologies can reduce the quality of sleep due to the inconvenience caused by body-worn devices and require periodic management of the device (e.g., charging, etc.).Recently, research has been progressing on contactless monitoring of the user's sleep to estimate the sleep state and manage the user's sleep according to the estimated sleep state.

[0010] Recently, a method for analyzing a user's sleep using a wearable device has been proposed. Korean Patent Publication No. 10-2022-0015835 relates to an electronic device for assessing sleep quality and an operating method thereof, and proposes a method for identifying sleep cycles based on sleep-related information acquired by a wearable device during sleep, thereby assessing sleep quality.

[0011] However, conventional sleep analysis methods using wearable devices have the problem that sleep analysis is impossible if the wearable device is not properly attached to the user's body or if the user is not wearing the wearable device. Also, when multiple users sleep in the same space, the movements of those not wearing the wearable device can interfere with the sleep analysis of those wearing the wearable device, and sleep analysis of those not wearing the wearable device is impossible.

[0012] Therefore, there may be a demand for a technology that can easily acquire acoustic information related to a sleep environment through a user terminal (e.g., a mobile terminal) carried by the user without requiring additional equipment, and that can analyze the user's sleep stage based on the acquired acoustic information to detect the sleep state.

[0013] There may also be a demand for a technology that can easily yet accurately analyze sleep states without special equipment in a general, unrestricted environment (e.g., a home environment) that is not a restricted environment such as a PSG environment, a hospital environment, or a laboratory environment.

[0014] Furthermore, there may be a need for a technology that attempts to detect a user's sleep state based on at least one of sleep acoustic information and other sleep environment information.

[0015] Furthermore, there may be a need for a technology that attempts to detect a user's sleep state in real time based on at least one of sleep acoustic information and other sleep environment information. Summary of the Invention [Problem to be solved by the invention]

[0016] The present invention has been devised in response to the above-mentioned background art, and aims to provide an artificial neural network model that determines a user's sleep state based on at least one of acoustic information sensed in the user's sleep environment and sleep environment information.

[0017] The problems to be solved by the present invention are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]

[0018] In one embodiment of the present invention to solve the above-mentioned problems, a method for generating a sleep analysis model for predicting sleep states based on acoustic information is disclosed.

[0019] The method for generating the sleep analysis model may include acquiring sleep acoustic information of a user, performing preprocessing on the sleep acoustic information, and analyzing the preprocessed sleep acoustic information to acquire sleep state information.

[0020] In one embodiment of the present invention, obtaining the sleep state information may include obtaining the sleep state information using a sleep analysis model that includes one or more network functions.

[0021] In one embodiment of the present invention, a method for generating a sleep analysis model may include converting raw time-domain acoustic information having amplitude, phase, and frequency into information including changes in frequency components over time, where the converted information may be visualized according to one embodiment of the present invention.

[0022] Alternatively, in one embodiment of the present invention, a step of converting raw acoustic information in the time domain into information in the frequency domain having amplitude and frequency may be included, where, according to one embodiment of the present invention, the converted information in the frequency domain may be visualized.

[0023] In one embodiment of the present invention, the method may include converting acoustic information into spectrogram information in the frequency domain.

[0024] In one embodiment of the present invention, the method may include applying a mel scale to the spectrogram to transform it into a mel spectrogram.

[0025] In one embodiment of the present invention, a step of performing pre-processing on raw acoustic information in the time domain or information in the frequency domain may be included.

[0026] In an embodiment of the present invention, performing pre-processing on the acoustic information may include performing spectral noise gating or deep learning-based noise reduction.

[0027] In one embodiment of the present invention, the method may include performing data augmentation on frequency domain information.

[0028] In one embodiment of the present invention, the data augmentation may include performing one or more of pitch shifting, Tile UnTile (TUT) augmentation, or noise-added augmentation.

[0029] In one embodiment of the present invention, the noise augmentation may include a method of converting noise information and sleep sound information into frequency domains, respectively, and adding information in the frequency domain.

[0030] In one embodiment of the present invention, the noise augmentation may include a method of converting noise information and sleep sound information into spectrograms, respectively, and adding information in the domain converted into the spectrograms.

[0031] In one embodiment of the present invention, the noise-added augmentation may include a method of adding information in a domain transformed into a mel spectrogram in which a mel scale is applied to the sleep acoustic information and the noise information.

[0032] In one embodiment of the present invention, a step of converting information in the frequency domain into a form that is close to a square may be included.

[0033] In an embodiment of the present invention, the converting into a shape close to a square may include at least one of reshaping, resizing, and split-cat methods.

[0034] In one embodiment of the present invention, the step of converting information in the frequency domain or spectrogram into a dB scale (log scale) may be included.

[0035] In one embodiment of the present invention, the method may include performing a normalization process on the information in the frequency domain or the spectrogram so that the average of all values becomes 0 and the standard deviation becomes 1.

[0036] In one embodiment of the present invention, the method may include a step of inputting frequency domain information, spectrogram or mel-spectrogram information as an image to an artificial intelligence model.

[0037] In one embodiment of the present invention, the method may include dividing the frequency domain information, spectrogram or mel spectrogram into 30-second units to generate a plurality of frequency domain information, spectrograms or mel spectrograms.

[0038] In one embodiment of the present invention, the method may include extracting sleep state information corresponding to information obtained by dividing frequency domain information, spectrograms, or mel-spectrograms into 30-second units.

[0039] In one embodiment of the present invention, the method may include extracting sleep state information using frequency domain information divided into 30-second units, a series of information consisting of multiple spectrograms or mel-spectrograms as input to a deep learning model.

[0040] In one embodiment of the present invention, the method may include a step of using frequency domain information, spectrograms or mel-spectrograms containing time series information as input to an artificial intelligence model and outputting a vector with reduced dimensions.

[0041] In one embodiment of the present invention, the method may include inputting the reduced-dimensional vector to an artificial intelligence model and outputting a vector containing time-series information.

[0042] In one embodiment of the present invention, the method may include inputting the reduced-dimensional vector to an intermediate layer and outputting a vector containing time-series information.

[0043] In one embodiment of the present invention, an intermediate layer to which a reduced-dimensional vector is input may include at least one model that performs linearization that implies vector information, normalization to input the mean and variance, or a dropout step that deactivates some nodes.

[0044] In one embodiment of the present invention, a method may be included that utilizes an unsupervised learning model that learns using unlabeled data in which a correct answer is not labeled.

[0045] In one embodiment of the present invention, the unsupervised learning model utilized may include a consistency training model using noise in the target environment.

[0046] In one embodiment of the present invention, the consistency training model utilized may include a step of performing learning using data to which noise has been intentionally added and data to which noise has not been intentionally added.

[0047] In one embodiment of the present invention, the unsupervised learning model utilized may include an Unsupervised Domain Adaptation (UDA) model.

[0048] The UDA model utilized in one embodiment of the present invention can perform a first learning using unlabeled data and labeled data, and a second learning using unlabeled data.

[0049] In the first learning of a UDA model according to an embodiment of the present invention, learning can be performed using labeled data acquired in a specific environment and unlabeled data acquired in another environment or target environment.

[0050] The primary learning of a UDA model according to one embodiment of the present invention may include a step of learning data acquired in a specific environment and data acquired in another environment or target environment as inputs to a sleep analysis model to extract commonalities between the data.

[0051] The primary learning of the UDA model according to one embodiment of the present invention may include a step of learning data acquired in a specific environment and data acquired in another environment or target environment, respectively, to classify differences between the input data by using commonalities between the output data as inputs to a discriminator model.

[0052] The secondary learning of the UDA model according to one embodiment of the present invention may include a step of learning the class information including the predicted value of the sleep state information output from the sleep analysis model using unlabeled data as an input of the deep learning model to make it more reliable.

[0053] Semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include a step of performing learning of a deep learning model by using unlabeled data as input to the deep learning model and output data as pseudo labels.

[0054] In one embodiment of the present invention, semi-supervised learning using utilized pseudo labels may include a step of performing augmentation preprocessing on images.

[0055] The augmentation preprocessing method performed by semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include at least one of a weakly-augmented method that modulates an image relatively less or a strongly-augmented method that modulates an image relatively more.

[0056] The augmentation preprocessing technique performed in semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include one or more of data augmentation, pitch shifting augmentation, Tile UnTile (TUT) augmentation, or noise-added augmentation techniques.

[0057] Semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include a method of performing learning based on image information that has been subjected to a weakly-augmented method that modulates an image relatively little, by utilizing the predicted values output as pseudo labels as input to a deep learning model.

[0058] Semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include one or more of a moving average technique, a weighted average technique, a weighted moving average technique, or an exponential weighted moving average technique.

[0059] Semi-supervised learning using pseudo labels utilized in one embodiment of the present invention may include a step of tuning such that a distribution of predicted values output from a deep learning model using data acquired from a target environment or a target subject population as input is formed in a direction consistent with a distribution of predicted values output from a deep learning model using data acquired from a specific environment or a comparison subject population as input.

[0060] The unsupervised learning and / or semi-supervised learning utilized in an embodiment of the present invention may include a method of performing pre-learning so that the reliability of prediction values for image information can be increased even if the image data does not have labels in the image domain.

[0061] The unsupervised learning and / or semi-supervised learning utilized in an embodiment of the present invention may include a method of performing learning so that a portion of image information can be predicted after damaging a portion of the image information.

[0062] Furthermore, according to one embodiment of the present invention, a method for analyzing sleep state information using both sleep acoustic information and sleep environment information in a multimodal manner may include: a first information acquiring step of acquiring time-domain acoustic information related to a user's sleep; a second information acquiring step of acquiring user sleep environment information related to the user's sleep; a step of combining the first information and the second information as multimodal data; a step of extracting features from the multimodal data as an input of a multimodally trained deep learning model; and a step of acquiring user sleep state information using the extracted features as an input of the deep learning model.

[0063] Also, the step of acquiring the first information may include a step of performing preprocessing on the acquired first information, and the step of acquiring the second information may include a step of performing preprocessing on the acquired second information.

[0064] The step of performing pre-processing of the first information may include a step of extracting first information features based on the first information, and the step of performing pre-processing of the second information may include a step of extracting second information features based on the second information.

[0065] In addition, the step of performing pre-processing of the first information may include the step of performing data augmentation of the first information, and the step of performing pre-processing of the second information may include the step of performing data augmentation of the second information.

[0066] The deep learning model in the step of acquiring the user's sleep state information may be a deep learning model based on natural language processing.

[0067] Also, performing pre-processing on the first information may include converting the first information in the time domain into information in the frequency domain.

[0068] According to one embodiment of the present invention, a non-transitory computer-readable storage medium may be provided that stores one or more programs configured to be executed by one or more processors to analyze sleep state information including sleep acoustic information and sleep environment information in a multimodal manner, the one or more programs including instructions for performing the above-described method.

[0069] According to one embodiment for achieving the object of the present invention, there is provided a smart device for analyzing multimodal sleep state information including sleep acoustic information and sleep environment information, the device including: a first information acquisition unit that acquires time-domain acoustic information related to a user's sleep; a second information acquisition unit that acquires user sleep environment information related to the user's sleep; a data combination unit that combines the first information and the second information as multimodal data; a feature inference unit that infers features using the multimodal representation as an input of a trained deep learning model; and a user sleep state information acquisition unit that acquires sleep state information using the inferred features as an input of the deep learning model.

[0070] wherein a first information preprocessing unit that performs preprocessing of the first information; The information processing device may further include a second information pre-processing unit that performs pre-processing of the second information.

[0071] The unit that performs pre-processing of the first information can extract first information features based on the first information, and the unit that performs pre-processing of the second information can extract second information features based on the second information.

[0072] In addition, the first information pre-processing unit may perform data augmentation of the first information, and the second information pre-processing unit may perform data augmentation of the second information.

[0073] The artificial intelligence model for analyzing the user's sleep state information may be an artificial intelligence model based on natural language processing.

[0074] In addition, the pre-processing unit for performing the first information pre-processing may convert the first information in the time domain into information in the frequency domain, where the information in the frequency domain may be information including a change in frequency components included in the first information in the time domain over a time axis.

[0075] To achieve the object of the present invention, according to one embodiment, a method for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information may include: a first information acquisition step of acquiring sleep acoustic information related to a user's sleep; a step of inferring first sleep state information using the sleep acoustic information as an input of a deep learning model; a second information acquisition step of acquiring user sleep environment information related to the user's sleep; a step of inferring second sleep state information using the user sleep environment information as an input of an inference model; and a user sleep state information acquisition step of combining the first sleep state information and the second sleep state information to acquire user sleep state information.

[0076] The first information obtaining step may convert the first information in the time domain into information in the frequency domain.

[0077] In addition, the step of inferring the first sleep state information may infer a hypnogram that displays sleep stages or a hypnodensity graph that displays the reliability of sleep stages as a probability as the first sleep state information.

[0078] The step of inferring the second sleep state information may be characterized by inferring a hypnogram that displays sleep stages or a hypnodensity graph that displays the reliability of sleep stages as a probability as the second sleep state information.

[0079] In addition, there may be provided a method for analyzing sleep state information using sleep acoustic information and sleep environment information in a multimodal manner, wherein the user sleep state information acquisition step further includes a sleep state information combination step of combining the first sleep state information and the second sleep state information.

[0080] The step of acquiring the user's sleep state information may further include a step of performing sleep state data augmentation using the first sleep state information and the second sleep state information.

[0081] Here, the inference model may be an artificial intelligence sleep information inference model.

[0082] The user sleep state information acquiring step may infer the user sleep state information through an artificial intelligence learning model to acquire the user sleep state information.

[0083] As one embodiment for achieving the object of the present invention, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information, wherein the one or more programs include instructions for performing the above-mentioned method, may be provided.

[0084] According to one embodiment of the present invention, there is provided a smart device for analyzing multimodal sleep state information including sleep acoustic information and sleep environment information, the smart device including a first information acquisition unit that acquires sleep acoustic information related to a user's sleep, a first information inference unit that infers first sleep state information using the sleep acoustic information as an input of a deep learning model, a second information acquisition unit that acquires user sleep environment information related to the user's sleep, a second sleep state information inference unit that infers second sleep state information using the user sleep environment information as an input of an inference model, and a user sleep state information acquisition unit that combines the first sleep state information and the second sleep state information to acquire user sleep state information.

[0085] Here, according to an embodiment of the present invention, the first information acquirer may be characterized in that it converts the first information in the time domain into information in the frequency domain.

[0086] Furthermore, according to one embodiment of the present invention, the first sleep state information inference unit may be characterized by inferring a hypnogram that displays sleep stages or a hypnodensity graph that displays the reliability of sleep stages as a probability as the first sleep state information.

[0087] According to one embodiment of the present invention, the second sleep state information inference unit may be characterized by inferring a hypnogram that displays sleep stages or a hypnodensity graph that displays the reliability of sleep stages as a probability as the second sleep state information.

[0088] According to an embodiment of the present invention, the user sleep state information acquisition unit may include a sleep state information combination unit that combines the first sleep state information and the second sleep state information.

[0089] According to an embodiment of the present invention, the user sleep state information acquisition unit may further include a sleep state data augmentation unit that performs data augmentation using the first sleep state information and the second sleep state information.

[0090] Furthermore, according to an embodiment of the present invention, the inference model may be an artificial intelligence sleep information inference model.

[0091] According to one embodiment of the present invention, the user sleep state information acquisition unit may infer the user sleep state information through an artificial intelligence learning model to acquire the user sleep state information.

[0092] Meanwhile, according to one embodiment of the present invention, there is provided a method for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information, the method including: a first information acquisition step of acquiring sleep acoustic information related to a user's sleep; a step of inferring first sleep state information using the sleep acoustic information as an input of a deep learning model; a second information acquisition step of acquiring user sleep environment information related to the user's sleep; and a user sleep state information acquisition step of combining the first sleep state information and the second information to acquire user sleep state information.

[0093] According to an embodiment of the present invention, the second information may be user information obtained via a smart watch.

[0094] According to an embodiment of the present invention, inferring the first sleep state information may include performing pre-processing on the acquired first sleep information.

[0095] According to an embodiment of the present invention, the step of pre-processing the first information may include converting the first information in the time domain into information in the frequency domain.

[0096] According to an embodiment of the present invention, the acquiring of the user's sleep state information may further include combining the first sleep state information and the second sleep state information.

[0097] According to an embodiment of the present invention, the step of acquiring the user's sleep state information may further include a step of performing sleep state data augmentation using the first sleep state information and the second information.

[0098] Furthermore, according to one embodiment of the present invention, the user sleep state information acquisition step may infer the user sleep state information through an artificial intelligence learning model to acquire the user sleep state information.

[0099] Meanwhile, in order to analyze sleep state information using multimodal sleep acoustic information and sleep environment information according to one embodiment of the present invention, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors may be provided, wherein the one or more programs include instructions to perform one or more of the above-mentioned methods.

[0100] In addition, according to one embodiment of the present invention, there is provided an apparatus for analyzing sleep state information in which sleep acoustic information and sleep environment information are multimodal, the apparatus including: a first information acquisition unit that acquires sleep acoustic information related to a user's sleep; a first sleep state information inference unit that uses the sleep acoustic information as an input of a deep learning model to acquire user sleep environment information related to the user's sleep; and a user sleep state information acquisition unit that combines the first sleep state information and the second information to acquire user sleep state information.

[0101] According to an embodiment of the present invention, the second information may be user information obtained via a smart watch.

[0102] According to an embodiment of the present invention, the first sleep state information inferring unit may include a first sleep information pre-processing unit for performing pre-processing on the acquired first sleep information.

[0103] According to an embodiment of the present invention, the first information pre-processing unit may convert the first information in the time domain into information in the frequency domain.

[0104] According to an embodiment of the present invention, the user sleep state information acquisition unit may further include a sleep state information combination unit that combines the first sleep state information and the second information.

[0105] According to an embodiment of the present invention, the user sleep state information acquisition unit may further include a sleep state data augmentation unit that performs data augmentation using the first sleep state information and the second information.

[0106] Furthermore, according to one embodiment of the present invention, the user sleep state information acquisition unit may infer the user sleep state information through an artificial intelligence learning model in order to acquire the user sleep state information.

[0107] According to an embodiment of the present invention, a method for detecting a real-time sleep event based on at least one of sleep acoustic information and sleep environment information may be provided.

[0108] Additionally, according to an embodiment of the present invention, a method for training a sleep analysis artificial intelligence model by unsupervised or semi-supervised learning methods may be provided.

[0109] Here, the learning method according to the embodiment of the present invention may include semi-supervised learning based on sequential consistency loss, and may be trained to take into account the time-series characteristics of the acoustic information through the semi-supervised learning based on sequential consistency loss according to the embodiment of the present invention.

[0110] Alternatively, a learning method according to an embodiment of the present invention may include learning based on a semi-supervised control loss. The learning based on the semi-supervised control loss according to an embodiment of the present invention may include setting a threshold value of class confidence and adjusting a position in a vector space based on anchor data based on the set threshold value of class confidence.

[0111] Here, the anchor data used in learning based on semi-supervised control loss according to an embodiment of the present invention may include at least one of labeling data to which labels for sleep states are assigned or pseudo-label data to which pseudo-labels for sleep states are assigned.

[0112] Furthermore, according to an embodiment of the present invention, a deep learning model may be provided for analyzing a user's sleep state information through multi-task learning based on acoustic information.

[0113] Here, for multi-task learning, the deep learning model may have multiple heads, and in this case, each head included in the multiple heads may perform a different task from the multiple tasks.

[0114] According to an embodiment of the present invention, the plurality of tasks may include tasks such as multimodal learning, sleep event analysis, and sleep stage analysis.

[0115] Other details of the invention are included in the detailed description and drawings. [Effects of the Invention]

[0116] The present invention has been devised in response to the above-mentioned background art, and can provide an artificial neural network model that determines a user's sleep state based on at least one of acoustic information sensed in the user's sleep environment and sleep environment information.

[0117] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]

[0118] [Figure 1a] FIG. 1a is a conceptual diagram illustrating a system in which various aspects can be implemented for generating a sleep analysis model that predicts a sleep state based on information related to an embodiment of the present invention. [Figure 1b] FIG. 1b is a conceptual diagram illustrating a system in which various aspects can be implemented for generating a sleep analysis model that predicts a sleep state based on information related to an embodiment of the present invention. [Figure 1c] FIG. 1c is a conceptual diagram illustrating a system in which various aspects can be implemented for generating a sleep analysis model that predicts a sleep state based on information related to an embodiment of the present invention. [Figure 2] FIG. 2 is an exemplary diagram illustrating a plurality of pieces of sleep acoustic information acquired in various environments. [Figure 3] FIG. 3 is a block diagram of a computing device for generating the sleep analysis model for predicting sleep states based on information associated with an embodiment of the present invention. [Figure 4] FIG. 4 is a diagram illustrating a process of acquiring sleep acoustic information in the sleep analysis method according to the present invention. [Figure 5] FIG. 5 is a conceptual diagram illustrating a privacy protection method using mel-spectrogram transformation for sleep acoustic information extracted from a user in the sleep analysis method according to the present invention. [Figure 6]FIG. 6 is a diagram illustrating a method for acquiring a spectrogram corresponding to sleep acoustic information in the sleep analysis method according to the present invention. [Figure 7] FIG. 7 is a schematic diagram illustrating one or more network functions for performing a sleep analysis method according to the present invention. [Figure 8] FIG. 8 is a diagram for explaining sleep stage analysis using a spectrogram in the sleep analysis method according to the present invention. [Figure 9] FIG. 9 is a diagram illustrating sleep event determination using a spectrogram in the sleep analysis method according to the present invention. [Figure 10] FIG. 10 is a diagram showing an experimental process for verifying the performance of the sleep analysis method according to the present invention. [Figure 11] FIG. 11 is a graph verifying the performance of the sleep analysis method according to the present invention, comparing polysomnography (PSG) results with analysis results using the AI algorithm according to the present invention. [Figure 12] FIG. 12 is a graph verifying the performance of the sleep analysis method according to the present invention, which compares polysomnography (PSG) results with analysis results using the AI algorithm according to the present invention in relation to sleep apnea and hypopnea. [Figure 13] FIG. 13 is a schematic diagram of data 3 according to one embodiment of the present invention. [Figure 14] FIG. 14 is a diagram illustrating noise reduction according to an embodiment of the present invention. [Figure 15] FIG. 15 is a diagram illustrating pitch shifting according to an embodiment of the present invention. [Figure 16] FIG. 16 is a diagram illustrating a pre-processing method for converting information in the frequency domain or a spectrogram into a form close to a square, according to an embodiment of the present invention. [Figure 17a] FIG. 17a is a diagram illustrating the overall structure of a sleep analysis model according to an embodiment of the present invention. [Figure 17b] FIG. 17b is a diagram illustrating the overall structure of a sleep analysis model according to an embodiment of the present invention. [Figure 18] FIG. 18 is a diagram illustrating a feature extraction model and a feature classification model according to an embodiment of the present invention. [Figure 19] FIG. 19 is a diagram for explaining in detail the operation of a sleep analysis model according to an embodiment of the present invention. [Figure 20] FIG. 20 is a diagram illustrating an unsupervised or semi-supervised learning model according to an embodiment of the present invention. [Figure 21] FIG. 21 is a diagram illustrating consistency training according to an embodiment of the present invention. [Figure 22] FIG. 22 is a diagram illustrating an Unsupervised Domain Adaptation (UDA) according to an embodiment of the present invention. [Figure 23] FIG. 23 is a diagram illustrating a TUT (Tile UnTile) augmentation method according to an embodiment of the present invention. [Figure 24] FIG. 24 is a diagram illustrating the structure of a sleep analysis model using a natural language processing model according to an embodiment of the present invention. [Figure 25] FIG. 25 is a flow chart illustrating an example of a method for analyzing a user's sleep state through acoustic information according to an embodiment of the present invention. [Figure 26] FIG. 26 is a flowchart illustrating a method for analyzing a sleep state according to an embodiment of the present invention, which includes a process of combining sleep acoustic information and sleep environment information as multimodal data. [Figure 27]FIG. 27 is a flowchart illustrating a method for analyzing sleep states according to an embodiment of the present invention, which includes combining inferred sleep acoustic information and inferred sleep environment information into multimodal data. [Figure 28] FIG. 28 is a flowchart illustrating a method for analyzing sleep state according to an embodiment of the present invention, which includes combining inferred sleep acoustic information with sleep environment information and multimodal data. [Figure 29a] FIG. 29a is a diagram illustrating the performance of sleep event determination of a noise addition and sleep analysis model in a sleep analysis method according to an embodiment of the present invention. [Figure 29b] 29b is a diagram illustrating the performance of sleep event determination of a noise addition and sleep analysis model in a sleep analysis method according to an embodiment of the present invention. [Figure 30] FIG. 30 is a diagram illustrating an example of a training method based on consistency loss or sequential consistency loss when the number of samples in a sequence is six, according to an embodiment of the present invention. [Figure 31] FIG. 31 is an exemplary diagram illustrating the operation mechanism of a learning method based on semi-supervised contrastive loss according to an embodiment of the present invention. [Figure 32] FIG. 32 is a table comparing the analysis results of a sleep analysis model according to an embodiment of the present invention with the analysis results of a PSG test in a home environment. [Figure 33] FIG. 33 is a table comparing the sleep analysis results based on PSG audio data with the analysis results of a sleep analysis model according to one embodiment of the present invention. [Figure 34] FIG. 34 is a diagram illustrating a linear regression analysis function utilized to analyze sleep events occurring during sleep, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0119] Overall structure

[0120] The advantages and features of the present invention, and methods for achieving them, will become apparent from the following detailed description of the embodiments in conjunction with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, and may be embodied in various different forms. The present embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully convey the scope of the present invention to those skilled in the art. The present invention is defined only by the scope of the claims.

[0121] The terms used in this specification are for the purpose of describing embodiments and are not intended to limit the present invention. In this specification, the singular includes the plural unless otherwise specified. The terms "comprises" and / or "comprising" used in this specification do not exclude the presence or addition of one or more other elements other than the elements listed. The same reference numerals refer to the same elements throughout this specification, and "and / or" includes each and every combination of one or more of the listed elements. Although terms such as "first," "second," etc. are used to describe various elements, these elements are not limited by these terms. These terms are used merely to distinguish one element from another. Therefore, a first element referred to below may of course be a second element within the technical spirit of the present invention.

[0122] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in the sense commonly understood by those of ordinary skill in the art to which the present invention pertains. Furthermore, commonly used and predefined terms are not to be interpreted ideally or excessively unless expressly defined otherwise.

[0123] The terms "module" and "module" used herein refer to software or hardware components, such as FPGAs or ASICs, that perform a certain function. However, "module" and "module" are not limited to software or hardware. A "module" or "module" may be configured to reside on an addressable storage medium or to execute on one or more processors. Thus, by way of example, a "module" or "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The components and functionality provided within a "module" or "module" may be combined into fewer components and "modules" or "modules" or further separated into additional components and "modules" or "modules."

[0124] In this specification, the term "computer" refers to any type of hardware device including at least one processor, and may also encompass software configurations operating on such hardware devices according to embodiments. For example, the term "computer" may refer to, but is not limited to, smartphones, tablet PCs, desktops, laptops, and user clients and applications running on each device.

[0125] Those skilled in the art will further appreciate that the various illustrative logical blocks, components, modules, circuits, means, logic, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or any combination of both. To clearly illustrate the interchangeability of hardware and software, the various illustrative components, blocks, components, means, logic, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in various ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0126] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0127] Although each step described in this specification is described as being performed by a computer, the subject of each step is not limited to this, and depending on the embodiment, at least some of each step may be performed by different devices.

[0128] 1a to 1c are conceptual diagrams illustrating a system in which various aspects of a method for generating a sleep analysis model for predicting a sleep state based on information related to an embodiment of the present invention can be implemented.

[0129] FIG. 3 is a block diagram of a computing device for generating the sleep analysis model for predicting sleep states based on information related to an embodiment of the present invention.

[0130] A system according to an embodiment of the present invention may include a computing device 100, a user terminal 10, an external server 20, and a network.

[0131] Here, the device shown in FIG. 1a is merely one example of a system for embodying the present invention, and its configuration is not limited to the embodiment shown in FIG. 1a, and may be added, changed, or deleted as necessary.

[0132] Meanwhile, FIGS. 1b and 1c are conceptual diagrams illustrating systems in which various aspects of a sleep analysis method according to another embodiment of the present invention may be implemented.

[0133] First, a system according to the embodiment shown in FIG. 1a will be described.

[0134] As shown in FIG. 1a, a computing device 100, a user terminal 10, and an external server 20 can mutually transmit and receive data for a system according to an embodiment of the present invention via a network.

[0135] According to an embodiment of the present invention, the computing device 100 or the external server 20 may be a server that provides a cloud computing service. More specifically, the computing device 100 or the external server 20 may be a server that provides a cloud computing service, which is a type of Internet-based computing, in which information is processed by another computer connected to the Internet, not the user's computer. The cloud computing service may be a service that stores data on the Internet and allows users to access the data or programs they need anytime and anywhere via an Internet connection without having to install them on their own computers. Data stored on the Internet can be easily shared and transmitted with simple operations and clicks.

[0136] In addition, cloud computing services are not just services that simply store materials on a server on the Internet, but also services that allow desired tasks to be performed using the functions of application programs provided on the web without the need to install a separate program, and that allow various people to share documents and work together at the same time.

[0137] Furthermore, the cloud computing service may be implemented in at least one form of Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Software as a Service (SaaS), a virtual machine-based cloud server, and a container-based cloud server. That is, the computing device 100 or the external server 20 of the present invention may be implemented in at least one form of the above-mentioned cloud computing services. The specific descriptions of the above-mentioned cloud computing services are merely examples, and may include any platform that constructs the cloud computing environment of the present invention.

[0138] Networks according to embodiments of the present invention can use various wired communication systems such as Public Switched Telephone Network (PSTN), x Digital Subscriber Line (xDSL), Rate Adaptive DSL (RADSL), Multi Rate DSL (MDSL), Very High Speed DSL (VDSL), Universal Asymmetric DSL (UADSL), High Bit Rate DSL (HDSL), and Local Area Network (LAN). In addition, the networks presented herein can use various wireless communication systems such as Code Division Multi Access (CDMA), Time Division Multi Access (TDMA), Frequency Division Multi Access (FDMA), Orthogonal Frequency Division Multi Access (OFDMA), Single Carrier-FDMA (SC-FDMA), and other systems.

[0139] A network according to an embodiment of the present invention may be configured regardless of its communication mode, such as wired or wireless, and may be configured as various communication networks, such as a short-range communication network (PAN: Personal Area Network) or a short-range communication network (WAN: Wide Area Network). The network may be the well-known World Wide Web (WWW), or may use wireless transmission technologies used for short-range communication, such as Infrared Data Association (IrDA) or Bluetooth. The technology described herein may be used in other networks as well as the networks mentioned above.

[0140] According to an embodiment of the present invention, the user terminal 10 may refer to a terminal carried by a user that can receive information related to the user's sleep through information exchange with the computing device 100. For example, the user terminal 10 may be a terminal associated with a user who wishes to improve their health through information related to their sleep habits.

[0141] The user terminal 10 may refer to any type of entity in a system having a mechanism for communicating with an external server 20 or a computing device 100. For example, the user terminal 10 may include a personal computer (PC), a notebook computer, a mobile terminal, a smartphone, a tablet PC, an AI speaker, an AI TV, a wearable device, a home appliance, etc., and may include all types of terminals that can be connected to a wired / wireless network. The user terminal 10 may also include any server implemented by at least one of an agent, an application programming interface (API), and a plug-in. The user terminal 10 may also include an application source and / or a client application.

[0142] According to an embodiment of the present invention, the external server 20 may be a server that stores information on a plurality of training data for training a neural network. Alternatively, the external server 20 may be a digital device equipped with a processor, memory, and computing capabilities, such as a laptop computer, notebook computer, desktop computer, web pad, or mobile phone. The external server 20 may be a web server that processes services. The above-mentioned types of servers are merely examples, and the present invention is not limited thereto. The plurality of training data may include, for example, sleep acoustic information acquired from a plurality of user terminals, or health checkup information and sleep screening information acquired in a hospital. A detailed description of the training dataset will be provided below.

[0143] According to an embodiment of the present invention, the external server 20 may be at least one of a hospital server and a government server, and may be a server that stores information on a plurality of sleep polymorphism test records, electronic health records, electronic medical records, etc. For example, the sleep polymorphism test records may include information on the breathing and movement of a sleep examination subject during sleep, and information on corresponding sleep diagnosis results (e.g., sleep stages, etc.). The information stored in the external server 20 may be used as training data, verification data, and test data for training the neural network of the present invention.

[0144] Furthermore, the external server 20 according to an embodiment of the present invention may store an artificial intelligence model for analyzing sleep state information. In this case, if sleep environment information is acquired from the user terminal 10 or the like and transmitted to the external server, sleep state information can be generated based on the sleep environment information through the artificial intelligence model installed in the external server.

[0145] Alternatively, according to an embodiment of the present invention, the user terminal 10 may acquire sleep environment information and then acquire sleep sound information through pre-processing of the sleep environment information. The acquired sleep sound information may then be transmitted to an external server, and the external server may generate sleep state information based on the received sleep sound information.

[0146] The computing device 100 of the present invention can receive a plurality of pieces of sleep acoustic information, health check information, sleep screening information, etc. from the external server 20 and construct a training data set based on the received information. The computing device 100 can generate a sleep analysis model that calculates sleep state information corresponding to the sleep acoustic information by performing training on one or more network functions using the training data set.

[0147] According to an embodiment of the present invention, at least one of the user terminal 10, the computing device 100, and the external server 20 may generate a sleep analysis model. The sleep analysis model may be a neural network model that predicts information about a user's sleep state based on information related to the user's sleep acoustics non-invasively acquired during the user's sleep. At least one of the electronic devices according to an embodiment of the present invention may generate a sleep analysis model that receives the user's sleep acoustic information as input and outputs the user's sleep state information. A configuration for constructing a learning dataset for neural network learning, a learning method using the learning dataset, and generation and learning of a sleep analysis model according to the present invention will be described in detail below.

[0148] According to an embodiment of the present invention, a user can obtain monitoring information related to their sleep through the user terminal 10. When at least one of the electronic devices according to an embodiment of the present invention obtains or receives sleep sound information, the electronic device processes the sleep sound information as an input of a sleep analysis model, and outputs sleep state information through the sleep analysis model.

[0149] Meanwhile, according to an embodiment of the present invention, sleep acoustic information acquired from an electronic device such as a user terminal 10 may have a low signal-to-noise ratio (SNR). Generally, a microphone module installed in a user terminal 10 carried by a user may be configured with a micro-electromechanical system (MEMS) since it must be installed in a relatively small user terminal 10. The microphone module installed in the user terminal 10 may be, for example, a common microphone (a low-performance, small microphone). While such a microphone module can be manufactured to be very small, it may have a lower signal-to-noise ratio than a condenser microphone or a dynamic microphone. A low signal-to-noise ratio means that a high proportion of noise, which is an acoustic that prevents the desired acoustic ratio from being identified, makes it difficult to identify the acoustic (i.e., unclear). Such sleep acoustic information is information about very small sounds (i.e., acoustics that are difficult to distinguish), such as the user's breathing and movement, and is acquired together with other sounds in the sleep environment. Therefore, when the sleep acoustic information is acquired through the above-described microphone module (i.e., a microphone module having a low signal-to-noise ratio), it may be very difficult to derive and analyze the information.

[0150] Therefore, according to one embodiment of the present invention, at least one of various electronic devices can convert and / or adjust sleep audio data that is unclearly acquired and contains a lot of noise into data that can be analyzed, and can perform learning on an artificial neural network using the converted and / or adjusted data.

[0151] When pre-training of the artificial neural network is completed, the trained neural network (e.g., an acoustic analysis model) can acquire the user's sleep state information based on (e.g., transformed and / or adjusted) data acquired corresponding to the sleep acoustic information (e.g., information including changes in frequency components over time contained in the raw sleep acoustic information, information in the frequency domain, or a spectrogram).

[0152] According to an embodiment of the present invention, when sleep sound information having a low signal-to-noise ratio is acquired through a commonly used user terminal for collecting sound (e.g., an AI speaker, a bedroom IoT device, a mobile phone, etc.), the computing device 100 processes the acquired sleep sound information into data suitable for analysis and processes the processed data to provide sleep state information related to changes in sleep stages. This eliminates the need for a microphone that contacts the user's body to acquire clear sound, and also provides the advantage of increasing convenience by enabling sleep state monitoring in a general home environment with only a software update, without the need to purchase a separate additional device with a high signal-to-noise ratio.

[0153] Sleep-related monitoring information may include, for example, sleep state information related to when a user falls asleep, how long they are asleep, when they wake up, etc., or sleep stage information specifically related to changes in sleep stages during sleep.

[0154] For example, the sleep stage information may refer to information indicating whether the user's sleep changed to light sleep, normal sleep, deep sleep, REM sleep, etc. The above-described specific descriptions of the sleep stage information are merely examples, and the present invention is not limited thereto.

[0155] Meanwhile, as shown in FIG. 1b, the sleep analysis according to the present invention may be performed in the user terminal 10 or the external server 20 without a separate computing device.

[0156] Meanwhile, FIG. 1c shows a conceptual diagram illustrating a system in which various aspects of various electronic devices related to yet another embodiment of the present invention can be implemented.

[0157] The electronic device shown in FIG. 1c may perform at least one of the operations performed by various devices according to embodiments of the present invention.

[0158] For example, operations performed by various devices according to embodiments of the present invention may include an operation of acquiring sleep environment information or environmental sensing information, an operation of learning a sleep analysis model, an operation of inferring a sleep state through the sleep analysis model, and an operation of acquiring sleep state information.

[0159] Or, for example, it may include operations such as receiving information related to the user's sleep or sleep environment information, sending or receiving environmental sensing information, discriminating environmental sensing information, processing or manipulating data, processing services, providing services, analyzing sleep states, constructing a learning data set based on information related to the user's sleep, storing acquired data or information on multiple learning data for neural network training, sending or receiving various information, and mutually transmitting and receiving data for the system according to an embodiment of the present invention via a network.

[0160] The electronic device shown in FIG. 1c may individually perform the operations performed by the various devices according to the embodiments of the present invention, or may perform one or more operations simultaneously or sequentially.

[0161] 1c, electronic devices (reference numerals 1a to 1d) may be electronic devices within a range of an area 11a capable of acquiring object state information such as information about a user's movement or breathing. Hereinafter, for convenience, the area 11a capable of acquiring object state information or environmental sensing information such as information about a user's movement or breathing will be referred to as "area 11a."

[0162] On the other hand, referring to FIG. 1c, the electronic devices (reference numerals 1a and 1d) may be a device consisting of a combination of two or more electronic devices.

[0163] Meanwhile, referring to FIG. 1c, electronic devices (reference numerals 1a and 1b) may be electronic devices connected to a network within an area 11a.

[0164] On the other hand, referring to FIG. 1c, electronic devices (reference numerals 1c and 1d) may be electronic devices that are not connected to a network within the area 11a.

[0165] On the other hand, referring to FIG. 1c, the electronic devices (reference numerals 2a-2b) may be electronic devices outside the range of the region 11a.

[0166] On the other hand, referring to FIG. 1c, there may be networks that interact with electronic devices within the area 11a, and there may be networks that interact with electronic devices outside the area 11a.

[0167] Here, the network interacting with electronic devices within the area 11a can serve to transmit and receive information for controlling smart home appliances.

[0168] Also, the network that interacts with the electronic devices within the area 11a may be, for example, a short-range network or a local network, whereas the network that interacts with the electronic devices within the area 11a may be, for example, a long-range network or a global network.

[0169] The detailed description of the operation of the network shown in FIG. 1c is the same as that described above, so the duplicated description will be omitted.

[0170] On the other hand, referring to FIG. 1c, there may be one or more electronic devices connected via a network outside the scope of area 11a, and in this case, the electronic devices may process data in a distributed manner or perform one or more operations separately.

[0171] Alternatively, if there are one or more electronic devices connected via a network outside the area 11a, the electronic devices may perform various operations independently of each other.

[0172] Meanwhile, according to an embodiment of the present invention, as shown in Fig. 3, a computing device 100 may include a network unit 110, a memory 120, and a processor 130. The components included in the computing device 100 described above are merely exemplary, and the scope of the present invention is not limited to the components described above. That is, additional components may be included or some of the components described above may be omitted depending on the implementation of the embodiment of the present invention.

[0173] According to an embodiment of the present invention, the computing device 100 may include a network unit 110 that transmits and receives data to and from the user terminal 10 and the external server 20. The network unit 110 may transmit and receive data, etc., for performing the method for analyzing a sleep state based on sleep sound information according to an embodiment of the present invention, to and from other computing devices, servers, etc. That is, the network unit 110 may provide a communication function between the computing device 100, the user terminal 10, and the external server 20. For example, the network unit 110 may receive sleep sound information from the user terminal 10 and transmit sleep state information corresponding to the received sleep sound information to the user terminal 10. Furthermore, for example, the network unit 110 may receive sleep examination records and electronic health records for multiple users from a hospital server. In addition, the network unit 110 may allow information transmission between the computing device 100, the user terminal 10, and the external server 20 by calling a procedure in the computing device 100.

[0174] The network unit 110 according to an embodiment of the present invention can use various wired / wireless communication systems such as the networks described above. The network has been described above, so a duplicate description will be omitted.

[0175] According to an embodiment of the present invention, the memory 120 may store computer programs for performing a method for generating a sleep analysis model for predicting a sleep state based on acoustic information and a method for analyzing a sleep state through sleep acoustic information, according to an embodiment of the present invention, and the stored computer programs may be read and driven by the processor 130. The memory 120 may also store any type of information generated or determined by the processor 130 and any type of information received by the network unit 110. The memory 120 may also store data related to the user's sleep. For example, the memory 120 may temporarily or permanently store input / output data (e.g., sleep acoustic information related to the user's sleep environment, sleep state information corresponding to the sleep acoustic information, etc.).

[0176] According to an embodiment of the present invention, the memory 120 may include at least one type of storage medium selected from the group consisting of flash memory, hard disk, micro multimedia card, card-type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. The computing device 100 may also operate in association with web storage that performs the storage function of the memory 120 over the Internet. The above description of memory is for illustrative purposes only, and the present invention is not limited thereto.

[0177] According to one embodiment of the present invention, processor 130 may be configured with one or more cores and may include a central processing unit (CPU) of a computing device, a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), or other processor for data analysis, machine learning, or deep learning.

[0178] According to an embodiment of the present invention, the processor 130 may read a computer program stored in the memory 120 and perform data processing for model training. According to an embodiment of the present invention, the processor 130 may perform calculations for neural network training. The processor 130 may perform calculations for neural network training, such as processing input data for training using machine learning or deep learning, extracting features from the input data, calculating errors, and updating weights of the neural network using backpropagation.

[0179] In addition, at least one of the CPU, GPGPU, and TPU of the processor 130 may process network function training. For example, the CPU and GPGPU may both process network function training and data classification using the network function. In addition, in an embodiment of the present invention, processors of multiple computing devices may be used together to process network function training and data classification using the network function. In addition, a computer program executed in a computing device according to an embodiment of the present invention may be a CPU-, GPGPU-, or TPU-executable program.

[0180] As used herein, a network function may be used interchangeably with an artificial neural network or a neural network. As used herein, a network function may include one or more neural networks, where the output of the network function may be an ensemble of outputs of the one or more neural networks. Also, as used herein, a model may include a network function. A model may include one or more network functions, where the output of the model may be an ensemble of outputs of the one or more network functions.

[0181] The processor 130 may read a computer program stored in the memory 120 to execute a sleep analysis model according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 130 may perform calculations to calculate sleep analysis information based on sleep sensing data. Alternatively, according to an embodiment of the present invention, the processor 130 may perform calculations to train the sleep analysis model.

[0182] According to an embodiment of the present invention, the processor 130 may generally process the overall operation of the computing device 100. The processor 130 may process signals, data, information, etc. input or output via the components described in detail above, or may run applications stored in the memory 120, thereby providing or processing appropriate information or functions to the user terminal.

[0183] According to one embodiment of the present invention, the processor 130 may acquire a plurality of training data sets to perform training on the neural network (or one or more network functions). The plurality of training data sets may be associated with a plurality of pieces of sleep acoustic information associated with each of a plurality of users. The processor 130 may acquire a plurality of pieces of sleep acoustic information associated with the sleep of a plurality of users, and may perform training on one or more network functions using a training data set including the plurality of pieces of sleep acoustic information to generate a sleep analysis model.

[0184] According to an embodiment of the present invention, acquiring a plurality of pieces of sleep sound information may involve acquiring or loading sleep sound information stored in memory 120. In one embodiment, a plurality of pieces of sleep sound information may be received from external server 20 via network unit 110, and the received sleep sound information may be stored in memory 120. Alternatively, acquiring the sleep sound information may involve receiving or loading data from another storage medium, another computing device, or a separate processing module within the same computing device via wired or wireless communication means.

[0185] Sleep status information

[0186] Meanwhile, in the present invention, the sleep state information may be information related to whether or not the user is asleep. Specifically, the sleep state information may include at least one of first sleep state information indicating that the user is before sleeping, second sleep state information indicating that the user is asleep, and third sleep state information indicating that the user has fallen asleep. In other words, when first sleep state information is inferred in relation to a user, processor 130 can determine that the user is in a pre-sleep state (i.e., before falling asleep), when second sleep state information is inferred, the processor can determine that the user is in a sleeping state, and when third sleep state information is acquired, the processor can determine that the user is in a post-sleep state (i.e., wake-up state).

[0187] Such sleep state information may be acquired based on environmental sensing information or actigraphy. Environmental sensing information may be sensing information acquired in a space where the user is located in a non-contact manner. For example, the processor 130 may extract sleep state information based on acquired environmental sensing information (such as acoustic information related to cleaning, cooking, watching TV, and sleep acoustic information acquired during sleep), actigraphy, biometric information, etc. In this case, sleep acoustic information acquired during the user's sleep may include sounds generated by the user turning over in bed during sleep, sounds related to muscle movements, breathing sounds during sleep, etc. In other words, sleep acoustic information in the present invention may refer to acoustic information related to the user's movement patterns and breathing patterns during sleep.

[0188] In addition, the sleep state information according to an embodiment of the present invention may include, in addition to sleep stage information, various information related to sleep events, such as information related to breathing during sleep, information on teeth grinding, whether or not a person coughs, the severity of the cough, whether or not a person sneezes, information on turning over in sleep, and information on talking in their sleep.

[0189] Sleep stage information

[0190] According to an embodiment, the processor 130 may extract sleep stage information. The sleep stage information may be extracted based on environmental sensing information of the user. Sleep stages may be classified into non-REM (NREM) sleep and rapid eye movement (REM) sleep, and NREM sleep may be further classified into multiple stages (e.g., two stages, Light and Deep, and four stages, N1 to N4). Sleep stages may be defined as general sleep stages or may be arbitrarily set by a designer to include various sleep stages. Through sleep stage analysis, not only the quality of sleep related to sleep but also various sleep events such as sleep disorders (e.g., sleep apnea) and their underlying causes (e.g., snoring) may be predicted.

[0191] Sleep environment information

[0192] In an embodiment, the sleep environment information of the present invention may be acquired through the user terminal 10. The sleep environment information may refer to information related to sleep acquired in a space where a user is located. The sleep environment information may be sensing information acquired in a space where a user is located using a non-contact method. The sleep environment information may be information related to a user's sleep acquired from a smart watch, smart home appliance, etc.

[0193] For example, the sleep environment information may be acoustic information acquired in the bedroom where the user sleeps. According to an embodiment, the sleep environment information acquired via the user terminal 10 may be information that serves as a basis for acquiring the user's sleep state information in the present invention. As a specific example, sleep state information related to whether the user is before sleep, during sleep, or after sleep can be acquired through the sleep environment information acquired in relation to the user's activity. As another specific example, the sleep environment information may include various information such as the user's heart rate, the user's respiration, illuminance and noise information related to the user's sleep environment.

[0194] In addition, the sleep environment information may be at least one of noise information that commonly occurs in daily life (acoustic information related to cleaning, acoustic information related to cooking food, acoustic information related to watching TV, cats' voices, dogs' voices, birds' voices, car sounds, wind sounds, rain sounds, etc.), or other biological information (electrocardiogram, brain waves, pulse information, information on muscle movements, etc.).

[0195] Data for sleep analysis

[0196] FIG. 13 is a schematic diagram of a data set according to one embodiment of the present invention.

[0197] The data according to an embodiment of the present invention may be raw acoustic information collected via a microphone, where the raw acoustic information may be information in the time domain having amplitude, phase, and frequency.

[0198] Furthermore, data according to an embodiment of the present invention may be obtained by converting low acoustic information into information including changes in frequency components of the low acoustic information along the time axis.

[0199] Alternatively, data according to an embodiment of the present invention may be raw acoustic information converted into information in the frequency domain rather than information in the time domain, where Fourier transform or wavelet transform may be performed to convert into information in the frequency domain.

[0200] Additionally, according to an embodiment of the present invention, the information transformed into the frequency domain may be information having an amplitude and a frequency.

[0201] Furthermore, data according to an embodiment of the present invention may correspond to a spectrogram, which is obtained by converting acoustic information into information in the frequency domain.

[0202] Alternatively, according to an embodiment of the present invention, data may be a mel spectrogram obtained by applying a mel scale to a spectrogram. Specifically, a mel spectrogram may be obtained by applying a mel filter bank to the spectrogram. Generally, the vibrating parts of the human cochlea may differ depending on the frequency of the audio data. Furthermore, the human cochlea has a characteristic of being sensitive to frequency changes in low frequency bands and insensitive to frequency changes in high frequency bands. Therefore, a mel spectrogram may be obtained from the spectrogram using a mel filter bank to have a recognition capability similar to the characteristics of the human cochlea for audio data. That is, the mel filter bank may apply fewer filter banks to low frequency bands and wider filter banks to higher frequency bands. In other words, the processor 130 may apply a mel filter bank to the spectrogram to obtain a mel spectrogram in order to recognize audio data in a manner similar to the characteristics of the human cochlea. The mel spectrogram may include frequency components that reflect the characteristics of human hearing. That is, in the present invention, the spectrogram generated in response to sleep acoustic information and subjected to analysis using a neural network may include the above-mentioned mel spectrogram.

[0203] Here, the information converted into information in the frequency domain, such as a spectrogram or mel spectrogram according to one embodiment of the present invention, may be converted into information that includes changes in the frequency components of the acoustic information over time, as a domain having amplitude and frequency.

[0204] Furthermore, data according to an embodiment of the present invention may be input to an image processing-based artificial intelligence model as a visualization of the above-described information. For example, low acoustic information may be converted into information including changes in frequency components over time and visualized to be input to the artificial intelligence model. Alternatively, information converted into the frequency domain may be visualized to be input to the artificial intelligence model. The artificial intelligence model to which the above information is input may be an image processing-based artificial intelligence model.

[0205] Meanwhile, according to an embodiment of the present invention, among the sleep state information, sleep acoustic information may be collected through polysomnography (PSG) in a hospital environment, or may be collected by a user in a home environment, etc., through a microphone built into a user terminal such as a wearable device or a smartphone.

[0206] A dataset according to an embodiment of the present invention may be collected and constructed via a microphone in a polysomnography (PSG) or electroencephalography (Video-EEG).

[0207] Alternatively, the dataset according to the embodiment of the present invention may be constructed by collecting acoustic signals generated during sleep via a microphone built into an electronic device such as a user terminal.

[0208] Sleep analysis results derivation method

[0209] The following describes how the processor 130 derives the final sleep analysis result using bio-signal, movement information (ACTIGRAPHY), and sleep acoustic information (SOUND).

[0210] First, the processor 130 may derive a final sleep analysis result using weighting values. Specifically, the processor 130 may derive a secondary sleep analysis result by applying the same weighting value to the primary sleep analysis result and the sleep analysis result using sleep acoustic information. Alternatively, the processor 130 may derive a secondary sleep analysis result by applying different weighting values to the primary sleep analysis result and the sleep analysis result using sleep acoustic information. For example, the processor 130 may derive a final secondary sleep analysis result by applying a weighting value of 30% to the contact-type primary sleep analysis based on HRV and actigraphy and a weighting value of 70% to the AI analysis using sleep acoustics.

[0211] In another embodiment, the processor 130 can determine that the user has entered a particular sleep stage and derive a final sleep analysis result only if the sleep stages in the primary sleep analysis result and the secondary sleep analysis result completely match.

[0212] In yet another embodiment, processor 130 may use a method for training an AI sleep analysis model that receives input of at least one of bio-signal, movement, and sound information. The AI sleep analysis model training method will be described in more detail below, but briefly, an AI sleep analysis model that performs sleep analysis based on one or more factors may be generated by inputting one or more pieces of information into the input layer of the artificial intelligence model.

[0213] In another embodiment, the processor 130 first performs a secondary sleep analysis using sleep acoustic information (SOUND) using an AI sleep analysis model (described below), and then additionally extracts an AI confidence factor for the sleep stage for each time period. If the extracted confidence factor is equal to or less than a predetermined value, the sleep stage result for that time period is adopted as the sleep stage result derived by the primary sleep analysis. In other words, by additionally adopting the primary sleep analysis result in addition to the secondary sleep analysis result, more reliable sleep analysis results can be derived.

[0214] In yet another embodiment, processor 130 first obtains statistics for portions of the AI sleep analysis model (described below) that are inconsistent with the actual analysis results. The statistics may be input by the user or may be obtained automatically based on multiple user data. Processor 130 may additionally use the primary sleep analysis results for portions of the obtained statistics that are inconsistent with the actual analysis results, centering on the secondary sleep analysis results (SOUND-based analysis).

[0215] In yet another embodiment, the processor 130 may use a method for training an AI sleep analysis model based on the primary sleep analysis result obtained from bio-signal and movement information (ACTIGRAPHY) and sleep acoustic information (SOUND). The method for training an AI sleep analysis model will be described in more detail below, but briefly, an AI sleep analysis model that performs sleep analysis based on the two factors may be generated by inputting two pieces of information (primary sleep analysis result and sleep acoustic information) into the input layer of the artificial intelligence model.

[0216] Sleep stages may be classified into NREM (non-REM) sleep and REM (rapid eye movement) sleep, and NREM sleep may be further classified into multiple stages (e.g., two stages: Light and Deep, and four stages: N1 to N4). Sleep stages may be defined based on commonly used sleep stages, or may be arbitrarily set in various ways by designers. Analysis of sleep stages can predict not only sleep quality but also sleep disorders (e.g., sleep apnea) and their underlying causes (e.g., snoring).

[0217] A method for learning and predicting a plurality of pieces of sleep state information according to an embodiment of the present invention will now be described using examples. However, the specific descriptions related to sleep states described below are merely examples, and the present invention is not limited thereto. It should be understood that learning can also be performed for at least one of other sleep state information not mentioned (e.g., sleep event information such as movement information, snoring information, and sleep disorder information) and / or differences in sleep state information due to differences in environment (e.g., race, etc.).

[0218] Learning or predicting sleep stage information may require acoustic information acquired over a long time interval, while learning or predicting sleep state information other than sleep stage information (e.g., sleep event information such as snoring or apnea information) may require acoustic information acquired over a relatively short time interval (e.g., 1 minute) around the time the sleep state occurs.

[0219] The present invention may include a method for converting acquired acoustic information into information including changes in frequency components over time, information in the frequency domain, or a spectrogram, and learning multiple pieces of sleep state information when the converted information is used as an input to a single artificial intelligence model. Alternatively, the present invention may include a method for learning multiple pieces of sleep state information when information including changes in frequency components over time of acquired acoustic information is visualized and displayed as an input to an image processing-based artificial intelligence model.

[0220] For example, the spectrogram information according to an embodiment of the present invention may be input to the same feature extraction model, and output information may be input to different feature classification models to perform learning.

[0221] Through this method, learning may be performed so that various sleep state information can be predicted based on one acoustic information.

[0222] Alternatively, through this method, various sleep state information can be complementarily learned based on one acoustic information. For example, an AI model that learns only sleep stages may mistakenly predict a state in which a loud noise, such as apnea or snoring, occurs as a wake state if such a state is recognized. On the other hand, an AI model designed to learn multiple sleep state information can complementarily learn other sleep state information, such as apnea or snoring, in addition to sleep stages, thereby preventing the above problem.

[0223] The specific description of the time interval and the specific description related to the sleep state information are merely examples for explaining the present invention, and the present invention is not limited thereto.

[0224] Acquisition of sleep acoustic information

[0225] FIG. 2 is an exemplary diagram illustrating a plurality of pieces of sleep acoustic information acquired in various environments.

[0226] The sleep sound information may include, as information related to sleep sounds, sounds generated by the user turning over in their sleep, sounds related to muscle movements, sounds related to the user's breathing during sleep, etc. In other words, the sleep sound information of the present invention may include sound information related to the movement patterns and breathing patterns associated with the user's sleep.

[0227] In the present invention, since sleep sound information is related to sounds associated with breathing and body movement, it may be very quiet. Therefore, the processor 130 can convert the sleep sound information into information or a spectrogram containing changes in frequency components of the raw sleep sound information over time, and then analyze the sound. In this case, as described above, the converted information includes information indicating how the frequency spectrum of the sound changes over time, making it possible to easily identify breathing or movement patterns associated with relatively quiet sounds (i.e., sounds whose characteristics are difficult to identify), thereby improving the efficiency of the analysis.

[0228] According to an embodiment of the present invention, as shown in Fig. 8, each spectrogram may be configured to have a frequency spectrum with different intensities according to various sleep stages. That is, it may be difficult to predict whether a sleep state is at least one of awake, REM sleep, light sleep, and deep sleep based solely on changes in the energy level of sleep sound information. However, by converting the sleep sound information into a spectrogram, changes in the spectrum of each frequency can be easily detected, making it possible to analyze quiet sounds (e.g., breathing and body movement). That is, by imaging sound and analyzing the image pattern, it may be possible to analyze low-quality sound.

[0229] According to an embodiment of the present invention, sleep acoustic information, which is information that serves as the basis for sleep state analysis, may include various types of noise. In this embodiment, the sleep acoustic information acquired for each of multiple users may be acquired in a different bedroom environment for each user, and may include different types of noise. Due to the influence of such various types of noise, it may be difficult to consistently analyze or predict the sleep acoustic information using a sleep analysis model.

[0230] Specifically, even if a user's sleep sounds are the same, the sleep sound information actually acquired may differ due to various noises generated in the bedroom environment or by the sound measurement device, and as a result, the neural network model (i.e., sleep analysis model) may output different prediction information (i.e., sleep state information). For example, even if the sleep sounds generated during a user's sleep are all the same, different sleep sound information may be acquired as shown in Figure 2 due to differences between environmental factors related to the space in which the user sleeps and the device that acquires the sleep sound. Figure 2 exemplarily illustrates that different types of sleep sound information may be acquired because the same sound information is acquired through various sleeping environments and / or various measurement devices.

[0231] More specifically, the background noise of the acquired sleep audio information may vary depending on the size and structure of the bedroom where the user sleeps. Also, the acquired sleep audio information may vary depending on various noises generated in the space where the user sleeps (e.g., the sound of an air conditioner, a fan, a pet, or a refrigerator).

[0232] For another example, the sleep sound information acquired in response to the same sleep sound may differ depending on the model of the sound measurement device used to acquire the sleep sound information. For a specific example, when first sleep sound information is acquired through a first user terminal and second sleep sound information is acquired through a second user terminal in response to the same sleep sound, the first sleep sound information and the second sleep sound information may not be completely identical due to differences between the microphone modules provided in the two devices. The sleep sound information may contain various types of noise due to the different microphone modules used by the devices.

[0233] That is, the sleep acoustic information that a user personally acquires and analyzes contains various noises because it is acquired through different bedroom environments and different acoustic measurement devices for each user, making it difficult to perform a consistent analysis using a sleep analysis model.

[0234] For example, a sleep analysis model may be generated by training a neural network using data acquired in various noise-free environments as training data. As a specific example, a user may sleep in a space (e.g., a hospital) where there is no noise or only predetermined noise and where an acoustic measurement device with predetermined performance is installed, and sleep acoustic information during sleep may be acquired. A specialized medical professional (e.g., a sleep technician) may label the sleep acoustic information acquired in this manner with a correct answer (i.e., sleep stage) corresponding to each point in the time series of sleep acoustics, and a neural network may be trained using the sleep acoustics and the labeled data to generate a sleep analysis model.

[0235] However, the sleep analysis model generated in the above manner is trained using sleep acoustic information acquired through a high-performance microphone module in a sleep environment with little noise or only predefined noise, and learning data in which a ground truth corresponding to the sleep acoustic information is labeled, so that it is difficult to expect robust performance for acoustic data acquired in various noise environments. Furthermore, since it is difficult to determine whether the sleep analysis model can be robust in various environments, it may not be easy to generalize the sleep analysis model.

[0236] In particular, since the present invention is intended to easily provide sleep state information in the user's real life, it must be possible to accurately analyze sleep acoustic information that contains various noises depending on each individual user's bedroom environment and acoustic measurement equipment.

[0237] According to an embodiment of the present invention, an electronic device according to an embodiment of the present invention may generate a sleep analysis model having robust performance even for sleep audio information including various background noises through adaptive learning of a neural network based on sleep audio information related to different domains. Sleep audio information related to different domains may refer to sleep audio information acquired in different environments and by different methods.

[0238] A specific method for generating a sleep analysis model having robust performance even for sleep acoustic information containing various background noises through adaptive learning of a neural network based on sleep acoustic information related to different domains, and a specific method for providing sleep state information through the sleep analysis model, according to an embodiment of the present invention, will be described in detail below. Also, according to an embodiment of the present invention, it is possible to perform pre-processing for removing or reducing noise, or to obtain sleep state information based on acoustic information to which noise has been added, and these methods will also be described in detail below.

[0239] Meanwhile, according to an embodiment of the present invention, the processor 130 may acquire sleep state information based on acoustic information, actigraphy, and biological information acquired from the user terminal 10. Specifically, the processor 130 may identify a singular point where pre-established pattern information is detected in the acoustic information. Here, the pre-established pattern information may be related to breathing and movement patterns associated with sleep.

[0240] For example, in a wakeful state, the entire nervous system is activated, resulting in irregular breathing patterns and frequent body movements. Furthermore, the neck muscles may not be relaxed, resulting in very little breathing noise. On the other hand, when a user sleeps, the autonomic nervous system stabilizes, resulting in regular breathing, less body movements, and louder breathing noise. That is, the processor 130 may identify, as singular points in the acoustic information, points at which acoustic information of a predetermined pattern associated with regular breathing, little body movements, or little breathing noise is detected. The processor 130 may also acquire sleep acoustic information based on the identified singular points. The processor 130 may identify singular points associated with the user's sleep time points in the acoustic information acquired in a time series, and acquire sleep acoustic information based on the singular points.

[0241] 4 is a diagram illustrating a process of acquiring sleep acoustic information in a sleep analysis method according to the present invention. Referring to FIG. 4, processor 130 may identify a singular point P associated with a user's sleep from acoustic information E. Based on the identified singular point P, processor 130 may acquire sleep acoustic information SS based on acoustic information acquired after the singular point P. The waveforms and singular points associated with acoustics in FIG. 4 are merely examples for understanding the present invention, and the present invention is not limited thereto.

[0242] That is, the processor 130 can identify a singular point P related to the user's sleep from the acoustic information, and extract and acquire only the sleep acoustic information SS from a huge amount of environmental sensing information (i.e., acoustic information) based on the singular point P. This can automate the process of the user recording their sleep time, providing convenience and contributing to improving the accuracy of the acquired sleep acoustic information.

[0243] Furthermore, in an embodiment, processor 130 may acquire sleep state information related to whether the user is asleep or not, based on singular point P identified from acoustic information E. Specifically, processor 130 may determine that the user is asleep if singular point P is not identified, and may determine that the user is asleep after singular point P if singular point P is identified. Furthermore, processor 130 may identify a time point (e.g., a wake-up time) at which a previously set pattern is not observed after singular point P is identified, and may determine that the user has fallen asleep, i.e., woken up, if the time point is identified.

[0244] That is, the processor 130 can obtain sleep state information related to whether the user is before sleep, during sleep, or after sleep based on whether a singular point P is identified in the acoustic information E and whether sleep is continuously sensed after the singular point is identified.

[0245] Meanwhile, the processor 130 may acquire sleep state information based on actigraphy or biometric information rather than acoustic information E. It may be advantageous to acquire user movement information through a sensor unit in contact with the body. In the present invention, the user's sleep state information is obtained in advance using actigraphy or biometric information during primary sleep analysis, thereby further improving the reliability of the sleep state analysis.

[0246] Meanwhile, the technical idea of acquiring sleep state information based on the above-mentioned sleep-related pattern information or singularities is merely an example, and the present invention is not limited to performing inference based on already-set pattern information or singularities, but may also include performing inference via an artificial intelligence model generated to acquire sleep state information.

[0247] Sleep acoustic information acquired in various environments

[0248] According to an embodiment of the present invention, the plurality of pieces of sleep sound information may include pieces of sleep sound information related to different domains. Specifically, the plurality of pieces of sleep sound information may include a plurality of source data and a plurality of target data. The plurality of source data and the plurality of target data may be characterized as being acquired in different sleeping environments in the information related to sleep sounds related to different domains.

[0249] For example, the plurality of source data may be acoustic data related to a first domain acquired in a professional sleep measurement environment (e.g., a polysomnography test), and the plurality of target data may be acoustic data related to a second domain acquired in an individual user's daily sleep environment. For example, the plurality of target data may be a large amount of data (i.e., sleep acoustic information) acquired from multiple users when the computing device 100 provides a sleep analysis service. As another example, the plurality of target data may be sleep acoustic information acquired via a microphone module of the user terminal 10.

[0250] In an embodiment of the present invention, the plurality of source data may be sleep acoustic information obtained in a space (e.g., a hospital) where there is no noise or only predetermined noise and where acoustic measurement equipment with predetermined performance is installed. Also, in an embodiment of the present invention, the plurality of source data may be data labeled with information about a plurality of sleep states by a professional medical professional (e.g., a sleep technician) and may be data containing predetermined noise.

[0251] In an embodiment of the present invention, the plurality of target data may include various types of noise depending on each individual's bedroom environment, or may be sleep acoustic information acquired through a plurality of user terminals each having a different microphone module. In an embodiment of the present invention, the plurality of target data may be data in which information regarding a plurality of sleep states is not labeled and includes undefined noise.

[0252] That is, the plurality of source data may be data related to a plurality of sleep acoustic data acquired through equipment set up in a noise-minimized state at a specialized institution (e.g., a hospital), and the plurality of target data may include data related to sleep acoustic data acquired individually from each individual user. The plurality of target data may be acquired in different ways depending on the bedroom environment of each user, and thus may contain various noises and may be related to acoustic data for which correct answers regarding sleep states (e.g., sleep stages) are not labeled.

[0253] In the case of multiple target data, various noises may be included depending on the bedroom environment of each user. For example, even for the same sleep sound, the sleep sound information obtained may differ depending on the size and structure of each user's bedroom and the distance between the user and the sound measuring device. That is, sleep sound information (i.e., multiple target data) including different background noises may be obtained depending on the size and shape of the sleeping space and the position of the sound measuring device when the user is sleeping.

[0254] For example, the plurality of target data may include various noises related to sounds generated in a space where a user sleeps. For example, the space where a user sleeps may include various noises such as the operating sounds of electronic appliances such as air conditioners and fans, and the sounds of pets.

[0255] As another example, the plurality of target data may contain various noises depending on the model of the acoustic measurement device used to acquire the sleep acoustic information. Specifically, when first sleep acoustic information is acquired through a first user terminal and second sleep acoustic information is acquired through a second user terminal corresponding to the same sleep acoustic, the first sleep acoustic information and the second sleep acoustic information may not be completely identical due to differences in the specifications of the microphone modules provided in the two devices. The sleep acoustic information may contain various types of noise due to differences in the microphone modules used by the devices.

[0256] That is, as described above, the plurality of target data is acoustic data acquired in the individual sleeping environments of the plurality of users, and may include a greater variety of noises.

[0257] According to an embodiment, each of the plurality of target data may be acquired through a user terminal 10 carried by a user. For example, sleep acoustic information related to the user's sleep environment may be acquired through a microphone module provided in the user terminal 10.

[0258] Generally, a microphone module installed in a user terminal 10 carried by a user may be configured with a micro-electromechanical system (MEMS) since it must be installed in a relatively small user terminal 10. Such a microphone module can be manufactured in a very small size, but may have a lower signal-to-noise ratio (SNR) than a condenser microphone or a dynamic microphone. A low SNR means that the ratio of noise, which is sound that makes it difficult to identify the sound to be identified, is high, and the sound is difficult to identify (i.e., unclear). Therefore, it is necessary to remove or mitigate noise, which will be described in detail below.

[0259] Noise reduction pre-processing

[0260] FIG. 14 is a diagram illustrating noise reduction according to an embodiment of the present invention. As shown in FIG. 14 or FIG. 5, audio information extracted from a user or sleep audio information (raw data) extracted therefrom undergoes a noise reduction preprocessing process. In the noise reduction process, noise (e.g., white noise) contained in the raw data is removed. The noise reduction process may be performed using algorithms such as spectral gating and spectral subtraction to remove background noise. Furthermore, in the present invention, the noise reduction process may be performed using a deep learning-based noise reduction algorithm. The deep learning-based noise reduction algorithm may be specialized for the user's breath and breathing sounds, in other words, a noise reduction algorithm learned through the user's breath and breathing sounds.

[0261] The pre-processing may be performed during a learning process of the sleep state information or during an inference process. An example of a pre-processing process for noise reduction will now be described.

[0262] Spectral Noise Gating

[0263] Spectral gating or spectral noise gating is a preprocessing method for audio information. Noise reduction can be performed on the entire acquired audio information, or it can be performed by splitting the audio information at regular time intervals (e.g., 5 minutes) and then performing noise reduction on each of the split audio information. To perform noise reduction on audio information split at regular time intervals, a method of first calculating a spectrum for each frame may be included.

[0264] Therefore, it is possible to identify the frame having the frequency spectrum with the smallest energy among the calculated spectrum frames.

[0265] The method may include a method of assuming that the frame having the smallest energy frequency spectrum among the respective spectrum frames is static noise, and subtracting the frequency of the frequency spectrum frame assumed to be static noise from the spectrum frame.

[0266] Meanwhile, according to an embodiment of the present invention, when performing noise reduction preprocessing on a plurality of data, sleep audio information may be classified into one or more audio frames. Here, a minimum audio frame having a minimum energy level may be identified based on the energy level of each of the one or more audio frames. Thus, noise removal or reduction may be performed on the audio data based on the minimum audio frame. For example, the processor 130 may classify 30 seconds of sleep audio information (e.g., target data) into one or more audio frames each having a very short amplitude of 40 ms. The processor 130 may also compare the amplitudes of each of the plurality of audio frames associated with the amplitude of 40 ms to identify the minimum audio frame having the minimum energy level.

[0267] The processor 130 may remove the identified minimum acoustic frame component from the entire sleep acoustic information (i.e., 30 seconds of sleep acoustic information). For example, referring to FIG. 14, the minimum acoustic frame component may be removed from the sleep acoustic information to obtain pre-processed sleep acoustic information. That is, the processor 130 may identify the minimum acoustic frame as a background noise frame and perform noise removal or reduction from the original signal (i.e., sleep acoustic information). The specific numerical values for the time intervals described above are merely examples and are not limiting.

[0268] Deep learning-based noise reduction

[0269] Meanwhile, to perform preprocessing for noise reduction according to an embodiment of the present invention, a deep learning-based noise reduction method performed on raw sound information in the time domain rather than the frequency domain may be used. For deep learning-based noise reduction, a method may be used in which information such as sleep sound information, which is information necessary for inputting a sleep analysis model, is maintained and other sounds are attenuated.

[0270] Noise reduction can be performed not only on acoustic information acquired through PSG test results but also on acoustic information acquired through a microphone built into a user terminal such as a smartphone.

[0271] Convert raw acoustic information into frequency domain information

[0272] FIG. 5 is a conceptual diagram illustrating a privacy protection method using mel-spectrogram transformation for sleep acoustic information extracted from a user in the sleep analysis method according to the present invention.

[0273] The sleep analysis method according to the present invention generates an inference model through deep learning of acoustic information, and the inference model extracts a user's sleep state and sleep stage. Briefly again, environmental sensing information (acoustic information) including sleep acoustic information may be converted into information including changes in frequency components of the acoustic information over time or into information in the frequency domain, and the inference model may be generated based on the converted information.

[0274] Alternatively, according to an embodiment of the present invention, environmental sensing information including sleep sound information is converted into frequency domain information or a spectrogram, and an inference model is generated based on the converted frequency domain information or the spectrogram. Here, the frequency domain information may be information including changes in frequency components of raw sleep sound information over time.

[0275] In sleep analysis using acoustic information, protecting the privacy of the user cannot be overlooked, and the present invention employs a process of pre-processing the acoustic information to protect the privacy of the user.

[0276] In this case, a method of converting raw acoustic information into frequency domain information or a spectrogram based only on the amplitude excluding the phase can be used, and this method not only protects privacy but also reduces data volume and improves processing speed. However, in other embodiments, a spectrogram can be generated using both the phase and the amplitude.

[0277] In one embodiment of the present invention, a sleep analysis model can be generated using a spectrogram SP converted based on sleep audio information SS. If sleep audio information expressed as audio data were used as is, the amount of information would be very large, resulting in a significant increase in the amount and time of calculations. In addition, the inclusion of unwanted signals would reduce calculation accuracy. Furthermore, if all of the user's audio signals were transmitted to a server, there would be a risk of privacy violations.

[0278] Therefore, in one embodiment of the present invention, noise is removed from sleep acoustic information using the above-described method, and then the information is converted into frequency domain information or a spectrogram, and the spectrogram is trained to generate a sleep analysis model, thereby reducing the amount of calculation and calculation time, and even protecting personal privacy.

[0279] For example, in acoustic information acquired through a microphone, sleep acoustic information (e.g., a user's breath, etc.) required for sleep stage analysis may be relatively quieter than other noises, but if converted into a spectrogram, the sleep acoustic information may be better identified than other surrounding noises.

[0280] On the other hand, when converting to a spectrogram according to an embodiment of the present invention, personal information cannot be identified by converting to a lower frequency domain resolution, but if the frequency resolution (frequency bins) is configured to a certain number (e.g., 20) or less, personal information cannot be identified from the restored signal.

[0281] Embodiments of the present invention may also include a method for converting acquired acoustic information into a spectrogram in real time.

[0282] In addition, by being able to compress the frequency resolution of the spectrogram on the user's smartphone rather than on a server or cloud, it is possible to prevent the leakage of personal information.

[0283] In this case, the de-identification of sound data may be performed on natural language and respiratory sounds, which may be converted into a natural language conversion spectrogram, a respiratory sound conversion spectrogram, etc. In the sleep analysis according to the present invention, only information required for the analysis model is used to improve the calculation speed and reduce the calculation load. Meanwhile, the spectrogram according to the embodiment of the present invention may be a mel spectrogram to which a mel scale is applied.

[0284] How to convert raw sleep acoustic information

[0285] FIG. 6 is a diagram illustrating a method for acquiring a spectrogram corresponding to sleep acoustic information in the sleep analysis method according to the present invention.

[0286] The processor 130 can generate a spectrogram SP corresponding to the sleep acoustic information SS, as shown in Fig. 6. Raw data (raw acoustic information in the time domain) that is the basis for generating the spectrogram SP can be input, and the raw data according to the present invention may be collected through polysomnography (PSG) in a hospital environment, or may be collected by a user in a home environment or the like through a microphone built into a user terminal such as a wearable device or a smartphone.

[0287] In addition, the raw data may be acquired from the start time to the end time entered by the user via a user terminal 10 such as a wearable device or a smartphone, or from the time when the user operates the device (e.g., setting an alarm) to the time corresponding to the device operation (e.g., the alarm setting time), or may be acquired by automatically selecting a time based on the user's sleep pattern, or may be acquired by automatically determining the user's sleep intention time based on sounds (such as the user's voice, breathing sounds, sounds from peripheral devices (TV, washing machine), etc.) or changes in illumination.

[0288] The processor 130 may perform a fast Fourier transform on the sleep audio information SS to convert it into information including changes in frequency components of the sleep audio information over time. Specifically, this information may be information in the frequency domain, such as a spectrogram or a mel spectrogram to which a mel scale is applied. The information including changes in frequency components over time, the frequency domain information, or the spectrogram SP is used to visualize and understand sounds or waves, and may be a combination of waveform and spectrum features.

[0289] In addition, this information may be visualized by representing the difference in amplitude according to the change in time axis and frequency axis as a difference in print density or display color, based on the information on the frequency components of the acoustic information on the time axis. When visualized in this way, sleep state information can be obtained as an input for an image processing-based artificial intelligence model. By converting the acoustic signal into an image signal in this way, time-series sleep analysis can be performed using data from a relatively long period of time, which has the advantage of enabling more accurate sleep analysis than analysis based on raw sleep acoustic information.

[0290] The preprocessed acoustic raw data may be divided into 30-second units and converted into spectrograms. As a result, a 30-second spectrogram has dimensions of 20 frequency bins x 1201 time steps. In the present invention, in order to convert a rectangular spectrogram into a form closer to a square, various methods such as reshaping, resizing, and split-cat may be used. Using these methods may also increase the amount of information that can be stored.

[0291] Meanwhile, the present invention can use a method of simulating breath measurements in various home environments by adding various noises generated in the home environment to clean breath sounds. Because sounds have an additive nature, they can be added to each other. However, adding original audio signals such as MP3 or PCM and converting them into spectrograms consumes a significant amount of computing resources. Therefore, the present invention proposes a method of converting breath and noise into spectrograms and adding them. By simulating breath measurements in various home environments and using them for deep learning model training, the robustness of sleep analysis in various home environments can be ensured.

[0292] Preprocessing of the transformed information

[0293] The purpose of converting data according to one embodiment of the present invention into information including changes in frequency components over time, or information in the frequency domain, or a spectrogram is to infer, via a trained model, what sleep state or sleep stage a pattern in the converted information corresponds to as input to a sleep analysis model; however, some preprocessing steps may be required before the converted information is used as input to the sleep analysis model.

[0294] Furthermore, according to one embodiment of the present invention, the converted information is converted so that acoustic information can be used as input for an image processing-based artificial intelligence model, so the acoustic information may be visualized through such a preprocessing process before being used as input.

[0295] Such a preprocessing step may be performed only in the learning step, or may be performed in both the learning step and the inference step, or may be performed only in the inference step.

[0296] Preprocessing for data augmentation

[0297] FIG. 15 is a diagram illustrating pitch shifting according to an embodiment of the present invention.

[0298] According to an embodiment of the present invention, a pre-processing method may be included that performs data augmentation on the spectrogram.

[0299] Data augmentation is used to ensure a sufficient amount of training data set or to conduct sufficient learning by assuming diverse and irregular environments.

[0300] Preprocessing methods for data augmentation according to embodiments of the present invention may include a pitch shifting method in which Gaussian noise is added to a spectrogram to expand the amount of data, or a pitch shifting method in which the pitch of the overall acoustic information is gradually raised or lowered, and a TUT (Tile UnTile) augmentation method in which a spectrogram or mel-spectrogram is converted into a vector during the learning process, and the converted vector is randomly tiled at the input stage of a node (neuron) and recombined after the output of the node (neuron).

[0301] In addition, data augmentation according to an embodiment of the present invention may include a noise addition augmentation method that adds noises that occur in various environments other than Gaussian noise (e.g., external sounds, natural sounds, fan sounds, door opening and closing sounds, animal sounds, human conversation sounds, movement sounds, etc.).

[0302] In order to shorten the learning time when a spectrogram is used as an input for a learning model, noise-added augmentation according to an embodiment of the present invention may include a method of converting noise information into a spectrogram and then artificially adding noise information to the sleep audio information on the spectrogram. In this case, there may be little difference between a spectrogram obtained by converting the entire sleep audio information into the original audio information domain and a spectrogram obtained by adding noise information to the sleep audio information and the noise information on the domains converted into spectrograms.

[0303] In addition, in order to make it difficult to convert a spectrogram back to an original signal and to protect the privacy of a user, the noise-added augmentation according to an embodiment of the present invention maintains only the amplitude of the amplitude and phase in the spectrograms of the sleep sound information and the noise information, and adds an arbitrary phase, thereby making it difficult to convert a spectrogram back to an original signal.

[0304] Alternatively, the noise-added augmentation according to the embodiment of the present invention may include not only a method of adding acoustic information in a domain converted into a spectrogram, but also a method of adding acoustic information in a domain converted into a mel spectrogram to which a mel scale is applied.

[0305] Additionally, the mel-scale addition method according to embodiments of the present invention can reduce the time it takes for hardware to process the data.

[0306] Meanwhile, the specific description of the types of noise mentioned above is merely an example for explaining the noise-added augmentation of the present invention, and the present invention is not limited thereto.

[0307] FIG. 23 is a diagram illustrating a TUT (Tile UnTile) augmentation method according to an embodiment of the present invention.

[0308] Tile / UnTile (TUT) augmentation according to an embodiment of the present invention may include randomly cutting and combining spectrograms or vectors at the node (neuron) input and output stages to increase the amount of training data for diverse patterns when a spectrogram is used as input to a learning model. Randomly cut spectrograms or vectors at the node (neuron) input stage may have data leakage or loss of information compared to uncut spectrograms or vectors input to the corresponding neural network layer. In this case, limited information may be input to the node (neuron) for learning. A node (neuron) that receives a cut spectrogram or vector as input may output a vector after calculation. The cut spectrogram or vector may be untiled again in the same manner as the cut before being input to a node (neuron) in the next neural network layer.

[0309] In addition, TUT augmentation according to an embodiment of the present invention can induce learning of data with information leakage by randomly cutting spectrograms or vectors at the input and output stages of nodes (neurons) and combining them in the same manner, thereby contributing to improving the accuracy or reliability of the learning model.

[0310] Preprocessing to convert to a square shape

[0311] FIG. 16 is a diagram illustrating a pre-processing method for converting information in the frequency domain or a spectrogram into a form close to a square, according to an embodiment of the present invention.

[0312] According to an embodiment of the present invention, after the preprocessing process of data augmentation of frequency domain information or spectrogram, such as the above-mentioned pitch shifting, noise-added augmentation, or TUT augmentation, a preprocessing method of converting frequency domain information or spectrogram into a form close to a square may be performed.

[0313] According to one embodiment of the present invention, the frequency domain information or spectrogram on which data augmentation has been performed can be converted into a form close to a square before being used as input to a deep learning model such as a CNN, a Transformer, a Vision Transformer (ViT), or a Mobile Vision Transformer (MobileViT), and then the frequency domain information or spectrogram converted into such a form close to a square can be used as input to an AI deep learning model.

[0314] According to an embodiment of the present invention, when performing preprocessing to convert information in the frequency domain or a spectrogram into a form close to a square, the form can be converted into a form close to a square using various methods such as reshaping, resizing, and split-cat.

[0315] According to an embodiment of the present invention, when preprocessing is performed to convert an image into a shape close to a square using a resizing method, a method of lowering the resolution of the x-axis and increasing the resolution of the y-axis by copying values may be performed.

[0316] In addition, according to an embodiment of the present invention, a preprocessing method is performed on the entire 30-second spectrogram having dimensions of 20 frequency bins x 1201 time steps, converting it into a nearly square shape at once using a resizing method, and missing information can be filled in using an interpolation method.

[0317] A split-cat preprocessing method according to an embodiment of the present invention may include a method of splitting a spectrogram into a certain size and then using a concatenation function to adjust the data to a shape close to a square. In other words, the split-cat method refers to a method of splitting a spectrogram to correspond to a patch and then adjusting the patch to a shape close to a square in order to perform learning in a patch unit in a deep learning model based on a Vit (Vision Transformer) or a Mobile Vit (Mobile Vision Transformer).

[0318] For example, according to one embodiment of the present invention, a 30-second spectrogram having dimensions of 20 frequency bins x 1201 time steps can be converted to dimensions of 150 frequency bins x 160 time steps. Then, a resizing technique can be used to convert the 30-second spectrogram having dimensions of 150 frequency bins x 160 time steps to dimensions close to 160 frequency bins x 160 time steps. Through this process, spectrograms corresponding to 30-second intervals can be converted into shapes close to squares.

[0319] When learning is performed using a spectrogram converted into a nearly square form as an input, a deep learning model based on a Transformer learning model can exhibit further improved learning performance. The specific numerical descriptions of the spectrogram bins, division time units, and number of divisions described above are merely examples, and the present invention is not limited thereto.

[0320] Scale conversion and normalization preprocessing

[0321] In the frequency domain information or spectrogram according to an embodiment of the present invention, if the value is very small and is not converted to another scale, the part where the value is greater than a certain level is expressed very brightly, while the remaining part is expressed very darkly, and therefore, it may be unsuitable as an input for a deep learning model. Therefore, before the frequency domain information or spectrogram according to an embodiment of the present invention is used as an input for a deep learning model, a preprocessing process of converting it into a dB scale (log scale) may be performed.

[0322] In performing preprocessing of the log scale conversion according to an embodiment of the present invention, the maximum value of the log value may be set to 0 as a basic base value, and the remaining values may be converted to log values.

[0323] According to an embodiment of the present invention, the spectrogram converted into log values can be subjected to an additional preprocessing step of normalization, which makes the overall mean value 0 and the standard deviation 1, before being used as input for a deep learning model.

[0324] The preprocessed data is used as an input for an image processing deep learning model, and information such as sleep state information can be learned or inferred through image analysis of the spectrogram. The specific numerical values for the maximum log values of the spectrograms described above are merely examples, and the present invention is not limited thereto.

[0325] Sleep Analysis Model

[0326] FIG. 25 is a flow chart illustrating an example of a method for analyzing a user's sleep state through acoustic information according to an embodiment of the present invention.

[0327] According to one embodiment of the present invention, the method may include a step (S10) of acquiring sleep acoustic information associated with a user's sleep.

[0328] According to an embodiment of the present invention, the method may include a step (S20) of performing pre-processing on sleep acoustic information.

[0329] According to an embodiment of the present invention, the method may include the step of acquiring sleep state information by performing an analysis on the pre-processed sleep acoustic information (S30).

[0330] 25 may be reordered and at least one step may be omitted or added as necessary. That is, the above-described steps are merely one embodiment of the present invention, and the scope of the present invention is not limited thereto.

[0331] 17a and 17b are diagrams illustrating the overall structure of a sleep analysis model according to an embodiment of the present invention.

[0332] In the present invention, the sleep state information may be obtained through a sleep analysis model that analyzes the sleep stage of the user based on acoustic information (sleep acoustic information).

[0333] In the present invention, since the sleep sound information SS is sound associated with breathing and body movements acquired during a user's sleep, it may be very quiet. Therefore, as described above, the present invention can convert the sleep sound information SS into a spectrogram SP to perform sound analysis. In this case, the spectrogram SP includes information indicating how the frequency spectrum of the sound changes over time, making it possible to easily identify breathing or movement patterns associated with relatively quiet sounds, thereby improving analysis efficiency. Specifically, it may be difficult to predict whether the sleep sound information is in at least one of an awake state, a REM sleep state, a light sleep state, and a deep sleep state based solely on changes in the energy level of the sleep sound information. However, by converting the sleep sound information into a spectrogram, changes in the spectrum of each frequency can be easily detected, making it possible to perform analysis that is appropriate for quiet sounds (e.g., breathing and body movements).

[0334] According to an embodiment of the present invention, the processor 130 may process the transformed frequency domain information or spectrogram SP as an input of a sleep analysis model to acquire sleep state information. Here, the sleep analysis model is a model for acquiring sleep state information related to changes in a user's sleep stage, and may input sleep acoustic information acquired during the user's sleep and output sleep state information. In an embodiment, the sleep analysis model may include a neural network model configured via one or more network functions.

[0335] Networking Functions

[0336] The sleep analysis model is composed of one or more network functions, which may be composed of a collection of interconnected computational units that may generally be referred to as "nodes." Such "nodes" may also be referred to as "neurons." One or more network functions are composed of at least one or more nodes. The nodes (or neurons) that make up one or more network functions may be interconnected by one or more "links."

[0337] FIG. 7 is a schematic diagram illustrating one or more network functions for performing a sleep analysis method according to the present invention. A deep neural network (DNN) may refer to a neural network including multiple hidden layers in addition to an input layer and an output layer. A deep neural network can be used to understand the latent structures of data. That is, it can understand the latent structures of photos, text, video, audio, and music (e.g., what objects are in the photo, what is the content and emotion of the text, what is the content and emotion of the audio, etc.). Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, generative adversarial networks (GANs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q-networks, U-networks, Siamese networks, Transformers, Vision Transformers (ViTs), Mobile Vision Transformers (Mobile ViTs), etc. The above description of deep neural networks is for illustrative purposes only and the present invention is not limited thereto.

[0338] In the present invention, the network function may include an autoencoder. An autoencoder may be a type of artificial neural network that outputs output data similar to input data. An autoencoder may include at least one hidden layer, and an odd number of hidden layers may be arranged between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding), and may be reduced and symmetrically expanded from the bottleneck layer to the output layer (symmetric to the input layer). The nodes in the dimensionality reduction layer and the dimensionality restoration layer may or may not be symmetric. An autoencoder may perform nonlinear dimensionality reduction. The number of input and output layers may correspond to the number of sensors remaining after preprocessing of the input data. In an autoencoder structure, the number of nodes in the hidden layer included in the encoder may decrease as it moves away from the input layer. If the number of nodes in the bottleneck layer (the layer with the fewest nodes located between the encoder and decoder) is too small, a sufficient amount of information may not be transmitted, so it may be maintained at a certain number or more (e.g., more than half of the number of nodes in the input layer).

[0339] Neural networks can be trained using at least one of supervised learning, unsupervised learning, and semi-supervised learning. Neural network training is performed to minimize output errors. Training involves repeatedly inputting training data into the neural network, calculating the error between the neural network's output and the target for the training data, and backpropagating the neural network's error from the output layer to the input layer in a direction that reduces the error, thereby updating the weights of each node in the neural network. In supervised learning, training data in which a correct answer is labeled is used for each piece of training data (i.e., labeled training data), while in unsupervised learning, the correct answer may not be labeled for each piece of training data. For example, in supervised learning for data classification, the training data may be data in which a category is labeled for each piece of training data. Labeled training data may be input to a neural network, and an error may be calculated by comparing the neural network output (category) with the label of the training data. As another example, in the case of unsupervised learning for data classification, an error may be calculated by comparing the input training data with the neural network output. The calculated error may be backpropagated from the neural network in the backward direction (i.e., from the output layer to the input layer), and the connection weights of each node in each layer of the neural network may be updated by backpropagation. The amount of change in the connection weights of each updated node may be determined according to a learning rate. The neural network calculation for input data and the backpropagation of the error may constitute a learning cycle (epoch). The learning rate may be applied differently depending on the number of iterations of the neural network learning cycle.For example, in the early stages of neural network learning, a high learning rate can be used to quickly ensure a certain level of performance, thereby increasing efficiency, and in the later stages of learning, a lower learning rate can be used to increase accuracy.

[0340] In neural network training, training data may generally be a subset of actual data (i.e., data to be processed using the trained neural network). Therefore, there may be a learning cycle in which errors decrease for training data but increase for actual data. Overfitting is a phenomenon in which excessive training on training data increases errors for actual data. For example, a neural network that has learned cats by showing yellow cats may be unable to recognize cats that are not yellow as cats, which may be a type of overfitting. Overfitting can cause an increase in errors in AI algorithms. Various optimization methods may be used to prevent overfitting. To prevent overfitting, methods such as increasing the training data, regularization, and dropout, which omits some network nodes during the training process, may be used.

[0341] Throughout this specification, the terms "computational model," "neural network," "network function," and "neural network" may be used interchangeably (hereinafter, they will be referred to interchangeably as "neural network"). A data structure may include a neural network. The data structure including a neural network may be stored on a computer-readable medium. The data structure including a neural network may also include data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and a loss function for training the neural network. The data structure including a neural network may include any of the components disclosed above. That is, the data structure including a neural network may include all or any combination of data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and a loss function for training the neural network. In addition to the above-mentioned components, the data structure including a neural network may include any other information that determines the characteristics of the neural network. Furthermore, the data structure may include all types of data used or generated in the computational process of a neural network, and is not limited to the foregoing. The computer-readable medium may include a computer-readable recording medium and / or a computer-readable transmission medium. A neural network may be composed of a collection of interconnected computational units, which may generally be referred to as nodes. Such nodes may also be referred to as neurons. A neural network is composed of at least one or more nodes.

[0342] Within a neural network, one or more nodes connected via links can form a relative relationship between input and output nodes. The concepts of input and output nodes are relative, and any node that is in an output node relationship with one node can also be in an input node relationship with another node, and vice versa. As mentioned above, the input node-to-output node relationship can be generated around links. One or more output nodes can be connected to one input node via links, and vice versa.

[0343] In a relationship between an input node and an output node connected via a link, the value of the output node may be determined based on data input to the input node. Here, the node interconnecting the input node and the output node may have a weight. The weight may be variable and may be varied by a user or an algorithm so that the neural network performs a desired function. For example, when one or more input nodes are interconnected to one output node via respective links, the output node may determine its output node value based on the value input to the input node connected to the output node and the weight assigned to the link corresponding to each input node.

[0344] As described above, a neural network has one or more nodes interconnected via one or more links to form input and output node relationships within the neural network. The characteristics of a neural network may be determined by the number of nodes and links, the relationships between the nodes and links, and the weights assigned to each link. For example, if two neural networks have the same number of nodes and links but different weights between the links, the two neural networks may be recognized as different.

[0345] Some of the nodes constituting a neural network can be configured as a layer based on their distance from the initial input node. For example, a set of nodes whose distance from the initial input node is n can constitute n layers. The distance from the initial input node can be defined by the minimum number of links that must be traversed to reach the node from the initial input node. However, this definition of a layer is arbitrary for the purpose of explanation, and the number of layers in a neural network can be defined in a different manner than that described above. For example, a node's layer can be defined by its distance from the final output node.

[0346] An initial input node may refer to one or more nodes in a neural network to which data is directly input without passing through a link in relation to other nodes. Alternatively, it may refer to a node in a neural network that does not have any other input nodes connected to it in relation to the links between nodes. Similarly, a final output node may refer to one or more nodes in a neural network that do not have any output nodes in relation to other nodes. Furthermore, a hidden node may refer to a node constituting a neural network that is neither an initial input node nor a final output node. A neural network according to an embodiment of the present invention may have more nodes in an input layer than in a hidden layer closer to the output layer, and the number of nodes may decrease as one moves from the input layer to the hidden layer.

[0347] A neural network may include one or more hidden layers. Hidden nodes in a hidden layer may receive the output of the previous layer and the output of surrounding hidden nodes as input. The number of hidden nodes in each hidden layer may be the same or different. The number of nodes in the input layer may be determined based on the number of data fields in the input data, and may be the same or different from the number of hidden nodes. Input data input to the input layer may be operated on by hidden nodes in the hidden layer and output by a fully connected layer (FCL), which is the output layer.

[0348] Feature extraction and classification models

[0349] FIG. 18 is a diagram illustrating a feature extraction model and a feature classification model according to an embodiment of the present invention.

[0350] FIG. 19 is a diagram for explaining in detail the operation of a sleep analysis model according to an embodiment of the present invention.

[0351] The sleep analysis model used in the present invention may include a feature extraction model that extracts one or more features for each predetermined epoch, and a feature classification model that classifies each of the features extracted by the feature extraction model into one or more sleep stages to generate sleep state information. The feature extraction model may analyze the time-series frequency pattern of the spectrogram SP to extract features related to respiratory sounds, breathing patterns, and movement patterns. In one embodiment, the feature extraction model may be configured as a part of a neural network model pre-trained using a training dataset.

[0352] The sleep analysis model used in the present invention may include a feature extraction model and a feature classification model. The feature extraction model may be a deep learning model based on a natural language processing model that can learn the time-series relevance of given data. The feature classification model may be a learning model based on a natural language processing model that can learn the time-series relevance of given data. Here, natural language processing model-based deep learning models that can learn the time-series relevance may include, but are not limited to, Turnsformer, ViT, MobileViT, and MobileViT2.

[0353] A training data set according to an embodiment of the present invention may be composed of data in the frequency domain and a plurality of pieces of sleep state information corresponding to each piece of data.

[0354] Alternatively, a training dataset according to an embodiment of the present invention may be composed of a plurality of spectrograms and a plurality of pieces of sleep state information corresponding to each spectrogram.

[0355] Alternatively, a training dataset according to an embodiment of the present invention may be composed of a plurality of mel spectrograms and a plurality of pieces of sleep state information corresponding to each mel spectrogram.

[0356] For ease of explanation, the configuration and execution of a sleep analysis model according to an embodiment of the present invention will be described in detail below based on a spectrogram data set. However, the learning data used in the sleep analysis model of the present invention is not limited to spectrograms, and information in the frequency domain, spectrograms or mel-spectrograms may be used as learning data.

[0357] The feature extraction model of the sleep analysis model according to an embodiment of the present invention may be pre-trained by a one-to-one proxy task in which one spectrogram is input and the model is trained to predict sleep state information corresponding to the spectrogram. When a CNN deep learning model is used as the feature extraction model according to an embodiment of the present invention, a fully connected layer (FC) or fully connected neural network (FCN) structure may be adopted for training. When a MobileViTV2 deep learning model is used as the feature extraction model according to an embodiment of the present invention, an intermediate layer structure may be adopted for training.

[0358] A feature classification model among the sleep analysis models according to an embodiment of the present invention may be trained to receive multiple consecutive spectrograms, predict sleep state information for each spectrogram, and analyze a sequence of multiple consecutive spectrograms to predict or classify overall sleep state information.

[0359] Furthermore, according to an embodiment of the present invention, after pre-training is performed through a one-to-one proxy task for a feature extraction model, fine-tuning can be performed through a many-to-many task for the pre-trained feature extraction model and feature classification model. For example, a sequence of 40 consecutive spectrograms can be input to multiple feature extraction models trained through a one-to-one proxy task, and 20 pieces of sleep state information can be output to infer sleep stages. The specific numerical values related to the number of spectrograms, the number of feature extraction models, and the number of pieces of sleep state information described above are merely examples, and the present invention is not limited thereto.

[0360] Hereinafter, a feature extraction model and a feature classification model that are generated or trained based on a transformed spectrogram according to an embodiment of the present invention will be described in detail. Meanwhile, the sleep analysis model of the present invention is not limited to being generated or trained based on a spectrogram, but may be generated or trained based on information including changes in frequency components of raw acoustic information over time or information in the frequency domain, as described above. Furthermore, inference of sleep state information through the sleep analysis model may also be performed based on information including changes in frequency components of raw acoustic information over time or information converted into information in the frequency domain.

[0361] 29a and 29b are diagrams illustrating the performance of sleep disorder determination and noise addition using a spectrogram in the sleep analysis method according to the present invention.

[0362] As shown in Figures 29a and 29b, a sleep analysis model according to an embodiment of the present invention can reliably detect apnea even when the spectrogram is corrupted by noise.

[0363] Feature Extraction Model

[0364] The feature extraction model may be configured as a unique deep learning model trained through a training dataset. The feature extraction model may be trained through supervised learning or unsupervised learning. The feature extraction model may be trained through a training dataset to output output data similar to the input data. In more detail, only the core feature data (or features) of the input spectrogram may be learned through the hidden layer. In this case, the output data of the hidden layer during the decoding process through the decoder may be an approximation of the input data (i.e., the spectrogram) rather than a perfect copy.

[0365] Each of the plurality of spectrograms included in the training dataset may be tagged with sleep state information. Each of the plurality of spectrograms may be input to a feature extraction model, and an output corresponding to each spectrogram may be matched with the tagged sleep state information and stored. Specifically, when a first training dataset (i.e., a plurality of spectrograms) tagged with first sleep state information (e.g., light sleep) is used as an input, features associated with the output for the input may be matched with the first sleep state information and stored. In an embodiment, one or more features associated with the output may be displayed in a vector space. In this case, feature data output corresponding to each of the first training datasets may be output via spectrograms associated with a first sleep stage, and therefore may be located relatively close to each other in the vector space. That is, training may be performed so that a plurality of spectrograms corresponding to each sleep stage output similar features.

[0366] The feature extraction model obtained through the above-described learning process can extract features corresponding to a spectrogram when the spectrogram (for example, a spectrogram converted in accordance with sleep acoustic information) is input.

[0367] In an embodiment, the processor 130 may extract features by processing the spectrogram SP generated corresponding to the sleep sound information SS as an input of a feature extraction model. Here, since the sleep sound information SS is time-series data acquired in a time-series manner during the user's sleep, the processor 130 may divide the spectrogram SP into predetermined epochs. For example, the processor 130 may divide the spectrogram SP corresponding to the sleep sound information SS into 30-second intervals to acquire a plurality of spectrograms. For example, if the sleep sound information is acquired during a user's 7-hour (i.e., 420-minute) sleep, the processor 130 may divide the spectrogram into 30-second intervals to acquire 840 spectrograms. The specific numerical values for the sleep duration, the time interval for dividing the spectrogram, and the number of divisions described above are merely examples, and the present invention is not limited thereto.

[0368] The processor 130 processes each of the divided spectrograms as an input of a feature extraction model to extract a plurality of features corresponding to each of the plurality of spectrograms. For example, if the number of spectrograms is 840, the number of features extracted by the feature extraction model may also be 840. The specific numerical values related to the number of spectrograms and features described above are merely examples, and the present invention is not limited thereto.

[0369] Meanwhile, the feature extraction model according to an embodiment of the present invention may be trained using a one-to-one proxy task. In addition, in the process of training to extract sleep state information for one spectrogram, the feature extraction model may be trained to extract sleep state information by combining it with another neural network (NN).

[0370] According to an embodiment of the present invention, if learning is performed through a pre-trained simple neural network, the learning time of the feature extraction model can be shortened or the learning efficiency can be improved.

[0371] For example, according to one embodiment of the present invention, a spectrogram divided into 30-second intervals can be trained so that the vector output as input to a feature extraction model can be used as input to a different NN to output sleep state information.

[0372] Meanwhile, according to an embodiment of the present invention, the processor 130 may generate a feature extraction model using multiple source data, and the feature extraction model may include a dimensionality reduction network function (e.g., an encoder).

[0373] In an embodiment, each of a plurality of pieces of information in frequency domains or spectrograms (i.e., a plurality of pieces of transformed information corresponding to a plurality of source data) used as training data may be labeled with sleep stage information. For example, the plurality of source data may be information about sleep acoustics acquired in a specific space (e.g., a hospital), and information about a plurality of sleep states (i.e., sleep stages) may be pre-labeled.

[0374] Each of the transformed pieces of information may be input to a dimension reduction network function, and an output corresponding to each transformed piece of information may be matched with the labeled sleep stage information. Specifically, when a first training dataset (e.g., a plurality of spectrograms associated with the source data) labeled with first sleep stage information (e.g., light sleep) is input to the dimension reduction network function, features associated with the output of the dimension reduction network function for that input may be matched with the first sleep stage information.

[0375] In an embodiment, one or more features associated with the output of the dimensionality reduction network function may be represented in a vector space, in which the feature data output corresponding to each of the first training datasets may be located relatively close to each other in the vector space because they are outputs via spectrograms associated with the first sleep stage (e.g., outputs via spectrograms corresponding to the same class).

[0376] That is, the dimension reduction network function may be trained so that a plurality of spectrograms corresponding to each sleep stage output similar features, but the specific training method of the dimension reduction network is not limited.

[0377] Through the above-described learning process, when information converted corresponding to sleep acoustic information is input, the feature extraction model can extract features corresponding to the converted information.

[0378] Feature Classification Model

[0379] FIG. 19 is a diagram for explaining in detail the operation of a sleep analysis model according to an embodiment of the present invention.

[0380] According to an embodiment of the present invention, processor 130 may acquire sleep state information by processing a plurality of features output by the feature extraction model as input to a feature classification model. In an embodiment, the feature classification model may be a neural network model modeled to predict a sleep stage corresponding to a feature. For example, the feature classification model may be configured with fully connected layers and classify a feature into at least one sleep stage. For example, when a first feature corresponding to a first spectrogram is input, the feature classification model may classify the first feature as light sleep. Or, for example, when a second feature corresponding to a second spectrogram is input, the feature classification model may classify the second feature as deep sleep. Or, when a third feature corresponding to a third spectrogram is input, the feature classification model may classify the third feature as REM sleep. Or, for example, when a fourth feature corresponding to a fourth spectrogram is input, the feature classification model may classify the fourth feature as wakefulness.

[0381] In addition, in an embodiment of the present invention, the feature classification model may be a neural network model modeled to predict sleep events corresponding to features. For example, the feature classification model may be configured to include a fully connected layer and classify features into at least one of sleep events. For example, when a first feature corresponding to a first spectrogram is input, the feature classification model may classify the first feature as a sleep apnea event. For example, when a second feature corresponding to a second spectrogram is input, the feature classification model may classify the second feature as a sleep hypopnea event. For example, when a third feature corresponding to a third spectrogram is input, the feature classification model may classify the third feature as a normal sleep state. For example, when a fourth feature corresponding to a fourth spectrogram is input, the feature classification model may classify the fourth feature as a snoring event during sleep. For example, when a fifth feature corresponding to a fifth spectrogram is input, the feature classification model may classify the fifth feature as a sleep talking event during sleep.

[0382] According to an embodiment of the present invention, a feature classification model may perform classification for a plurality of features. The feature classification model may classify each of the plurality of features into at least one of a plurality of sleep stages. Furthermore, according to an embodiment of the present invention, a feature extraction model may extract features to facilitate classification of which sleep stage the feature classification model corresponds to, and the feature classification model may be trained to better classify features transferred from the feature extraction model (i.e., to better classify specific sleep stages). In other words, through adversarial learning, the feature classification model may be able to better classify features transferred from the feature extraction model. In one embodiment, the feature classification model may be trained to facilitate classifying features to better perform sleep stage classification or sleep event classification corresponding to the features.

[0383] According to an embodiment of the present invention, the processor 130 may update the feature extraction model through first learning information related to a first loss between the feature extraction model and the feature classification model in accordance with the results of learning the feature extraction model and the feature classification model. This allows the updated feature extraction model to extract features that enable the feature classification model to properly classify classes (i.e., sleep stages). In other words, the updated feature extraction model may extract features such that features associated with the same sleep stage are clustered together.

[0384] According to one embodiment, the discriminator model may be a neural network model that distinguishes whether each of the plurality of features transferred from the feature extraction model is a feature associated with the source data or a feature associated with the target data. For example, the discriminator model may distinguish whether an input-related feature corresponds to sleep acoustic information acquired in a hospital or to sleep acoustic information acquired in the real life of an individual user. That is, the discriminator model may use at least one of the plurality of features as an input and distinguish whether the input-related feature is a feature associated with the source data or a feature associated with the target data. Training of the sleep analysis model using the discriminator model will be described in detail below.

[0385] The feature classification model may perform multi-epoch classification, which predicts sleep stages of various epochs using spectrograms associated with various epochs as input. Multi-epoch classification does not provide analysis information for a single sleep stage corresponding to a spectrogram of a single epoch (i.e., one spectrogram corresponding to 30 seconds), but may instead estimate various sleep stages (e.g., changes in sleep stages over time) simultaneously using spectrograms corresponding to multiple epochs (i.e., a combination of spectrograms each corresponding to 30 seconds). For example, because breathing patterns or movement patterns change more slowly than electroencephalograms or other biological signals, accurate sleep stage estimation may be possible only by observing how the patterns change between past and future points in time. For example, the feature classification model may perform prediction for the central 20 spectrograms using 40 spectrograms (e.g., 40 spectrograms each corresponding to 30 seconds) as input. That is, by examining all spectrograms 1 to 40 in detail, it is possible to predict the sleep stage through classification corresponding to spectrograms corresponding to 10 to 20. The specific numerical values for the number of spectrograms described above are merely examples, and the present invention is not limited thereto.

[0386] That is, in the process of inferring sleep state information, rather than performing prediction of sleep state information corresponding to each single spectrogram, spectrograms corresponding to multiple epochs are used as input so that all information related to the past and future can be taken into consideration, thereby improving the accuracy of the output. Meanwhile, the accuracy of the output can be improved by performing inference using not only spectrograms but also information including changes in frequency components corresponding to multiple epochs over time or information in the frequency domain as input.

[0387] FIG. 8 is a diagram for explaining sleep stage analysis using a spectrogram in the sleep analysis method according to the present invention.

[0388] According to an embodiment of the present invention, after the primary sleep analysis based on actigraphy and HRV, the secondary analysis based on sleep acoustic information uses the sleep analysis model as described above, and when the user's sleep acoustic information is input, the corresponding sleep stage (Wake, REM, Light, Deep) may be immediately inferred, as shown in Fig. 8. In addition, the secondary analysis based on sleep acoustic information may extract the time point at which a sleep disorder (sleep apnea, hyperventilation), snoring, etc. occurs through the singular points of the Mel spectrum corresponding to the sleep stage.

[0389] FIG. 9 is a diagram illustrating sleep event determination using a spectrogram in the sleep analysis method according to the present invention.

[0390] 9, by analyzing a breathing pattern in a transformed frequency domain or a spectrogram or mel-spectrogram, if a characteristic corresponding to a sleep apnea or hyperpnea event is detected, the time point can be determined as a time point at which a sleep event occurred. In this case, a step of classifying the event as snoring rather than sleep apnea or hyperpnea through frequency analysis may be further included.

[0391] FIG. 10 is a diagram showing an experimental process for verifying the performance of the sleep analysis method according to the present invention.

[0392] As shown in Figure 10, a user's sleep video and sleep sounds are acquired in real time, and the acquired sleep sound information is immediately converted into a spectrogram. At this time, a pre-processing process of the sleep sound information may be performed. The spectrogram may be input into a sleep analysis model to immediately analyze the sleep stages.

[0393] Furthermore, when a CNN or Transformer-based deep learning model is adopted as a feature classification model according to an embodiment of the present invention, the operation may be performed as follows.

[0394] According to an embodiment of the present invention, a spectrogram including time series information may be input to a CNN-based deep learning model to output a vector with reduced dimensionality. In this manner, the vector with reduced dimensionality may be input to a Transformer-based deep learning model to output a vector including time series information.

[0395] According to an embodiment of the present invention, an output vector of a Transformer-based deep learning model can be input to a 1D Convolutional Neural Network (1D CNN) so that an average pooling technique can be applied to the output vector, and a process of converting the output vector into an N-dimensional vector containing time series information through an averaging operation on the time series information can be performed. In this case, the N-dimensional vector containing time series information is still data containing time series information, although there is a difference in resolution from the input data.

[0396] According to an embodiment of the present invention, it is possible to perform multi-epoch classification on a combination of N-dimensional vectors including output time series information, and to perform prediction of various sleep stages. In this case, it is also possible to perform prediction of continuous sleep state information by using output vectors of a Transformer-based deep learning model as inputs of multiple Fully Connected layers (FCs).

[0397] Furthermore, when a deep learning model based on ViT or Mobile ViT is adopted as the feature classification model according to an embodiment of the present invention, the operation may be performed as follows: Figure 24 is a diagram illustrating the structure of a sleep analysis model using a natural language processing model according to an embodiment of the present invention.

[0398] According to an embodiment of the present invention, a spectrogram containing time-series information can be used as an input for a Mobile ViT-based deep learning model to output a vector with reduced dimensionality.

[0399] Furthermore, according to an embodiment of the present invention, features can be extracted from each spectrogram at the output of the Mobile ViT-based deep learning model.

[0400] According to an embodiment of the present invention, a vector with reduced dimension may be input to an intermediate layer, and a vector containing time series information may be output. The intermediate layer model may include at least one of a linearization step that implies vector information, a layer normalization step that inputs the mean and variance, and a dropout step that deactivates some nodes.

[0401] According to an embodiment of the present invention, a process of outputting a vector containing time series information using a vector with reduced dimensions as an input to an intermediate layer is performed, thereby preventing overfitting.

[0402] According to an embodiment of the present invention, the output vector of the intermediate layer can be used as an input to a ViT-based deep learning model to output sleep state information. In this case, sleep state information corresponding to frequency domain information, spectrogram, or mel-spectrogram including time-series information can be output.

[0403] Furthermore, an embodiment of the present invention can output sleep state information corresponding to a series of frequency domain information, spectrograms, or mel-spectrograms containing time series information.

[0404] Meanwhile, in the feature extraction model or feature classification model according to the embodiment of the present invention, various artificial intelligence models other than the AI models mentioned above may be adopted to perform learning or inference, and the specific description related to the types of artificial intelligence models mentioned above is merely an example, and the present invention is not limited thereto.

[0405] Unsupervised or semi-supervised learning method for sleep analysis models according to embodiments of the present invention

[0406] FIG. 20 is a diagram illustrating an unsupervised or semi-supervised learning model according to an embodiment of the present invention.

[0407] A dataset according to an embodiment of the present invention may be composed of labeled data acquired in a specific environment (preferably, a polymorphic sleep test environment), or may be composed of unlabeled data acquired in another environment (preferably, an environment other than a polymorphic sleep test). The specific descriptions of the environments described above are merely examples, and the present invention is not limited thereto.

[0408] Supervised learning is possible when learning using labeled data, which is learning data in which the correct answer is labeled. However, unsupervised learning is required when learning using unlabeled data in which the correct answer is not labeled. The following describes an unsupervised learning model according to an embodiment of the present invention.

[0409] Consistency Training using noise in the target environment

[0410] FIG. 21 is a diagram illustrating consistency training according to an embodiment of the present invention.

[0411] Consistency training is a type of semi-supervised learning model, and consistency training according to an embodiment of the present invention may be a method of performing learning using data to which noise has been intentionally added and data to which noise has not been intentionally added.

[0412] Furthermore, consistency training according to an embodiment of the present invention may be a method of performing learning by generating data of a virtual sleep environment using noise of a target environment.

[0413] According to an embodiment of the present invention, the intentionally added noise may be noise of a target environment, where the noise of the target environment may be noise acquired in an environment other than a polymorphic sleep study, for example.

[0414] Hereinafter, for convenience, data to which noise has been intentionally added will be referred to as “corrupted data.” Corrupted data may preferably mean data to which noise of the target environment has been intentionally added.

[0415] For convenience, data to which noise has not been intentionally added will be referred to as "clean data." However, clean data may actually contain noise, but the noise has not been intentionally added.

[0416] The clean data used in consistency training according to one embodiment of the present invention may be data acquired in a specific environment (preferably, a sleep polymorphism test environment), and the corrupted data may be data acquired in another environment or a target environment (preferably, an environment other than a sleep polymorphism test).

[0417] Corrupted data according to an embodiment of the present invention may be data in which noise acquired in another environment or target environment (preferably an environment other than a polymorphic sleep test) is intentionally added to clean data.

[0418] In consistency training, when clean data and corrupted data are input to the same deep learning model, a loss function or consistency loss may be defined so that the outputs are the same, and training may be performed to achieve consistent predictions.

[0419] In accordance with an embodiment of the present invention, in the process of adding noise to obtain corrupted data, a problem may arise in that the length of each acquired noise varies. In this case, a noise sampling method will be described for performing learning on various spectrograms.

[0420] According to an embodiment of the present invention, there may be at least nine types of noise, and several thousand pieces of acoustic information may be applied to each type of noise. When multiple spectrograms according to an embodiment of the present invention are input, each spectrogram may be divided into 30-second intervals, and 40 pieces (data corresponding to a total of 20-minute time intervals) may be input to a deep learning model. To match the time intervals with the input data, the noise may be arbitrarily sampled to correspond to a time interval corresponding to 20 minutes (e.g., 5 minutes, 9 minutes, 4 minutes, 7 minutes, etc.). If the total time interval of the sampled noise exceeds 20 minutes, the portion exceeding 20 minutes may be omitted.

[0421] According to an embodiment of the present invention, noise corresponding to the same time interval (e.g., 20 minutes) as the time interval of the spectrogram of clean data can also be converted into a spectrogram, and the corrupted data can be obtained by optionally adding the noise to the spectrogram of clean data.

[0422] In the process of optionally adding noise to a spectrogram according to an embodiment of the present invention, the spectrogram may be a mel spectrogram to which a mel scale is applied. Meanwhile, in the process of adding noise according to an embodiment of the present invention, the noise may be noise in the same domain as the acoustic information, or noise in the same domain as the spectrogram or the mel spectrogram to which a mel scale is applied. Here, the domain of the acoustic information may be a domain having amplitude, phase, and frequency.

[0423] Furthermore, the domain of a spectrogram or mel-spectrogram according to an embodiment of the present invention may be a domain having amplitude and frequency.

[0424] According to an embodiment of the present invention, when noise is converted into the same domain as a spectrogram or a mel spectrogram and added, an arbitrary phase can be assigned to the noise to proceed with the addition process. This makes it difficult to back-calculate data in the mel state, thereby maintaining data anonymity and protecting personal privacy, while reducing the amount of learning calculations and shortening the learning time.

[0425] Additionally, the mel-scale addition method according to the embodiment of the present invention can reduce the time required for hardware to process the data.

[0426] When the acquired corrupted data and cleaned data are input to the same deep learning model, training is performed so that the outputs are the same. Meanwhile, the specific descriptions related to the time interval and the number of spectrograms described above are merely examples to help understand the present invention, and the present invention is not limited thereto.

[0427] Unsupervised Domain Adaptation (UDA)

[0428] 22 is a diagram illustrating Unsupervised Domain Adaptation (UDA) according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 130 may perform learning on an artificial intelligence model including a feature extraction model, a feature classification model, and a discriminator model.

[0429] In the UDA according to an embodiment of the present invention, after the AI model has been fully trained through supervised learning using labeled data, additional training can be performed using only additionally provided unlabeled data.

[0430] Alternatively, the UDA according to an embodiment of the present invention may be configured and performed through primary learning and secondary learning.

[0431] In the first learning of the UDA according to an embodiment of the present invention, unlabeled data and labeled data can be utilized.

[0432] In the secondary learning of UDA according to an embodiment of the present invention, unlabeled data can be utilized.

[0433] The primary learning of the UDA according to an embodiment of the present invention may include performing learning to extract commonalities between data acquired in different environments as input to a single sleep analysis model. Alternatively, the primary learning may include performing learning to distinguish and classify differences between input data by using commonalities extracted from data acquired in different environments as input to a single sleep analysis model as input to a deep learning model. The primary learning of the UDA according to an embodiment of the present invention may include learning to extract common data (e.g., human sleep acoustic information) between labeled data and unlabeled data using a feature extraction model.

[0434] The labeled data used for the primary learning of the UDA according to one embodiment of the present invention may be data acquired in a specific environment (preferably, a sleep polymorphism test environment), and the unlabeled data may be data acquired in another environment or a target environment (preferably, an environment other than a sleep polymorphism test).

[0435] The labeled data used for the primary learning of UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a specific race (e.g., Koreans), and the unlabeled data may be sleep acoustic information acquired from other races (e.g., Asians, Blacks, Caucasians, Hispanics, etc.).

[0436] The labeled data used for the primary learning of the UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a specific gender (e.g., male), and the unlabeled data may be sleep acoustic information acquired from another gender (e.g., female).

[0437] The labeled data used for the primary learning of UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a specific age group (e.g., people in their 20s), and the unlabeled data may be sleep acoustic information acquired from other age groups (e.g., people in their teens, 30s, 40s, etc.).

[0438] The labeled data used for the primary learning of UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a specific body composition index population (e.g., a population with a body mass index (BMI) of 25 or more), and the unlabeled data may be sleep acoustic information acquired from another body composition index population (e.g., a population with a body mass index (BMI) of less than 25).

[0439] The labeled data used for the primary learning of UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a population with a sleep disorder (e.g., a population of people with sleep apnea), and the unlabeled data may be sleep acoustic information acquired from a population without a sleep disorder (e.g., a population without sleep apnea).

[0440] The labeled data used for the primary learning of UDA according to one embodiment of the present invention may be sleep acoustic information acquired from a population with a respiratory disease (e.g., a population of asthma patients), and the unlabeled data may be sleep acoustic information acquired from a population without a respiratory disease (e.g., a population without asthma).

[0441] The labeled data used for the primary learning of the UDA according to one embodiment of the present invention is not limited to applying each of the above-mentioned environments or characteristics individually, but may be acoustic information obtained from a combination of subject populations exhibiting one or more environments or characteristics.

[0442] Furthermore, the unlabeled data used for the primary learning of the UDA according to one embodiment of the present invention is not limited to applying each of the above-mentioned environments or characteristics individually, but may be acoustic information acquired from a combination of subject populations exhibiting one or more environments or characteristics.

[0443] In addition, the primary learning of the UDA according to an embodiment of the present invention may include learning to classify whether the input data is acquired from a specific environment or another environment or the target environment by using data acquired in a specific environment and data acquired in a different environment or a target environment as inputs to a feature extraction model and common data extracted as inputs to a discriminator model.

[0444] In this case, in the first learning of UDA according to one embodiment of the present invention, the feature extraction model is trained to output only commonalities between input data, and therefore can play a role in weakening classification between data acquired from a specific environment and data acquired from another environment or target environment, and the discriminator model can play a role in strengthening classification between data, so the losses applied to each model can be set inversely.

[0445] In an embodiment of the present invention, unlike data acquired from a specific environment, data acquired from another environment or a target environment may not have labeling associated with sleep state information. Therefore, data labeled with sleep state information acquired from a specific environment (e.g., a sleep multidimensional test environment) may be separately subjected to learning of sleep state information through a feature extraction model or a feature classification model. Here, for convenience, the feature extraction model or feature classification model that receives the labeled data and performs learning on the sleep state information will be referred to as a classifier.

[0446] If such learning is performed appropriately, the feature extraction model is trained to extract only commonalities between data acquired from a specific environment and data acquired from another environment or target environment as input. Therefore, even if the output value of unlabeled data acquired from another environment or target environment among the output data of the feature extraction model is input to the classifier, sleep state information may be output.

[0447] In summary, the primary learning of the UDA according to an embodiment of the present invention extracts or outputs common features of data acquired from a specific environment (e.g., a sleep polymorphism test environment) or another environment or target environment (e.g., an environment other than a sleep polymorphism test) through a feature extraction model, and inputs the output common features into a classifier model to perform learning so as to classify differences between data acquired from a specific environment and data acquired from another environment or target environment. Meanwhile, by training a classifier that acquires sleep state information from data acquired from a specific environment, data acquired from another environment or target environment can be input into the feature extraction model and the extracted information input into the classifier, so that learning can be performed so as to output sleep state information even if the data is unlabeled.

[0448] In addition, the output data of the classifier model according to the embodiment of the present invention may be recycled for the first learning of the UDA, or may be utilized in various ways, such as being used as a correction value for the final output of the classifier.

[0449] During the first learning of the UDA according to an embodiment of the present invention, in the process in which the classifier learns sleep state information from data acquired from a specific environment, data acquired from another environment or target environment (e.g., a home environment) is not input, and since the data acquired from the other environment or target environment does not have a label for the sleep stage, learning using the data acquired from the other environment or target environment may not be performed.

[0450] When the first learning of the UDA according to an embodiment of the present invention is ideally performed, the feature extraction model can extract commonalities between data. Therefore, when a classifier that uses the output commonalities as input outputs sleep state information, it can accurately classify whether the sleep stage corresponding to the input data is REM sleep, light sleep, wake state, or deep sleep through clustering. However, to further improve such clustering, a second learning can be performed using a technique such as conditional entropy.

[0451] In the secondary learning of UDA according to an embodiment of the present invention, learning can be performed using unlabeled data as input to a deep learning model.

[0452] When unlabeled data is input into a deep learning model, information such as prediction and confidence may be output. Here, a higher confidence may mean that the class information contained in the prediction value can be more reliable.

[0453] In the secondary learning of UDA according to an embodiment of the present invention, a process of learning using loss is performed to make the class information contained in the predicted value of the classifier's sleep state information or sleep stage information more reliable, and in this case, a method of self-learning using unlabeled output data may be included.

[0454] The specific descriptions of various environments or characteristics set forth above are merely examples and are not intended to limit the present invention.

[0455] Additionally, according to one embodiment of the present invention, the processor 130 may utilize a feature extraction model to extract a plurality of features corresponding to each of the plurality of source data and the plurality of target data. The processor 130 may process a plurality of transformed information (e.g., frequency domain information or spectrograms) corresponding to the plurality of source data and a plurality of transformed information (e.g., frequency domain information or spectrograms) corresponding to the plurality of target data at the input of the feature extraction model to extract a plurality of features. That is, a plurality of features associated with the output of the feature extraction model may include a plurality of features associated with the source data and a plurality of features associated with the target data.

[0456] In one embodiment, since the plurality of source data and the plurality of target data are data related to different domains, the transformed information (e.g., information in the frequency domain or spectrogram) corresponding to the plurality of source data may have information about sleep stages labeled in features extracted for each epoch, but the information corresponding to the plurality of target data may not have information about sleep stages labeled in features extracted for each epoch. In other words, since features corresponding to target data containing various noises do not have labeled information, sleep stage classification may be difficult.

[0457] Thus, according to one embodiment of the present invention, the processor 130 can arrange each of the plurality of first features corresponding to the plurality of source data and each of the plurality of second features corresponding to the plurality of target data adjacent to each other in the vector space, but also arrange each of the plurality of first features and each of the plurality of second features in clusters by class.

[0458] According to an embodiment of the present invention, a first feature corresponding to a plurality of source data and a second feature corresponding to a plurality of target data may each include information on various features (e.g., sleep stage features such as light sleep and REM sleep). Here, if a second feature corresponding to a plurality of target data does not have labeled information but blends well with the first feature corresponding to the plurality of source data, classification of the first feature may enable classification of the second feature, thereby enabling analysis of sleep acoustic information (i.e., multiple target data) containing various noises. In other words, although the second feature blends well with the first feature with labeled information, it may be important that the second feature be mapped to facilitate classification between classes.

[0459] Specifically, the second feature may be easily mapped for classification between classes. For example, if the first feature and the second feature are far apart in the vector space, when classifying the classes based on the labeling information of the first feature (e.g., when classifying the class of the first feature using a virtual line), features corresponding to different classes of the second feature may be classified into the same class, or features corresponding to the same class may be classified into different classes. In other words, the present invention allows the first feature and the second feature to be arranged adjacent to each other in the vector space, but allows the feature extraction model to be trained so that the features are arranged clustered by class.

[0460] To this end, the processor 130 may transmit a plurality of features to the feature classification model and the discriminator model, respectively.

[0461] According to one embodiment, the feature classification model may be a neural network model that classifies a plurality of features into one or more sleep stages. The feature classification model may be a neural network model trained to predict a sleep stage corresponding to a feature. In an embodiment, the processor 130 may generate the feature classification model by training a neural network using label information matched to each feature. For example, the feature classification model may include a fully connected layer and may be a model that classifies features into at least one of the sleep stages. For example, when a feature corresponding to a first spectrogram is input, the feature classification model may classify the feature into light sleep (e.g., the first sleep stage).

[0462] Meanwhile, according to one embodiment of the present invention, the processor 130 may update the feature extraction model via first learning information related to a first loss between the feature extraction model and the feature classification model according to the results of the learning of the feature extraction model and the feature classification model.

[0463] According to one embodiment, the processor 130 may obtain second training information from the discriminator model, which may be related to adversarial training between the feature extraction model and the discriminator model.

[0464] The processor 130 may perform training via a second loss between the feature extraction model and the discriminator model. The second loss may refer to a loss associated with adversarial training between the feature extraction model and the discriminator model.

[0465] For example, the discriminator model may be trained to output a probability value close to 1 when a first feature corresponding to source data is input, and to output a probability value close to 0 when a second feature corresponding to target data is input. The sum of the difference between the output value when the first feature is input and 1, and the difference between the output value when the second feature is input and 0, may be the loss (or loss function) of the discriminator model. The purpose of the feature extraction model is to deceive the discriminator model (i.e., to make it difficult to distinguish between the first feature and the second feature), and the feature extraction model may be trained to output a value close to 1 when the feature generated by the feature extraction model is input to the discriminator model. The error between the output value and 1 may be the loss of the feature extraction model. That is, each model may be trained by the processor 130 in a direction that minimizes the loss. In other words, the processor 130 may perform training on the adversarial neural network by updating the parameters of the feature extraction model and the discriminator model in a direction that minimizes the adversarial loss.

[0466] In addition, in one embodiment of the present invention, the processor 130 generates a second feature that is as close as possible to the first feature through the feature extraction model, and updates the parameters of each model so that the second feature is more likely to be identified as a feature related to the target data through the discriminator model.

[0467] That is, the processor 130 may update the parameters of the feature extraction model using the second adversarial loss between the feature extraction model and the discriminator model as second learning information, so that the feature extraction model can output similar features (i.e., features located close to each other in the vector space) corresponding to spectrograms related to both domains. In other words, the feature extraction model updated through the second learning information may extract the first feature related to the source data and the second feature related to the target data in the vector space in a manner that is well-mixed and without distinction, regardless of the domain.

[0468] As described above, when the processor 130 receives spectrograms related to the source data and target data as input, the feature extraction model updated through the first learning information and the second learning information may arrange the first and second features adjacent to each other in the vector space, but may also arrange each feature clustered by class.

[0469] According to an embodiment of the present invention, by appropriately positioning a second feature without labeling information based on a first feature with labeling information, the first feature may be classified, and the second feature may also be classified, thereby enabling analysis of sleep acoustic information (i.e., multiple target data) containing various noises.

[0470] According to an embodiment of the present invention, the processor 130 may divide each of the plurality of source data and the plurality of target data into predetermined sample units to generate a plurality of source sub-data and a plurality of target sub-data.

[0471] Additionally, in one embodiment of the present invention, the processor 130 may process frequency domain information, or spectrograms, corresponding to each of the plurality of source sub-data and the plurality of target sub-data as input to a feature extraction model to generate one or more sample features.

[0472] Specifically, instead of processing frequency domain information or spectrograms corresponding to each of the source data and target data as inputs to a feature extraction model, processor 130 may divide each data into sample units and process multiple pieces of frequency domain information or spectrograms corresponding to the divided sample units as inputs to multiple feature extraction models, respectively. In an embodiment, each of the multiple feature extraction models may share parameters. That is, the multiple feature extraction models may be updated to have the same performance.

[0473] The spectrogram corresponding to each sample may be passed through each feature extraction model independently to generate features, and each generated feature may be transmitted to a discriminator model, in which case the discriminator model receives each feature corresponding to each sample unit.

[0474] According to an embodiment of the present invention, raw sleep acoustic information is sequential (or time-series) data over time, which may be a large volume of data. When generating information (e.g., a spectrogram) including changes in frequency components over time or information in the frequency domain (e.g., a spectrogram) based on the sequential data and transmitting the generated spectrogram to a discriminator model, the discriminator model must divide the transmitted spectrogram into epoch units and perform a judgment corresponding to each epoch unit (e.g., whether the feature is related to the source data or the target data), so the learning information (or learning amount) to be learned may be weighted. In other words, if features are extracted corresponding to an entire spectrogram that is not divided into sample units and transmitted to the discriminator model, the learning efficiency of the discriminator model may be reduced.

[0475] Therefore, the processor 130 may divide data (e.g., spectrograms) into sample units, extract features corresponding to each sample, and input the features for each sample to the discriminator model. This allows the discriminator model to be trained with less data through sample-to-sample, and may result in improved overall model performance through efficient training.

[0476] According to an embodiment of the present invention, the processor 130 may refine a feature classification model through decision boundary iterative refinement learning using multiple target data. Refining a feature classification model may refer to using features corresponding to the target data as classification criteria in preference to features corresponding to the source data during the process of classifying features into classes. In other words, it may refer to transforming the decision boundary (i.e., classification boundary) of the features corresponding to the source data into a boundary based on the features corresponding to the target data. The processor 130 may gradually push the decision boundary out of the data density region by minimizing the cluster assumption violation loss on the target side.

[0477] Specifically, the processor 130 may perform iterative decision boundary refinement training using the teacher network. The iterative decision boundary refinement training may be training to improve the placement of the decision boundary based on minimizing the conditional entropy associated with the outputs of the teacher network and the student network, respectively. In a specific embodiment, the processor 130 may input spectrograms corresponding to multiple target data to the student network and the teacher network, respectively. In this case, the student network and the teacher network may each include a feature extraction model and a feature classification model. The processor 130 may train to improve the placement of the decision boundary via the conditional entropy associated with the outputs of the student network and the teacher network.

[0478] Thus, the feature classification model may be refined by modifying the decision boundary so that classification performed based on the decision boundary of the feature corresponding to the source data is performed using the feature corresponding to the target data as a reference for classification. Since the sleep analysis model of the present invention must be equipped to provide analysis information for noisy sounds in the real life of a typical user, as described above, when features corresponding to the target data are used as the basis for the decision boundary, the accuracy of calculating sleep state information may be improved. In other words, through iterative refinement learning of the decision boundary, the feature classification model may output sleep state information with improved accuracy in response to features related to sleep sound information containing various noises.

[0479] According to an embodiment of the present invention, processor 130 may generate a sleep analysis model through the learning model corresponding to the time point at which learning is completed. Specifically, a sleep analysis model may be generated based on the feature extraction model and feature classification model that are trained based on the learning model updated through adversarial learning between the feature extraction model and the feature classification model and adversarial learning between the feature extraction model and the discriminator model. That is, the sleep analysis model may be constructed through the feature extraction model and the feature classification model in the updated learning model, as shown in FIG. 17a or 17b.

[0480] The sleep analysis model according to one embodiment of the present invention is generated through an adaptive learning process, i.e., a training model updated through first learning information and second learning information, and can predict sleep states with improved accuracy even from acoustic data containing various noises.

[0481] In one embodiment of the present invention, the processor 130 may acquire sleep acoustic information from the user terminal 10 and provide sleep state information corresponding to the acquired sleep acoustic information. The processor 130 may generate sleep state information corresponding to the sleep acoustic information by utilizing a sleep analysis model.

[0482] In this case, the sleep analysis model can perform robust prediction even for sound data containing various noises through the adaptive learning described above, which has the advantage of allowing users to easily obtain analytical information about their own sleep state in their own daily environment. In other words, users can receive analytical information about their own sleep state in a general home environment without having to visit a specialized medical institution, install additional special equipment other than sound acquisition equipment, or create a special sleeping environment, which is inexpensive.

[0483] Semi-supervised learning using pseudo labels

[0484] Semi-supervised learning according to an embodiment of the present invention may refer to performing learning of a deep learning model by utilizing data output by a deep learning model to which unlabeled data is input as a pseudo label.

[0485] Unlabeled data may be input into a deep learning model, and information such as prediction and confidence may be output. Here, a higher confidence may mean that the class information contained in the prediction is more reliable.

[0486] In semi-supervised learning according to an embodiment of the present invention, when unlabeled data is input to a deep learning model, a prediction output from the deep learning model is treated as a new label (pseudo label), and learning of sleep state information can be performed based on the new label (pseudo label).

[0487] Semi-supervised learning according to an embodiment of the present invention may perform augmentation preprocessing on images to perform learning. The augmentation preprocessing may be a weakly augmented method or a strongly augmented method.

[0488] According to an embodiment of the present invention, the weakly-augmented method may use one or more augmentation techniques to modulate an image relatively little. The augmentation techniques of the weakly-augmented method may include data augmentation or pitch shifting augmentation (preferably, a technique in which the pitch is shifted in the range of 10% to 20%).

[0489] According to an embodiment of the present invention, the strongly augmented method may use one or more augmentation techniques to modulate an image relatively more. As the augmentation techniques of the strongly augmented method, data augmentation may include one or more of TUT augmentation and noise-added augmentation.

[0490] The learning method according to an embodiment of the present invention may include a method of performing an augmentation preprocessing process on an image and then using the image as an input to a deep learning model for learning. The learning method may also include a method of using weakly-augmented image information as an input to a deep learning model and using output predictions as pseudo labels to perform re-learning based on the pseudo labels. While learning is performed using a weakly-augmented augmentation technique according to an embodiment of the present invention, a moving average technique, a weighted average technique, a weighted moving average technique, an exponential weighted moving average technique, or the like may be used to reflect intermediate learning information in a final learning model.

[0491] According to an embodiment of the present invention, it is possible to obtain highly reliable pseudo labels that can be generated by performing inference using data that utilizes a weakly-augmented augmentation technique as input to a deep learning model.

[0492] In a learning method according to an embodiment of the present invention, pseudo labels obtained by performing unsupervised learning through weakly-augmented data augmentation are used as data labels to perform supervised learning through strongly-augmented data augmentation. By performing supervised learning using pseudo labels, it is possible to learn about more information by using relatively more modulation of an image as input. In addition, it is possible to perform supervised learning on more data without the need for data labeling.

[0493] In addition, the learning method according to an embodiment of the present invention may include a method of performing tuning so that the distribution of predicted values of sleep states (e.g., predicted values of REM sleep stages, predicted values of wake states, predicted values of light sleep stages, predicted values of deep sleep stages, etc.) output as a result of learning using data acquired from a target environment (e.g., an environment other than a sleep polymorphism test) or a target subject population (e.g., a subject population without sleep disorders) as input to a deep learning model is formed to match the distribution of predicted values of sleep states output as a result of learning using data acquired from a specific environment (e.g., a sleep polymorphism test environment) or a comparison subject population (e.g., a subject population with sleep disorders) as input to a deep learning model.

[0494] In this case, the distribution of predicted values of sleep states output using data acquired in a specific environment (e.g., a sleep polymorphism test environment) as input to a deep learning model may not completely match the distribution of predicted values acquired in another environment or target environment (e.g., an environment other than a sleep polymorphism test), but a tuning method may be included so that the modeling data are formed in a direction that matches each other.

[0495] Un / Self-Supervised learning

[0496] According to an embodiment of the present invention, unsupervised learning or self-supervised learning may refer to a method of performing pre-learning to increase the reliability of a predicted value for image information by performing deep learning even when image information in an image domain does not have a label. In this case, it may include a method of performing learning to predict the damaged part of image information after damaging a part of image information.

[0497] A learning method according to an embodiment of the present invention may include a method of first performing pre-learning using unlabeled data as input to a deep learning model, and then performing additional learning using labeled data as input.

[0498] By inputting a smaller amount of labeled data into a deep learning model that has undergone pre-training using the unsupervised learning or self-supervised learning method according to an embodiment of the present invention and training it to perform its intended task, the reliability of the predictions of the pre-trained deep learning model may be further increased.

[0499] Meanwhile, analysis of sleep state information based on acoustic information may include a pattern identification step for sleep sounds, such as breathing and body movements. However, since the characteristics of sleep sound patterns are reflected over time, it may be difficult to fully understand them with just a short snapshot of acoustic data at a specific point in time. Therefore, in order to model acoustic information, analysis must be performed based on the time-series characteristics of the acoustic information, and there has been a demand for applying semi-supervised learning methods to such time-series data.

[0500] Furthermore, acoustic information without a specified label may include environmental sensing information, which is acoustic information acquired by a user through the user terminal 10. However, such environmental sensing information may include data without a specified label (e.g., everyday noise, sounds caused by various people, musical sounds, natural sounds, etc.), and there has been a demand for a new approach to modeling data acquired in an environment where quality control may not be properly performed.

[0501] Therefore, according to one embodiment of the present invention, a semi-supervised learning method based on sequential consistency loss that takes into account the time-series characteristics of sleep acoustic information may be provided.

[0502] In addition, in an embodiment, a semi-supervised contrastive learning (SSCL) method may be provided that processes out-of-distribution (OOD) data from unlabeled information.

[0503] The semi-supervised learning and semi-supervised contrast learning methods according to embodiments of the present invention have been evaluated based on various datasets, including a labeled audio dataset and a PSG audio dataset, and have been found to have the effect of providing robust performance in a real environment, demonstrating the generalizability of the sleep analysis model according to embodiments of the present invention. Hereinafter, semi-supervised learning and semi-supervised contrast learning methods for processing sleep acoustic information in a real environment, including both noise and time series characteristics, will be described in detail with reference to drawings and equations.

[0504] JPEG2025526211000002.jpg41164

[0505] The sleep stage label yi may be represented by four possible classes of sleep stage information: wake stage, REM sleep stage, light sleep stage, and deep sleep stage, and can be expressed as a one-hot label. Meanwhile, according to one embodiment of the present invention, a sequence-to-sequence model can be used, which is composed of a backbone network for extracting low-level features and a head for learning the temporal correlation between mel-spectrograms.

[0506] JPEG2025526211000003.jpg41164

[0507] A learning method based on sequential consistency loss

[0508] JPEG2025526211000004.jpg66164

[0509] JPEG2025526211000005.jpg1526

[0510] On the other hand, consistency loss (LC) can make a sleep analysis model more generalizable for predicting sleep stages for each sample, but it may be difficult to utilize the temporal correlation or time series information between sleep acoustic information and the corresponding labels. Therefore, according to one embodiment of the present invention, sequential consistency loss (LSC) can be used to match the similarity of predicted sequences, as shown in [Equation 2] below. In [Equation 2], ° is the Hadamard power, and ℓ denotes the element-wise product of two matrices averaged by value.

[0511] JPEG2025526211000006.jpg1538

[0512] According to an embodiment of the present invention, in order to predict the degree of sleep stage variation over time, cosine similarity may be employed between the logits of the i-th and j-th samples in sequence. The cosine similarity is calculated by dividing the value of the dot product of two vectors by the product of the magnitudes of the two vectors, and may indicate the degree of similarity or relationship between the vectors measured using the cosine value of the angle between the two vectors in the dot product space.

[0513] JPEG2025526211000007.jpg34164

[0514] JPEG2025526211000008.jpg2089

[0515] In [Number 3], w min means the value of the smallest weight of the most distant pair. Thus, according to one embodiment of the present invention, loss may be constrained so that predictions for two different augmentation results from the same sequence have similar sequence trends.

[0516] 30 is a diagram illustrating an example of a training method based on consistency loss or sequential consistency loss when the number of samples in a sequence is 6, according to an embodiment of the present invention. In FIG. 30, for the two consistency losses, C a and the upper triangular matrices of W are displayed.

[0517] JPEG2025526211000009.jpg58164

[0518] JPEG2025526211000010.jpg58164

[0519] JPEG2025526211000011.jpg26163

[0520] Meanwhile, the upper triangular matrix represented as W in FIG. 30 indicates weights based on importance. Since the further apart samples are in the same sequence, the less relevant they may be, the smaller the weights can be set. In one embodiment of the present invention, weights ranging from 0 to 1 are set as represented in the upper triangular matrix of W, but this is merely an example and the present invention is not limited thereto. For example, the greater the number of samples included in the sequence, the smaller the difference between the weights can be set. Furthermore, the minimum weight value can be set to a value other than 0.5, and the difference between the weights may not necessarily be constant.

[0521] 30 is merely an example, and the present invention is not limited thereto. For example, the number of samples in the sequence may be 40 samples or 14 samples. This number may vary depending on what sleep state information the sleep analysis model performs prediction on, and is not absolute.

[0522] Semi-supervised Contrastive Loss-based learning method

[0523] Hereinafter, a semi-supervised contrastive learning method according to an embodiment of the present invention will be described in detail using equations, drawings, etc. According to an embodiment of the present invention, a class-aware contrastive semi-supervised learning (CCSSL) method may be adopted to fully utilize unlabeled data that may include out-of-distribution (OOD) samples.

[0524] Meanwhile, according to an embodiment of the present invention, pushing and pulling can be performed not only between unlabeled data (data without a designated label) but also between labeled data. When performing such a CCSSL method, learning can be performed to ensure that data determined to be out-of-distribution (OOD) is pushed. Meanwhile, according to an embodiment, when labeled data is included in the process of performing pushing and pulling between data, the labeled data can act as anchors that remain stationary, and only the unlabeled data can move. In this way, when performing the CCSSL method using labeled data, data determined to belong to in-distribution (OOD) can be more effectively pulled and data determined to be out-of-distribution (OOD) can be more effectively pushed, thereby further improving the learning performance of the sleep analysis model.

[0525] FIG. 31 is an exemplary diagram illustrating the operation mechanism of a learning method based on semi-supervised contrastive loss according to an embodiment of the present invention.

[0526] 31, x1 can serve as an anchor as information labeled as a deep sleep stage. In the semi-supervised contrastive learning method according to an embodiment of the present invention, u1 can be selected as the anchor x1 because u1 has a sufficiently high reliability for the deep sleep stage class due to a clear and regular breathing pattern. In this case, a certain threshold can be regarded as a critical value in the in-distribution of the pseudo label, and it can be determined that u1 should be selected when the reliability exceeds the critical value.

[0527] On the other hand, since u3 has no evidence in the deep sleep stage class, it may be data from another class or out-of-distribution (OOD), and u3 can be pushed out from anchor x1. In this case, a certain threshold is regarded as the in-distribution threshold of the pseudo label, and if the threshold cannot be exceeded, it can be determined to be pushed out. In other words, data determined to belong to out-of-distribution (OOD) through this method can be pushed out.

[0528] On the other hand, if the similarity with the anchor is judged to be neither high nor low, as in the case of u2, there is a risk of pushing data belonging to the same class or pulling data belonging to a different class, so pushing and pulling may not be performed.

[0529] The CCSSL method utilizes labeled data or pseudo labels as above to calculate supervised contrastive loss between unlabeled data samples and can eliminate OOD samples with clusters that represent features within a class.

[0530] JPEG2025526211000012.jpg51164

[0531] JPEG2025526211000013.jpg2167

[0532] JPEG2025526211000014.jpg41164

[0533] JPEG2025526211000015.jpg26163

[0534] JPEG2025526211000016.jpg90164

[0535] JPEG2025526211000017.jpg57164

[0536] As can be seen from Equation 4, for samples that are not labeled according to the above method but are assigned pseudo labels, learning can be performed in a subtractive manner.

[0537] On the other hand, in the presence of heavily contaminated unlabeled data, CCSSL may be unreliable because OOD samples in the unlabeled data may be sampled with high confidence, causing confusion in class clustering. Therefore, according to one embodiment of the present invention, to solve this problem, a semi-supervised contrastive learning (SSCL) method may be provided that utilizes reliable labeled data as anchor points for class clustering.

[0538] In a semi-supervised contrast learning (SSCL) method according to an embodiment of the present invention, reliable positive and negative samples can be used by considering labeled samples as anchors. Here, a positive sample refers to a sample that exceeds a threshold, known as in-distribution data, and corresponds to a sample determined to be in the same class. Positive samples may be trained by being attracted to labeled samples. Negative samples refer to samples that do not exceed a threshold, known as in-distribution data, and correspond to a sample determined to be in a different class from the labeled samples. Negative samples may be trained by being pushed by labeled samples. Thus, the SSCL method according to an embodiment of the present invention may train by pushing OOD samples out of an embedding cluster within a class. However, if it is not clear whether a sample's class is positive or negative relative to the anchor, push-pull training may not be performed.

[0539] JPEG2025526211000018.jpg29164

[0540] JPEG2025526211000019.jpg2185

[0541] JPEG2025526211000020.jpg47164

[0542] JPEG2025526211000021.jpg1537

[0543] According to one embodiment of the present invention, SSCL uses labeled embedding z mThis is because the goal of SSCL is not to train the properties of labeled samples, but to push OOD samples in unlabeled sequences farther away from labeled samples.

[0544] JPEG2025526211000022.jpg15163

[0545] JPEG2025526211000023.jpg1387

[0546] Performance of a sleep analysis model using semi-supervised learning

[0547] A sleep analysis model according to one embodiment of the present invention was learned or trained on labeled data from approximately 3,000 laboratory-based PSG studies and approximately 3,000 self-collected, unlabeled home data.

[0548] Additionally, a sleep analysis model according to an embodiment of the present invention was evaluated based on laboratory PSG, home PSG, and PSG audio data. Here, the generalization ability of the sleep analysis model was tested by evaluating its performance in comparison with PSG audio data (PSG-Auido), an open dataset primarily composed of apnea patient data.

[0549] Specifically, to evaluate the results of a sleep analysis model according to one embodiment of the present invention, the NS value, which is the number of samples in a sequence, was set to 40, and the evaluation was performed in the same manner, taking into consideration that when a sleep analysis technician assigns a label every 30 seconds in a PSG test performed in a hospital environment, they usually check for a margin of ±10 minutes.

[0550] JPEG2025526211000024.jpg41164

[0551] FIG. 32 is a table comparing the analysis results of a sleep analysis model according to an embodiment of the present invention with the analysis results of a PSG test in a home environment. The SoundSleepNet row in FIG. 32 represents the results of applying a sequence-to-sequence sleep analysis model, which is composed of a backbone network for low-level feature extraction and a head for learning the time-series correlation between melspectrograms, to a PSG test in a home environment according to an embodiment of the present invention. The SleepFormer row in FIG. 32 represents a sleep analysis model reflecting unsupervised learning and / or semi-supervised learning methods according to an embodiment of the present invention. Meanwhile, C, SC, CC, SS, and WA represent consistency, sequential consistency, CCSSL, SSCL, and weighted average, respectively.

[0552] 33 is a table comparing sleep analysis results based on PSG audio data with analysis results of a sleep analysis model according to an embodiment of the present invention. The upper table of FIG. 33 is a table comparing sleep analysis results based on PSG audio, and the lower table of FIG. 33 is a table comparing sleep analysis results based on PSG data in a laboratory environment.

[0553] As shown in the comparison table of analysis results in Figure 32, the semi-supervised learning method according to an embodiment of the present invention was applied one by one to evaluate the change in performance. The SleepFormer model, which is the supervised baseline, was constructed using a Transformer-based artificial intelligence model, and it was confirmed that the F1 score was 0.6332, which was an improvement of 0.0614 compared to SoundSleepNet. Furthermore, it was confirmed that the addition of consistency loss (C) and sequential consistency loss (SC) resulted in a much greater improvement in the F1 score, reaching 0.6597 and 0.6751, respectively.

[0554] In addition, by introducing SS (SSCL) according to one embodiment of the present invention, the F1 score improved significantly to 0.6780, and when the weights of three models trained with different seeds were averaged and compared (WA), the final score was 0.6804, an improvement of 0.1085.

[0555] Meanwhile, sleep analysis results based on PSG audio data were compared with analysis results of a sleep analysis model according to an embodiment of the present invention, as shown in the table in Figure 33. The first row, "Supervised," in each table in Figure 33 is the result of predicting sleep state information by inputting PSG audio data into a sleep analysis model according to an embodiment of the present invention, and the second row, "Ours," is the result of predicting sleep state information by inputting PSG audio data into a sleep analysis model based on semi-supervised learning according to an embodiment of the present invention.

[0556] 33, the PSG-Audio dataset that serves as the input for the sleep analysis model is a data distribution that the sleep analysis model has not encountered during training, i.e., has not been exposed to, and such a dataset may be used to evaluate the generalization performance of the sleep analysis model for new data and unexposed data. On the other hand, the PSG-Audio dataset is mainly composed of severe apnea patients, and the distribution of sleep stage classes may be unbalanced.

[0557] Furthermore, the PSG dataset in a laboratory environment that serves as input to the sleep analysis model in the table shown in the lower part of Figure 33 may be representative of the distribution of labeled sources.

[0558] 33, when the PSG-Audio dataset was input to the Supervised model and the Semi-supervised model and the accuracy of the prediction results was compared, it was confirmed that the accuracy of the Semi-supervised model improved by 0.0437. On the other hand, the accuracy improvement for the PSG dataset in a laboratory environment was relatively small. This is because the supervised baseline according to an embodiment of the present invention already achieved good performance of 0.7000 points on labeled source distribution data.

[0559] Such semi-supervised learning methods may further improve sleep analysis models by processing time-series sleep acoustic information in a real sleep environment. Briefly, sequential consistency loss improves the temporal correlation of the sleep analysis model, and semi-supervised contrastive loss improves feature representation clusters with labeled samples, effectively removing out-of-distribution (OOD) samples, thereby improving accuracy. Furthermore, it has been confirmed that sleep analysis models according to embodiments of the present invention can show consistent improvement effects across all data sets, including data from a home environment, unexposed data, and labeled datasets.

[0560] A method for analyzing multimodal sleep state information

[0561] One embodiment of a multimodal sleep state information analysis method (CONCEPT-A)

[0562] FIG. 26 is a flowchart illustrating a method for analyzing sleep state information including a process of combining sleep acoustic information and sleep environment information as multimodal data according to an embodiment of the present invention.

[0563] In order to achieve the object of the present invention, according to one embodiment, a method for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information may include a first information acquisition step (S100) of acquiring time-domain acoustic information related to a user's sleep, a step (S102) of performing preprocessing of the first information, a second information acquisition step (S110) of acquiring user sleep environment information related to the user's sleep, a step (S112) of performing preprocessing of the second information, a combination step (S120) of combining data multimodally, a step (S130) of inputting the multimodal data to a deep learning model, and a step (S140) of acquiring sleep state information as an output of the deep learning model.

[0564] According to an embodiment of the present invention, the first information acquisition step (S100) may acquire time-domain acoustic information related to the user's sleep in the user terminal 10. The time-domain acoustic information related to the user's sleep may include sound source information obtained by a sound source detection unit of the user terminal 10.

[0565] According to an embodiment of the present invention, in the step of performing data preprocessing of the first information (S102), the time-domain sleep sound information may be converted into information including changes in frequency components along a time axis or into frequency-domain information. The frequency-domain information may be expressed as a spectrogram, which may be a mel-spectrogram to which a mel scale is applied. Converting the information into a spectrogram can protect the user's privacy and reduce the amount of data processing. The converted time-domain sleep sound information is visualized, and in this case, it can be used as an input for an image processing-based artificial intelligence model to obtain sleep state information through image analysis.

[0566] According to an embodiment of the present invention, the step of performing data pre-processing of the first information (S102) may further include a step of extracting features based on the acoustic information. For example, a user's sleep breathing pattern may be extracted based on the acquired time-domain acoustic information. For example, the acquired time-domain acoustic information may be converted into information including changes in frequency components over time, and the user's sleep breathing pattern may be extracted based on the converted information. Alternatively, the time-domain acoustic information may be converted into frequency-domain information, and the user's sleep breathing pattern may be extracted based on the frequency-domain acoustic information.

[0567] In this case, the converted information is visualized, and information such as the user's breathing pattern may be output as input for an image processing-based artificial intelligence model. According to an embodiment of the present invention, the step of performing data pre-processing of the first information (S102) may include a data augmentation process for obtaining a sufficient amount of meaningful data for inputting the sleep acoustic information into the deep learning model. Data augmentation techniques may include pitch shifting augmentation, Tile UnTile (TUT) augmentation, and noise-added augmentation. The above-mentioned augmentation techniques are merely examples, and the present invention is not limited thereto.

[0568] According to an embodiment of the present invention, the mel-scale addition method can reduce the time it takes for the hardware to process the data.

[0569] Meanwhile, the specific description of the types of noise mentioned above is merely an example for explaining the noise-added augmentation of the present invention, and the present invention is not limited thereto.

[0570] According to an embodiment of the present invention, the second information acquiring step (S110) of acquiring user sleep environment information related to the user's sleep may acquire the user sleep environment information via the user terminal 10, an external server, or a network. The user sleep environment information may refer to information related to sleep acquired in the space where the user is located. The sleep environment information may be sensing information acquired in the space where the user is located using a non-contact method. The sleep environment information may be respiratory movement and body movement information measured using radar. The sleep environment information may be information related to the user's sleep acquired using a smart watch, smart home appliance, etc. The sleep environment information may be photoplethysmography (Photoplethysmography). The sleep environment information may be heart rate variability (HRV) and heart rate obtained through photoplethysmography (PPG), and the photoplethysmography may be measured by a smart watch or a smart ring. The sleep environment information may be electroencephalography (EEG) signals. The sleep environment information may be actigraphy signals measured during sleep.

[0571] According to an embodiment of the present invention, the step of pre-processing the second information (S112) may include a data augmentation process to obtain a sufficient amount of meaningful data for inputting the user's sleep environment information data into the deep learning model.

[0572] According to an embodiment of the present invention, the step of pre-processing the second information (S112) may include a step of processing data of the user's sleep environment information to extract features. For example, if the second information is a photoplethysmogram (PPG), a heart rate variability (HRV) and a heart rate (HR) may be extracted from the photoplethysmogram.

[0573] According to an embodiment of the present invention, the step of pre-processing the second information (S112) may include TUT (Tile UnTile) augmentation and noise-added augmentation of the image information when the user's sleep environment information data is obtained as image information. The above-described augmentation techniques are merely examples of augmentation techniques for image information, and the present invention is not limited thereto. The user's sleep environment information may be information stored in various formats. Various methods may be adopted for augmenting the user's sleep environment information.

[0574] According to an embodiment of the present invention, the step of combining the first information and the second information that have undergone the data preprocessing process as multimodal data (S120) combines the data to input the multimodal data into a deep learning model.

[0575] According to an embodiment of the present invention, a method for combining multimodal data may combine preprocessed first information and preprocessed second information in the same data format. Specifically, the first information may be acoustic image information in the frequency domain, and the second information may be heartbeat image information in the time domain obtained by a smartwatch. In this case, since the first information and the second information are not in the same domain, they can be converted into the same domain and combined.

[0576] According to an embodiment of the present invention, a method for combining multimodal data can also combine preprocessed first information and preprocessed second information in the same data format. Specifically, the first information may be acoustic image information in the frequency domain, and the second information may be heartbeat image information in the time domain obtained by a smartwatch. In this case, since the first information and the second information are not in the same domain, each piece of data can be labeled as relating to the first information and the second information for use as input for a deep learning model.

[0577] According to an embodiment of the present invention, the step of combining using multimodal data (S120) may involve augmenting first information and then augmenting second information before combining. For example, the first information may be time-domain acoustic information of a user, and the second information may be photoplethysmogram (PPG), which may be combined using multimodal data. For example, the first information may be time-domain acoustic information of a user or a spectrogram obtained by converting time-domain acoustic information into frequency-domain acoustic information, and the second information may be photoplethysmogram (PPG), which may be combined using multimodal data.

[0578] According to one embodiment of the present invention, the step of combining multimodal data (S120) may involve augmenting first information, augmenting second information, and extracting features to combine the data. For example, the first information may be a user's time-domain acoustic information or a spectrogram obtained by converting time-domain acoustic information into frequency-domain acoustic information, and the second information may be heart rate variability (HRV) or heart rate obtained by photoplethysmography (PPG), and these may be combined as multimodal data. According to one embodiment of the present invention, the step of combining multimodal data (S120) may involve augmenting first information and extracting features, and augmenting second information to combine the data. For example, the first information may be a user's breathing pattern extracted based on the user's acoustic information, and the second information may be heart rate variability (HRV) or heart rate obtained by photoplethysmography (PPG), and these may be combined as multimodal data.

[0579] According to an embodiment of the present invention, the step of combining using multimodal data (S120) may involve augmenting first information and extracting features, and then augmenting second information and extracting features to combine them. For example, the first information may be a user's breathing pattern extracted based on the user's acoustic information, and the second information may be heart rate variability (HRV) or heart rate obtained by photoplethysmography (PPG), which may be combined using multimodal data.

[0580] According to an embodiment of the present invention, the step of inputting multimodal combined data into a deep learning model (S130) may process the data into a consistent form required for inputting the multimodal combined data into the deep learning model.

[0581] According to an embodiment of the present invention, in the step of acquiring sleep state information as an output of a deep learning model (S140), the sleep state information may be inferred by using the multimodal combined data as an input of a deep learning model for inferring sleep state information. The sleep state information may be information regarding a sleep state of a user.

[0582] According to an embodiment of the present invention, the user's sleep state information may include sleep stage information that expresses the user's sleep as stages. Sleep stages may be classified into NREM (non-REM) sleep and REM (rapid eye movement) sleep, and NREM sleep may be further classified into multiple stages (e.g., two stages: light and deep, and four stages: N1 to N4). The sleep stages may be defined as general sleep stages, or may be arbitrarily set as various sleep stages by a designer.

[0583] According to an embodiment of the present invention, the user's sleep state information may include sleep event information representing sleep-related disorders and sleep behaviors occurring during the user's sleep. Specifically, the sleep event information occurring during the user's sleep may include information on sleep apnea and hypopnea due to the user's sleep disorders. More specifically, the sleep event information occurring during the user's sleep may include whether the user snores, the duration of the snoring, whether the user talks in their sleep, the duration of the sleep talking, whether the user turns over in their sleep, and the duration of the turning over in their sleep. The user's sleep event information described above is merely an example representing events occurring during the user's sleep, and is not limited thereto.

[0584] One embodiment of a multimodal sleep state information analysis method (CONCEPT-B)

[0585] FIG. 27 is a flowchart illustrating a method for analyzing sleep state information, including a step of combining inferred sleep acoustic information and inferred sleep environment information as multimodal data, according to an embodiment of the present invention.

[0586] In order to achieve the object of the present invention, according to one embodiment, a method for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information may include a first information acquisition step (S200) of acquiring time-domain acoustic information related to a user's sleep, a step (S202) of performing preprocessing of the first information, a step (S204) of inferring information related to sleep using the first information as an input of a deep learning model, a second information acquisition step (S210) of acquiring user sleep environment information related to the user's sleep, a step (S212) of performing preprocessing of the second information, a step (S214) of inferring information related to sleep using the second information as an input of a deep learning model, a combination step (S220) of combining data multimodally, and a step (S230) of acquiring sleep state information by combining multimodal data.

[0587] According to an embodiment of the present invention, the first information acquisition step (S200) may acquire time-domain acoustic information related to the user's sleep in the user terminal 10. The time-domain acoustic information related to the user's sleep may include sound source information obtained by a sound source detection unit of the user terminal 10.

[0588] According to an embodiment of the present invention, in the step of performing data pre-processing of the first information (S202), time-domain acoustic information may be converted into information including changes in frequency components along the time axis or into information in the frequency domain. Furthermore, the information in the frequency domain may be expressed as a spectrogram, and may be a mel spectrogram to which a mel scale is applied. Converting the information into a spectrogram can protect user privacy and reduce the amount of data processing.

[0589] According to an embodiment of the present invention, the step of performing data pre-processing of the first information (S202) may include a data augmentation process to obtain a sufficient amount of meaningful data for inputting the sleep acoustic information into a deep learning model. Data augmentation techniques may include pitch shifting augmentation, Tile UnTile (TUT) augmentation, and noise-added augmentation. The above-mentioned augmentation techniques are merely examples, and the present invention is not limited thereto.

[0590] According to an embodiment of the present invention, the mel-scale addition method can reduce the time it takes for the hardware to process the data.

[0591] Meanwhile, the specific description of the types of noise mentioned above is merely an example for explaining the noise-added augmentation of the present invention, and the present invention is not limited thereto.

[0592] According to an embodiment of the present invention, the second information acquiring step (S210) of acquiring user sleep environment information related to the user's sleep may acquire the user sleep environment information via the user terminal 10, an external server, or a network. The user sleep environment information may refer to information related to sleep acquired in the space where the user is located. The sleep environment information may be sensing information acquired in the space where the user is located using a non-contact method. The sleep environment information may be respiratory movement and body movement information measured using radar. The sleep environment information may be information related to the user's sleep acquired using a smart watch, smart home appliance, etc. The sleep environment information may be heart rate variability (HRV) and heart rate obtained through photoplethysmography (PPG), which may be measured using a smart watch or smart ring. The sleep environment information may be electroencephalography (EEG) signals. The sleep environment information may be actigraphy signals measured during sleep.

[0593] According to an embodiment of the present invention, the step of pre-processing the second information (S212) may include a data augmentation process to obtain a sufficient amount of meaningful data for inputting the user's sleep environment information data into the deep learning model.

[0594] According to an embodiment of the present invention, in the step S212 of pre-processing the second information, if the user's sleep environment information data is obtained as image information, the image information may include Tile UnTile (TUT) augmentation and noise-added augmentation. The above-described augmentation techniques are merely examples of image information augmentation techniques, and the present invention is not limited thereto. The user's sleep environment information may be information stored in various formats. Various methods may be adopted for augmenting the user's sleep environment information.

[0595] According to an embodiment of the present invention, the step of inferring information about sleep using the preprocessed first information as an input of a deep learning model (S204) can infer information about sleep using the preprocessed first information as an input of an already trained deep learning model.

[0596] Embodiments of the present invention allow an already trained deep learning model to use inferred data as input for self-training via inferred data.

[0597] According to an embodiment of the present invention, a deep learning sleep analysis model that takes first information about sleep acoustics as input to infer information about sleep may include a feature extraction model and a feature classification model.

[0598] The feature extraction model of the deep learning sleep analysis model according to an embodiment of the present invention may be pre-trained using a one-to-one proxy task in which one spectrogram is input and the model is trained to predict sleep state information corresponding to the spectrogram. When a CNN deep learning model is used as the feature extraction model according to an embodiment of the present invention, a fully connected layer (FC) or fully connected neural network (FCN) structure may be adopted for training. When a MobileViTV2 deep learning model is used as the feature extraction model according to an embodiment of the present invention, an intermediate layer structure may be adopted for training.

[0599] Among the deep learning sleep analysis models according to embodiments of the present invention, the feature classification model can be trained to receive multiple consecutive spectrograms, predict sleep state information for each spectrogram, and analyze a sequence of multiple consecutive spectrograms to predict or classify overall sleep state information.

[0600] According to an embodiment of the present invention, in the step S214 of inferring information about sleep using the preprocessed second information as an input of an inference model, information about sleep may be inferred using the preprocessed second information as an input of a pre-trained inference model. The pre-trained inference model may be, but is not limited to, the above-described deep learning sleep analysis model, and may be an inference model of various types to achieve a purpose. The pre-trained inference model may adopt various methods.

[0601] According to an embodiment of the present invention, the step of combining the first information and the second information that have undergone the data pre-processing process as multimodal data (S220) combines the information to determine sleep state information.

[0602] According to one embodiment of the present invention, a method for combining multimodal data may combine sleep information inferred via pre-processed first information and information inferred via pre-processed second information in the same format of data.

[0603] According to an embodiment of the present invention, in the step of acquiring sleep state information by combining multimodal data (S230), data acquired in a multimodal manner may be combined to determine sleep state information of the user. The sleep state information may be information regarding the sleep state of the user.

[0604] According to an embodiment of the present invention, the step of acquiring sleep state information by multimodal data combination (S230) may combine a hypnogram related to the user's sleep inferred in the step of inferring information about sleep using the preprocessed first information as an input to a deep learning model (S204) with a hypnogram related to the user's sleep inferred in the step of inferring information about sleep using the preprocessed second information as an input to an inference model (S214). For example, the sleep state information may be acquired by overlapping the hypnograms and using information about sleep stages for matching parts and assigning a weight to information about sleep stages for non-matching parts to determine whether to use them.

[0605] According to an embodiment of the present invention, the step of acquiring sleep state information through multimodal data combination (S230) may combine a hypnodensity graph related to the user's sleep inferred in the step of inferring sleep information using the preprocessed first information as an input to a deep learning model (S204) with a hypnodensity graph related to the user's sleep inferred in the step of inferring sleep information using the preprocessed second information as an input to an inference model (S214). For example, the probability of each hypnodensity graph may be substituted into a formula to obtain the sleep stage with the highest reliability for each time as the user's sleep stage information. For example, if the reliability of each hypnodensity graph by time exceeds a predetermined reliability threshold, it may be adopted as the user's sleep stage information. If there is no sleep stage information whose reliability by time exceeds the predetermined reliability threshold, it may be adopted as the sleep stage information using a weighted value to obtain sleep state information.

[0606] According to an embodiment of the present invention, the step of acquiring sleep state information through multimodal data combination (S230) may combine a hypnogram related to the user's sleep inferred in the step of inferring sleep information using the preprocessed first information as an input to a deep learning model (S204) with a hypnodensity graph related to the user's sleep inferred in the step of inferring sleep information using the preprocessed second information as an input to an inference model (S214). For example, if the reliability of the sleep stages displayed on the hypnogram and the hypnodensity graph exceeds a predetermined threshold, the reliability is adopted as the user's sleep stages, thereby acquiring the user's sleep state information. For example, if the reliability of the sleep stages displayed on the hypnogram and the hypnodensity graph does not exceed a predetermined threshold, a weighted value is added and calculated, and the reliability is adopted as the user's sleep stages, thereby acquiring the user's sleep state information with high reliability.

[0607] According to an embodiment of the present invention, the user's sleep state information may include sleep stage information that expresses the user's sleep as stages. Sleep stages may be classified into NREM (non-REM) sleep and REM (rapid eye movement) sleep, and NREM sleep may be further classified into multiple stages (e.g., two stages: light and deep, and four stages: N1 to N4). The sleep stages may be defined as general sleep stages, or may be arbitrarily set as various sleep stages by a designer.

[0608] According to an embodiment of the present invention, the user's sleep state information may include sleep stage information that displays the user's sleep as stages. Methods for displaying the sleep stages may include, but are not limited to, a hypnogram that displays the sleep stages in a graph and a hypnodesity graph that displays the probability of each sleep stage in a graph.

[0609] According to an embodiment of the present invention, the user's sleep state information may include sleep event information representing sleep-related disorders and sleep behaviors occurring during the user's sleep. Specifically, the sleep event information occurring during the user's sleep may include information on sleep apnea and hypopnea due to the user's sleep disorders. More specifically, the sleep event information occurring during the user's sleep may include whether the user snores, the duration of the snoring, whether the user talks in their sleep, the duration of the sleep talking, whether the user turns over in their sleep, and the duration of the turning over in their sleep. The user's sleep event information described above is merely an example representing events occurring during the user's sleep, and is not limited thereto.

[0610] One embodiment of a multimodal sleep state information analysis method (CONCEPT-C)

[0611] FIG. 28 is a flowchart illustrating a method for analyzing sleep state information including combining inferred sleep acoustic information with sleep environment information as multimodal data according to an embodiment of the present invention.

[0612] In order to achieve the object of the present invention, according to one embodiment, a method for analyzing sleep state information using multimodal sleep acoustic information and sleep environment information may include a first information acquisition step (S300) of acquiring time-domain acoustic information related to a user's sleep, a step (S302) of performing preprocessing of the first information, a step (S304) of inferring information related to sleep using the first information as input to a deep learning model, a second information acquisition step (S310) of acquiring user sleep environment information related to the user's sleep, a combination step (S320) of combining data multimodally, and a step (S330) of acquiring sleep state information by combining the multimodal data.

[0613] According to an embodiment of the present invention, the first information acquisition step (S300) may acquire time-domain acoustic information related to the user's sleep in the user terminal 10. The time-domain acoustic information related to the user's sleep may include sound source information obtained by a sound source detection unit of the user terminal 10.

[0614] According to an embodiment of the present invention, in the step of performing data pre-processing of the first information (S302), time-domain acoustic information may be converted into frequency-domain information. Furthermore, the frequency-domain information may be expressed as a spectrogram, which may be a mel-spectrogram to which a mel scale is applied. Converting the information into a spectrogram can protect user privacy and reduce the amount of data processing.

[0615] According to an embodiment of the present invention, the step of performing data pre-processing of the first information (S302) may include a data augmentation process for obtaining a sufficient amount of meaningful data for inputting the sleep acoustic information into a deep learning model. Data augmentation techniques may include pitch shifting augmentation, Tile UnTile (TUT) augmentation, and noise-added augmentation. The above-mentioned augmentation techniques are merely examples, and the present invention is not limited thereto.

[0616] According to an embodiment of the present invention, the mel-scale addition method can reduce the time it takes for the hardware to process the data.

[0617] Meanwhile, the specific description of the types of noise mentioned above is merely an example for explaining the noise-added augmentation of the present invention, and the present invention is not limited thereto.

[0618] According to an embodiment of the present invention, the second information acquiring step (S310) of acquiring user sleep environment information related to the user's sleep may acquire the user sleep environment information via the user terminal 10, an external server, or a network. The user sleep environment information may refer to information related to sleep acquired in the space where the user is located. The sleep environment information may be sensing information acquired in the space where the user is located using a non-contact method. The sleep environment information may be respiratory movement and body movement information measured using radar. The sleep environment information may be information related to the user's sleep acquired using a smart watch, smart home appliance, etc. The sleep environment information may be heart rate variability (HRV) and heart rate obtained through photoplethysmography (PPG), which may be measured by a smart watch or smart ring. The sleep environment information may be electroencephalography (EEG) signals. The sleep environment information may be actigraphy signals measured during sleep. The sleep environment information may be labeling data representing information about the user. Specifically, the labeling data may include the user's age, presence or absence of illness, physical condition, race, height, weight, and body mass index. These are merely examples of labeling data representing information about the user, and are not limited thereto. The above-mentioned sleep environment information is merely an example of information that may affect the user's sleep, and is not limited thereto.

[0619] According to an embodiment of the present invention, the step of inferring information about sleep using the preprocessed first information as an input of a deep learning model (S304) can infer information about sleep using the preprocessed first information as an input of an already trained deep learning model.

[0620] According to an embodiment of the present invention, a deep learning sleep analysis model that takes first information about sleep acoustics as input to infer information about sleep may include a feature extraction model and a feature classification model.

[0621] The feature extraction model of the deep learning sleep analysis model according to an embodiment of the present invention may be pre-trained using a one-to-one proxy task in which one spectrogram is input and the model is trained to predict sleep state information corresponding to the spectrogram. When a CNN deep learning model is used as the feature extraction model according to an embodiment of the present invention, a fully connected layer (FC) or fully connected neural network (FCN) structure may be adopted for training. When a MobileViTV2 deep learning model is used as the feature extraction model according to an embodiment of the present invention, an intermediate layer structure may be adopted for training.

[0622] A feature classification model among deep learning sleep analysis models according to embodiments of the present invention may be trained to receive a plurality of consecutive spectrograms, predict sleep state information for each spectrogram, and analyze a sequence of the plurality of consecutive spectrograms to predict or classify time-series sleep state information.

[0623] According to an embodiment of the present invention, the step of combining the first information and the second information that have undergone the data preprocessing process as multimodal data (S320) combines the data to input the multimodal data into a deep learning model.

[0624] According to one embodiment of the present invention, the method for combining multimodal data can also combine sleep information inferred via preprocessed first information and information inferred via preprocessed second information in the same data format.

[0625] According to an embodiment of the present invention, the step of acquiring sleep state information by combining multimodal data (S330) may determine sleep state information of a user by combining data acquired multimodally. The sleep state information may be information about a sleep state of a user.

[0626] According to an embodiment of the present invention, the user's sleep state information may include sleep stage information that expresses the user's sleep as stages. Sleep stages may be classified into NREM (non-REM) sleep and REM (rapid eye movement) sleep, and NREM sleep may be further classified into multiple stages (e.g., two stages: light and deep, and four stages: N1 to N4). The sleep stages may be defined as general sleep stages, or may be arbitrarily set as various sleep stages by a designer.

[0627] According to an embodiment of the present invention, the user's sleep state information may include sleep event information representing sleep-related disorders and sleep behaviors occurring during the user's sleep. Specifically, the sleep event information occurring during the user's sleep may include information on sleep apnea and hypopnea due to the user's sleep disorders. More specifically, the sleep event information occurring during the user's sleep may include whether the user snores, the duration of the snoring, whether the user talks in their sleep, the duration of the sleep talking, whether the user turns over in their sleep, and the duration of the turning over in their sleep. The user's sleep event information described above is merely an example representing events occurring during the user's sleep, and is not limited thereto.

[0628] Real-time sleep analysis according to embodiments of the present invention

[0629] Analysis of sleep state information based on acoustic information according to an embodiment of the present invention may include a detection step for sleep events (e.g., apnea, hypopnea, snoring, sleep talking, etc.). However, since the characteristics of sleep acoustic patterns are reflected over time, they may be difficult to understand using only short acoustic data at a specific point in time. Therefore, in order to model the acoustic information, analysis must be performed based on the time-series characteristics of the acoustic information.

[0630] In addition, sleep events that occur during sleep (e.g., apnea, hypopnea, snoring, sleep talking, etc.) have various characteristics associated with them. For example, there may be no sound during an apnea event, but after the apnea event ends, air may pass through again, causing a loud noise. Therefore, sleep events can be detected by learning the characteristics of apnea events in a chronological order.

[0631] Differences in Deep Neural Networks for Real-Time Sleep Event Detection

[0632] According to an embodiment of the present invention, the deep neural network structure for analyzing sleep stages described above can be modified and used to detect sleep events occurring during sleep. Specifically, sleep stage analysis requires time-series learning of sleep sounds. However, since sleep event detection typically occurs between 10 and 60 seconds, accurate detection of one or two epochs, each 30 seconds long, is sufficient. Therefore, the deep neural network structure for analyzing sleep stages according to an embodiment of the present invention can reduce the number of inputs and outputs of the deep neural network structure for analyzing sleep stages. For example, if a deep neural network structure for analyzing sleep stages processes 40 mel spectrograms and outputs sleep stages for 20 epochs, a deep neural network structure for detecting sleep events can process 14 mel spectrograms and output sleep event labels for 10 epochs. Here, sleep event labels may include, but are not limited to, no event, apnea, hypopnea, snoring, and tossing and turning.

[0633] Furthermore, a deep neural network architecture for detecting sleep events occurring during sleep according to an embodiment of the present invention may include a feature extraction model and a feature classification model. Specifically, the feature extraction model extracts features of sleep events found in each mel-spectrogram, and the feature classification model detects multiple epochs, finds epochs containing sleep events, and analyzes adjacent features to predict and classify the types of sleep events in a time series.

[0634] Class Weights for Real-Time Sleep Event Detection

[0635] According to an embodiment of the present invention, a method for detecting sleep events occurring during sleep may assign class weights to each sleep event to solve the class imbalance problem. Specifically, among sleep events occurring during sleep, "no event" may have a dominant influence on the overall sleep length, which may reduce the efficiency of sleep event learning. Therefore, by assigning a weight higher than "no event" to other sleep events, learning efficiency and accuracy may be improved. For example, if sleep event classes are classified into three, "no event," "apnea," and "hypopnea," a weight of 1.0 may be assigned to "no event," a weight of 1.3 may be assigned to "apnea," and a weight of 2.1 may be assigned to "hypopnea" to reduce the influence of "no event" on learning.

[0636] Consistency Training for Real-Time Sleep Event Detection

[0637] 21 is a diagram illustrating consistency training according to an embodiment of the present invention. In order to detect sleep events occurring during sleep in a home environment and a noisy environment, the step of detecting sleep events occurring during sleep according to an embodiment of the present invention may utilize consistency training as described above, as shown in FIG. 21. Consistency training is a type of semi-supervised learning model, and consistency training according to an embodiment of the present invention may be a method of performing learning using data to which noise has been intentionally added and data to which noise has not been intentionally added.

[0638] Furthermore, consistency training according to an embodiment of the present invention may be a method of performing learning by generating data of a virtual sleep environment using noise of a target environment.

[0639] According to an embodiment of the present invention, the intentionally added noise may be noise of a target environment, and here, the noise of the target environment may be noise obtained in an environment other than a polysomnography test. Specifically, in order to detect sleep events, various noises may be added by adjusting the SNR and the type of noise to resemble the actual user's environment. In this way, various types of noise obtained in a laboratory and noise occurring in an actual home environment can be collected and learned.

[0640] According to an embodiment of the present invention, for convenience, data to which noise has been intentionally added is referred to as corrupted data. Corrupted data may preferably refer to data to which noise of a target environment has been intentionally added.

[0641] For convenience, data to which noise has not been intentionally added will be referred to as clean data. However, clean data may actually contain noise even if noise has not been intentionally added.

[0642] The clean data used in consistency training according to one embodiment of the present invention may be data acquired in a specific environment (preferably, a sleep polymorphism test environment), and the corrupted data may be data acquired in another environment or a target environment (preferably, an environment other than a sleep polymorphism test).

[0643] Corrupted data according to an embodiment of the present invention may be data in which noise acquired in another environment or target environment (preferably an environment other than a polymorphic sleep test) is intentionally added to clean data.

[0644] In consistency training, when clean data and corrupted data are input to the same deep learning model, a loss function or consistency loss may be defined so that the outputs are the same, and training may be performed to achieve consistent predictions.

[0645] Home Noise Consistency Training

[0646] According to an embodiment of the present invention, detecting sleep events (e.g., apneas, hypopneas, snoring, sleep talking, etc.) occurring during sleep may include home noise consistency training. Home noise consistency training may enable a model to perform robustly in the presence of noise in the home. Home noise consistency training may be performed by advancing consistency training so that the model outputs similar predictions both in the presence and absence of noise, thereby becoming robust to noise.

[0647] According to an embodiment of the present invention, detection of sleep events occurring during sleep can facilitate consistency learning in a home environment. Consistency learning in a home environment may include a consistency loss function. For example, consistency loss may be defined as the mean squared error (MSE) between a prediction of the sound of clean sleep breathing and a prediction of an impaired version of that sound.

[0648] Consistent learning in a home environment, according to one embodiment of the present invention, can randomly sample data from training noises and add noise to clean sleep breathing sounds at random SNRs between -20 and 5 to generate impaired sounds.

[0649] According to an embodiment of the present invention, consistency learning in a home environment may be performed so that the length of the input sequence is 14 epochs and the total length of sampled noise is 7 minutes or more. This allows sleep event detection according to the present invention to detect information within a shorter period of time than sleep stage analysis according to the present invention, thereby increasing the accuracy of sleep event detection.

[0650] Regression analysis for estimating AHI values from event detection

[0651] FIG. 34 is a diagram illustrating a linear regression analysis function used to analyze AHI, which is a sleep apnea occurrence index, through sleep events occurring during sleep, according to an embodiment of the present invention.

[0652] According to an embodiment of the present invention, the AHI index, which indicates the number of respiratory events per unit time (e.g., one hour), can be analyzed independently of the length of an epoch for a sleep stage analysis, separate from sleep stage analysis. Specifically, two or three short sleep events may occur during one epoch, and one long sleep event may occur during multiple epochs. According to an embodiment of the present invention, a regression analysis function can be used to estimate the number of actual events from the number of epochs in which sleep events occurred during sleep. For example, a random sample consensus (RANSAC) regression analysis model can be used. The RANSAC regression analysis model is one method for estimating parameters of a fitting model, and randomly selects sample data and then selects the model that best matches the data.

[0653] Multi-headed multi-task analysis

[0654] A method for analyzing sleep states according to an embodiment of the present invention may include analysis through a deep learning model. The deep learning model according to an embodiment of the present invention may perform multi-task learning and / or multi-task analysis. Specifically, the multi-task learning and multi-task analysis may simultaneously learn tasks according to the embodiments of the present invention described above (e.g., multimodal learning, real-time sleep event analysis, sleep stage analysis, etc.).

[0655] A deep learning model for analyzing a sleep state according to an embodiment of the present invention may perform multi-task learning and multi-task analysis. Specifically, for multi-task learning and analysis, the deep learning model may adopt a structure having multiple heads. Each of the multiple heads may be responsible for a specific task (e.g., multimodal learning, real-time sleep event analysis, sleep stage analysis, etc.). For example, the deep learning model may have a structure having three heads, including a first head, a second head, and a third head. The first head may perform inference and / or classification of sleep stage information, the second head may perform detection and / or classification of sleep apnea and hypopnea among sleep events, and the third head may perform detection and classification of snoring among sleep events. The specific descriptions of the specific tasks of the heads described above are merely examples for explaining the present invention and are not limited thereto. The deep learning model according to the present invention may perform multi-task learning and analysis through a structure having multiple heads, and may optimize multiple tasks or specific tasks by improving data efficiency.

[0656] Effects of the sleep analysis method according to the present invention

[0657] When compared with the results of polysomnography (PSG), it was confirmed that the results of the sleep analysis model using sleep acoustic information as input were highly accurate.

[0658] While existing sleep analysis models predict sleep stages using ECG (Electrocardiogram) or HRV (Heart Rate Variability) as input, the present invention can analyze and infer sleep stages by converting sleep acoustic information into frequency domain information, spectrogram or mel spectrogram, and inputting the converted information. Therefore, since the present invention converts sleep acoustic information into frequency domain information, spectrogram or mel spectrogram, and inputs the converted information, sleep stages can be sensed or acquired in real time through analysis of sleep pattern specificity, unlike existing sleep analysis models.

[0659] FIG. 11 is a graph verifying the performance of the sleep analysis method according to the present invention, comparing polysomnography (PSG) results with analysis results using the AI algorithm according to the present invention.

[0660] As shown in Figure 11, the sleep analysis results obtained by the present invention not only closely match the sleep polymorphism test, but also contain more precise and meaningful information related to sleep stages (Wake, Light, Deep, REM). The hypnogram shown at the bottom of Figure 10 shows the probability of which of four classes (Wake, Light, Deep, REM) the user belongs to in 30-second increments when predicting sleep stages based on input of the user's sleep acoustic information. Here, the four classes represent awake, light sleep, deep sleep, and REM sleep, respectively.

[0661] Figure 12 is a graph verifying the performance of the sleep analysis method according to the present invention, comparing polysomnography (PSG) results with analysis results using the AI algorithm according to the present invention in relation to sleep apnea and hypopnea. The hypnogram at the bottom of Figure 12 shows the probability of a user's sleep disorder being classified as one of two disorders (sleep apnea or hypopnea) in 30-second intervals when predicting a sleep disorder based on input sleep acoustic information. When using the sleep analysis according to the present invention, as shown in Figure 12, the sleep state information obtained according to the present invention not only closely matches the polysomnography test but also includes more precise analysis information related to apnea and hypopnea.

[0662] The present invention can identify the point at which a sleep disorder (sleep apnea, sleep hyperpnea, sleep hypopnea) occurs by analyzing a user's sleep in real time. By providing a stimulus (tactile stimulus, auditory stimulus, olfactory stimulus, etc.) to the user at the moment a sleep disorder occurs, the sleep disorder can be temporarily alleviated. That is, the present invention can interrupt a user's sleep disorder based on accurate event detection related to a sleep disorder and reduce the frequency of the sleep disorder. Furthermore, the present invention has the advantage of enabling highly accurate sleep analysis by performing sleep analysis in a multimodal manner.

[0663] Inference Stage Post-Processing of the Invention

[0664] In the inference step of inferring sleep stages through learning, sleep may be configured not only as a fixed time (e.g., 30 seconds, 20 minutes, etc.) as in the learning data, but also as a sleep duration (e.g., 5 hours, 8 hours, etc.). To perform accurate inference of sleep stages, post-processing may be performed to improve the accuracy of inference based on sleep duration. The specific values related to the above-mentioned time intervals are merely examples, and the present invention is not limited thereto.

[0665] According to an embodiment of the present invention, inferences regarding the depth of sleep based on sleep duration can be post-processed using medical information.

[0666] According to an embodiment of the present invention, post-processing can be performed through artificial intelligence learning using sleep stage information data based on sleep duration.

[0667] The steps of a method or algorithm described in connection with the embodiments of the present invention may be embodied directly in hardware, in a software module executed by hardware, or in a combination thereof. The software module may reside in Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art to which the present invention pertains.

[0668] Components of the present invention may be embodied as a program (or application) stored on a medium for execution in conjunction with a computer, which is hardware. Components of the present invention may be implemented as software programming or software elements. Similarly, embodiments include various algorithms embodied in a combination of data structures, processes, routines, or other programming constructs, and may be implemented in programming or scripting languages such as C, C++, Java, assembler, etc. Functional aspects may be embodied as algorithms executed by one or more processors.

[0669] Those skilled in the art will appreciate that the various illustrative logic blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented by electronic hardware, various forms of programs (referred to herein for convenience as "software"), or design code, or a combination of all of these. To clearly illustrate this interoperability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art will appreciate that the described functionality may be implemented in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0670] Various embodiments presented herein may be embodied as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" includes a computer program, carrier, or medium accessible by any computer-readable device. For example, computer-readable media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Additionally, various storage media presented herein include one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, wireless channels and various other media capable of storing, carrying, and / or transmitting instructions and / or data.

[0671] It is understood that the specific order or hierarchy of steps in the processes presented is an example of a sample approach. Based on design priorities, it is understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of the present invention. The accompanying method claims present elements of the various steps in a sample order, but are not meant to be limited to the specific order or hierarchy presented.

[0672] The description of the embodiments presented is provided to enable any person skilled in the art to use or practice the invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the scope of the invention. The present invention is not intended to be limited to the embodiments presented herein, but rather is to be accorded the widest scope consistent with the principles and novel features disclosed herein. [Explanation of symbols]

[0673] 10: User terminal 100: Computing equipment 110: Network Department 120:Memory 130: Processor 20: External server 11a: Area where object state information or environmental sensing information can be acquired 1a: Electronic device connected to a network within area 11a 1b: Electronic devices connected to a network within the area 11a 1c: Electronic devices not connected to a network within area 11a 1d: Electronic devices not connected to the network within area 11a 2a: Electronic devices outside the scope of area 11a 2b: Electronic devices outside the scope of area 11a E: Acoustic information P: A singularity related to the user's sleep SS:Sleep acoustic information SP: Spectrogram

Claims

1. In a method for analyzing a user's sleep state based on acoustic information, A step of acquiring acoustic information in the time domain related to the user's sleep, wherein the acoustic information includes sounds related to the user's breathing and body movements. A step of performing preprocessing on the acquired acoustic information via a processor in order to obtain preprocessed information, wherein the preprocessed information is obtained by visualizing the time-axis changes of the frequency components of the acquired acoustic information in order to generate a spectrogram. The process involves using the pre-processed information via a processor to perform at least one of the following: extraction or classification of sleep state information, Includes, The aforementioned preprocessing includes converting the spectrogram into a form that is close to a square, The deep learning model includes a Vision Transformer-based neural network configured to process the preprocessed information. A method for analyzing a user's sleep state based on acoustic information.

2. The method for analyzing a user's sleep state based on acoustic information according to Claim 1, wherein the neural network of the Vision Transformer platform includes at least one of Vision Transformer (ViT) and Mobile Vision Transformer (MobileViT).

3. A method for analyzing a user's sleep state based on acoustic information according to Claim 1, wherein the deep learning model is trained using augmentation preprocessing which includes at least one of pitch shifting, TUT (Tile UnTile) augmentation, and noise augmentation.

4. The deep learning model includes a feature extraction model and a feature classification model, The feature extraction model includes one or more neural networks, including encoders pre-trained via autoencoders. The feature classification model includes one or more neural networks, and at least one of the one or more neural networks includes a fully connected layer. A method for analyzing a user's sleep state based on acoustic information as described in claim 1.

5. The feature extraction model extracts a plurality of features based on one or more patterns associated with at least one of the respiratory sounds, respiratory patterns, and movement patterns, The feature classification model infers at least one of the sleep stages and sleep events based on the plurality of features, The aforementioned sleep stages correspond to at least one of the following: wakefulness, light sleep, deep sleep, and REM sleep. The aforementioned sleep events correspond to at least one of the following: sleep apnea, sleep hypopnea, teeth grinding, sleep talking, tossing and turning, sleep disorders, and snoring. A method for analyzing a user's sleep state based on acoustic information as described in claim 4.

6. The sleep state information includes sleep stage information related to the depth of the user's sleep, The pre-processed information includes multiple spectrograms corresponding to each of the already set epochs. A method for analyzing a user's sleep state based on acoustic information as described in claim 1.

7. In a method for analyzing a user's sleep state based on acoustic information, The process involves acquiring acoustic information in the time domain related to the user's sleep, The steps include performing preprocessing on the acquired acoustic information via a processor, The process involves using the pre-processed information as input to a deep learning model via a processor to perform at least one of the following: extraction or classification of sleep state information. Includes, The aforementioned deep learning model is trained by learning based on semi-supervised control loss. The learning based on the aforementioned semi-supervised controlled loss includes setting a critical value for cluster confidence and adjusting the position in the vector space relative to the anchor data based on the set critical value for cluster confidence, The anchor data includes labeled data to which labels have been assigned for sleep states and unlabeled data to which labels have not been assigned for sleep states. A method for analyzing a user's sleep state based on acoustic information.

8. The labeled data includes acoustic information acquired in conjunction with a multi-source sleep study in a restricted environment including at least one of a multi-source sleep study (PSG) environment, a hospital environment, and a laboratory environment, and the labels of the labeled data are based on the results of the multi-source sleep study. The aforementioned unlabeled data includes acoustic information acquired in an unrestricted environment, including the user's home environment, without the use of sleep multi-source testing. A method for analyzing a user's sleep state based on acoustic information as described in claim 7.

9. A method for analyzing a user's sleep state based on acoustic information according to claim 7, wherein adjusting the position in the vector space is such that the anchor data is closer to data belonging to the same cluster in the vector space and pushes the anchor data away from data belonging to different clusters.

10. A method for analyzing a user's sleep state based on acoustic information according to claim 7, wherein the deep learning model is further trained by semi-supervised learning based on sequential consistency loss.

11. A method for analyzing a user's sleep state based on acoustic information according to claim 10, wherein the semi-supervised learning based on sequential consistency loss is performed to take into account the time-series characteristics of the acoustic information.

12. The method for analyzing a user's sleep state based on acoustic information according to claim 7, wherein the deep learning model is further trained based on the pseudo-labels by treating the predicted values ​​based on the unlabeled data as pseudo-labels.

13. A method for analyzing a user's sleep state based on acoustic information according to claim 12, wherein the learning based on the pseudo-labels includes performing an augmentation preprocessing that includes at least one of a weakly-augmented method that modulates the image relatively less and a stronglyly-augmented method that modulates the image relatively more.

14. In a method for analyzing a user's sleep state based on acoustic information, The process involves acquiring acoustic information in the time domain related to the user's sleep, The steps include performing preprocessing on the acquired acoustic information via a processor, The process involves using the pre-processed information as input to a deep learning model via a processor to perform at least one of the following: extraction or classification of sleep state information. Includes, The deep learning model is trained by multi-task learning based on the acoustic information. The deep learning model has a structure that includes multiple heads learned through the multi-task learning process, Each of the aforementioned heads performs one distinct task from among several tasks, including at least sleep stage analysis and sleep event analysis. A method for analyzing a user's sleep state based on acoustic information.

15. The multitask learning is a method for analyzing a user's sleep state based on acoustic information as described in Claim 14, which is based on multimodal learning.

16. The multimodal learning is The first stage involves acquiring acoustic information in the time domain related to the user's sleep, The first stage involves acquiring second-level information, which is information about the user's sleep environment related to the user's sleep. The steps include combining the first information and the second information as multimodal data, The steps include: using the multimodal data as input to the deep learning model to extract features related to at least one of the multiple tasks; including, A method for analyzing a user's sleep state based on acoustic information as described in claim 15.

17. A method for analyzing a user's sleep state based on acoustic information according to claim 14, wherein the sleep event analysis includes detecting sleep events including at least one of sleep apnea, sleep hypopnea, teeth grinding, sleep talking, tossing and turning, sleep disorders, and snoring over a plurality of epochs.

18. A method for analyzing a user's sleep state based on acoustic information according to claim 14, wherein the deep learning model is trained by assigning class weights to sleep event classes.

19. A method for analyzing a user's sleep state based on acoustic information according to claim 18, wherein assigning the class weighting value includes assigning a higher class weighting value to sleep event classes other than the no-event class than the class weighting value of the no-event class.

20. A method for analyzing a user's sleep state based on acoustic information according to claim 14, wherein the sleep stage analysis infers a plurality of sleep stages, including at least one of awake, light sleep, deep sleep, and REM sleep.