Abnormal sleep audio segment identification method, electronic device, and program product

By identifying and analyzing audio segments collected by sensors in a smartphone and using snoring information to determine confidence values, the problem of inaccurate sleep monitoring in existing technologies is solved, enabling accurate identification of abnormal sleep audio segments without affecting sleep quality.

CN116327115BActive Publication Date: 2026-01-02BAIDU INT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111603381.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2026-01-02
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

Existing technologies struggle to monitor sleep patterns over the long term without compromising sleep quality, and sleep monitoring results based on general models are not accurate enough.

Method used

By acquiring multiple initial audio segments collected by sensors, the target audio segment that matches the preset abnormal sleep state is identified, and the confidence value is determined by using the snoring information before and after the target audio segment, thereby accurately identifying the abnormal sleep audio segment.

Benefits of technology

It enables accurate identification of abnormal sleep audio segments without affecting sleep quality, thus improving the accuracy and efficiency of sleep monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116327115B_ABST
    Figure CN116327115B_ABST
Patent Text Reader

Abstract

The method for identifying abnormal sleep audio clips, the electronic device and the program product provided by the present disclosure relate to a deep learning technology, and include: obtaining a plurality of initial audio clips, and determining target audio clips meeting a preset sleep state in the plurality of initial audio clips; determining first snoring sound information before each initial audio clip determines the target audio clip, and second snoring sound information after the target audio clip; determining a confidence value of the target audio clip according to the first snoring sound information and the second snoring sound information; and determining whether the target audio clip is an abnormal sleep audio clip according to the confidence value of each target audio clip. In the scheme provided by the present disclosure, the target audio clip that may be an abnormal sleep audio clip can be initially identified in the plurality of initial audio clips, and then whether the target audio clip is indeed an abnormal sleep audio clip is determined by using the snoring sound information before and after the target audio clip, so that the abnormal sleep audio clip can be accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a deep learning technology in the artificial intelligence technology, and in particular, to an abnormal sleep audio segment identification method, an electronic device, and a program product. BACKGROUND

[0002] Sleep quality is very important to people. In order to enable a user to fully understand his / her own sleep condition, the user can be monitored for sleep.

[0003] In one scheme, a professional can use a dedicated device to monitor the user for sleep to obtain a polysomnography (PSG). In another scheme, a smart phone can be used to monitor the user for sleep. A general model is set in the smart phone, and data collected by the smart phone is input into the model to obtain a sleep monitoring result of the user.

[0004] However, in the first scheme, it is difficult to monitor the sleep state for a long time without affecting the sleep quality in daily life. In the second scheme, only a general model is used to determine the sleep condition of the user, and the output result is not accurate enough. SUMMARY

[0005] The present disclosure provides an abnormal sleep audio segment identification method, an electronic device, and a program product, which are used to accurately identify an abnormal sleep audio segment of a user.

[0006] According to a first aspect of the present disclosure, an abnormal sleep audio segment identification method is provided, comprising:

[0007] obtaining a plurality of initial audio segments collected by a sensor, and determining a target audio segment meeting a preset sleep state in the plurality of initial audio segments; the preset sleep state represents an abnormal sleep state;

[0008] determining first snoring information before the target audio segment and second snoring information after the target audio segment according to each initial audio segment;

[0009] determining a confidence value of the target audio segment according to the first snoring information and the second snoring information; the confidence value is used to represent a possibility that the target audio segment belongs to an abnormal sleep audio segment;

[0010] determining whether the target audio segment is the abnormal sleep audio segment according to the confidence value of each target audio segment.

[0011] According to a second aspect of the present disclosure, an abnormal sleep audio segment identification device is provided, comprising:

[0012] An acquisition unit is configured to acquire a plurality of initial audio clips collected by a sensor;

[0013] A target determination unit is configured to determine, from the plurality of initial audio clips, a target audio clip that meets a preset sleep state, the preset sleep state representing an abnormal sleep state;

[0014] A snoring sound determination unit is configured to determine, for each of the initial audio clips, first snoring sound information before the target audio clip and second snoring sound information after the target audio clip;

[0015] A confidence determination unit is configured to determine, according to the first snoring sound information and the second snoring sound information, a confidence value of the target audio clip, the confidence value representing a possibility that the target audio clip belongs to an abnormal sleep audio clip;

[0016] An abnormality determination unit is configured to determine, according to the confidence value of each of the target audio clips, whether the target audio clip is the abnormal sleep audio clip.

[0017] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein

[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.

[0021] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.

[0022] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to perform the method of the first aspect.

[0023] The method for identifying abnormal sleep audio segments, the electronic device, and the program product provided by the present disclosure include: obtaining a plurality of initial audio segments collected by a sensor, and determining a target audio segment that meets a preset sleep state from the plurality of initial audio segments; the preset sleep state represents an abnormal sleep state; determining first snoring sound information before each initial audio segment and second snoring sound information after the target audio segment; determining a confidence value of the target audio segment according to the first snoring sound information and the second snoring sound information; the confidence value is used to represent the possibility that the target audio segment belongs to an abnormal sleep audio segment; and determining whether the target audio segment is an abnormal sleep audio segment according to the confidence value of each target audio segment. In the scheme provided by the present disclosure, the target audio segment that may be an abnormal sleep audio segment can be initially identified from the plurality of initial audio segments, and then the snoring sound information before and after the target audio segment is used to determine whether the target audio segment is indeed an abnormal sleep audio segment, so that the abnormal sleep audio segment can be accurately determined from the plurality of initial audio segments.

[0024] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0026] Figure 1 A flowchart of a method for identifying abnormal sleep audio segments is shown for an exemplary embodiment of the present disclosure;

[0027] Figure 2 A flowchart of a method for identifying abnormal sleep audio segments is shown for another exemplary embodiment of the present disclosure;

[0028] Figure 3 A sleep curve diagram is shown for an exemplary embodiment of the present disclosure;

[0029] Figure 4 A structural diagram of an abnormal sleep audio segment identification device is shown for an exemplary embodiment of the present disclosure;

[0030] Figure 5 A structural diagram of an abnormal sleep audio segment identification device is shown for another exemplary embodiment of the present disclosure;

[0031] Figure 6 A block diagram of an electronic device for implementing the method of the present disclosure is shown. DETAILED DESCRIPTION

[0032] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are cited as illustrative examples. Various details of the embodiments of the present disclosure are described herein in order to provide a thorough understanding of the present disclosure. It will be understood by those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in the following description, descriptions of well-known functions and constructions are omitted for clarity and conciseness.

[0033] The data, which can be audio data, can be collected by the smart phone, and the collected data can be input into a model set in the smart phone, and a sleep state can be output by the model, so as to achieve the purpose of monitoring the sleep of the user.

[0034] However, it is not accurate enough to determine the sleep state of the user based on the output result of the model, and it is also difficult to accurately identify abnormal sleep data. Therefore, in the scheme provided by the present disclosure, a target audio segment that may belong to an abnormal sleep state is determined in the collected multiple audio segments, and whether the target audio segment is indeed an abnormal sleep audio segment is determined according to snoring information before and after the target audio segment, so that the sleep of the user can be accurately monitored, and the abnormal sleep audio segment can also be identified.

[0035] Figure 1 The flow chart of the method for identifying the abnormal sleep audio segment is shown for an exemplary embodiment of the present disclosure.

[0036] As shown in Figure 1 The method for identifying the abnormal sleep audio segment provided by the present disclosure comprises:

[0037] In step 101, multiple initial audio segments collected by a sensor are obtained, and a target audio segment conforming to a preset sleep state is determined in the multiple initial audio segments; the preset sleep state represents an abnormal sleep state.

[0038] The scheme provided by the present disclosure is executed by an electronic device with computing capability, for example, a smart phone.

[0039] Multiple sensors can be set in the smart phone, for example, a microphone can be included, and the microphone can be used to collect audio segments, for example, initial audio segments of 10 seconds each can be collected.

[0040] The smart phone can execute the method provided by the present disclosure when the user is sleeping, for example, whether the user is sleeping can be determined by other schemes, and if the user is sleeping, the method provided by the present disclosure can be executed. Alternatively, the method provided by the present disclosure can be executed when a preset time point is reached, for example, the method of the present disclosure can be executed when 10:00 PM is reached, and various sleep states of the user can be determined by this method, and abnormal sleep audio segments can be identified.

[0041] Specifically, the smart phone executes the method provided by the present disclosure, and can acquire an initial audio segment collected by a sensor, specifically, an initial audio segment collected by a microphone.

[0042] Further, the smart phone can acquire the data collected by the microphone every certain period of time, and then acquire the initial audio segment. For example, the smart phone can acquire the data collected by the microphone every 10 seconds, thereby obtaining an initial audio segment of 10 seconds.

[0043] In actual application, the smart phone can process each initial audio segment to determine the sleep state of each initial audio segment. For example, the initial audio segment is determined to be a low ventilation state, an apnea state, a continuous breathing state, or a continuous snoring state.

[0044] The low ventilation refers to that the intensity (amplitude) of respiratory airflow during sleep is reduced by more than 50% of the basic level, and the blood oxygen saturation is reduced by more than 4% of the basic level or micro-awakening. The apnea syndrome refers to that loud snoring is suddenly interrupted, the user breathes strongly but does not work, cannot breathe completely, wakes up after a few seconds or even dozens of seconds, gasps loudly, the airway is forced to open, and then continues to breathe.

[0045] If the user is in a continuous breathing state or a continuous snoring state during sleep, it is a normal sleep state. If the user is in a low ventilation state or an apnea state during sleep, it indicates an abnormal sleep state.

[0046] If the initial audio segment meets the preset abnormal sleep state, it is determined that the initial audio segment is a target audio segment.

[0047] Specifically, a model for identifying a sleep state can be set in advance, each initial audio segment is input into the model, the sleep state of each initial audio segment is obtained, and then the target audio segment is obtained.

[0048] Step 102, determining first snoring information before the target audio segment and second snoring information after the target audio segment according to each initial audio segment.

[0049] Further, after the target audio segment is determined, the target audio segment can be considered as a suspected abnormal sleep audio segment, and whether the target audio segment is indeed an abnormal sleep audio segment can be further determined.

[0050] In actual application, the first snoring information of the user before the target audio segment and the second snoring information of the user after the target audio segment can be determined, and whether the target audio segment is indeed an abnormal sleep segment can be determined through the first snoring information and the second snoring information.

[0051] For example, a first initial audio segment before the target audio segment can be determined, and then the first snoring information can be determined according to the initial audio segment, a second initial audio segment after the target audio segment can be determined, and then the second snoring information can be determined according to the initial audio segment.

[0052] For example, the first snoring information can be determined according to the snoring intensity at multiple time points in the first initial audio segment, and the second snoring information can be determined according to the snoring intensity at multiple time points in the second initial audio segment. The snoring intensity at the last time point in the first initial audio segment can be determined as the first snoring information, and the snoring intensity at the starting time point in the second initial audio segment can be determined as the second snoring information.

[0053] In step 103, the confidence value of the target audio segment is determined according to the first snoring information and the second snoring information; the confidence value is used to represent the possibility that the target audio segment belongs to the abnormal sleep audio segment.

[0054] In step 103, the confidence value of the target audio segment is determined according to the first snoring information and the second snoring information; the confidence value is used to represent the possibility that the target audio segment belongs to the abnormal sleep audio segment.

[0055] For example, if the intensity difference between the first snoring information and the second snoring information is large, the determined confidence value is high, representing that the possibility that the target audio segment belongs to the abnormal sleep audio segment is greater. If the intensity difference between the first snoring information and the second snoring information is small, the determined confidence value is low, representing that the possibility that the target audio segment belongs to the abnormal sleep audio segment is smaller.

[0056] In step 104, whether the target audio segment is an abnormal sleep audio segment is determined according to the confidence value of each target audio segment.

[0057] Specifically, the smart phone can determine whether the target audio segment is an abnormal sleep audio segment according to the confidence value of the target audio segment.

[0058] For example, a confidence threshold can be set in advance, if the determined confidence value of the target audio segment is greater than the confidence threshold, it can be determined that the target audio segment is an abnormal sleep audio segment, otherwise, it is determined that the target audio segment is not an abnormal sleep audio segment.

[0059] The method for identifying abnormal sleep audio segments provided in this disclosure includes: acquiring multiple initial audio segments collected by sensors, and identifying a target audio segment that conforms to a preset sleep state among the multiple initial audio segments; the preset sleep state represents an abnormal sleep state; determining first snoring information before the target audio segment and second snoring information after the target audio segment based on each initial audio segment; determining a confidence value of the target audio segment based on the first snoring information and the second snoring information; the confidence value is used to characterize the probability that the target audio segment belongs to an abnormal sleep audio segment; and determining whether the target audio segment is an abnormal sleep audio segment based on the confidence value of each target audio segment. In the method for identifying abnormal sleep audio segments provided in this disclosure, a target audio segment that may be an abnormal sleep audio segment can be initially identified among multiple initial audio segments, and then the snoring information before and after the target audio segment can be used to determine whether the target audio segment is indeed an abnormal sleep audio segment, thereby accurately identifying the abnormal sleep audio segment among multiple initial audio segments.

[0060] Figure 2 A flowchart illustrating a method for identifying abnormal sleep audio segments, as shown in another exemplary embodiment of this disclosure.

[0061] like Figure 2 As shown, the method for identifying abnormal sleep audio segments provided in this disclosure includes:

[0062] Step 201: Acquire multiple initial audio segments using sensors.

[0063] Step 201 is similar to obtaining the initial audio segment in step 101.

[0064] Step 202: Input the initial audio segment into the preset sleep event recognition model to obtain the sleep event recognition result corresponding to the initial audio segment.

[0065] A sleep event recognition model can be set up in a smartphone. This model can be pre-trained using machine learning techniques. For example, multiple training audio data sets can be collected in advance, and each set can be labeled with a sleep event. The model can then be trained using the labeled audio data to obtain a sleep event recognition model that can identify sleep events in the audio data.

[0066] The initial audio segment can be input into the preset sleep event recognition model, which can then output the sleep event of the initial audio segment.

[0067] Specifically, sleep events can include snoring events, sleep talking events, breathing events, and ambient noise events.

[0068] In step 203, if the sleep event recognition result of the initial audio segment conforms to the preset sleep event, the sleep state recognition result of the initial audio segment is determined; the sleep state recognition result includes the preset sleep state.

[0069] Further, if it is determined that the sleep event recognition result of the initial audio segment conforms to the preset sleep event, it indicates that there is a possibility that an abnormal sleep segment exists in the initial audio segment, and therefore, the sleep state of the initial audio segment can be continuously recognized.

[0070] If it is determined that the sleep event recognition result of the initial audio segment does not conform to the preset sleep event, it indicates that there is no possibility that an abnormal sleep segment exists in the initial audio segment, and therefore, the sleep state of the initial audio segment does not need to be detected.

[0071] Through this embodiment, the audio segment that needs to recognize the sleep state in the initial audio segment can be determined, only the sleep state of the initial audio segment conforming to the preset sleep event needs to be recognized, and therefore, the amount of audio data that needs to recognize the sleep state is reduced, and the data processing speed is improved.

[0072] If the sleep event recognition result of the initial audio segment conforms to the preset sleep event, the sleep state recognition result of the initial audio segment can be further determined.

[0073] The audio feature of the initial audio segment can be extracted; and then the sleep state recognition result of the initial audio segment is determined according to the audio feature. The audio feature in the initial audio segment is processed, and the sleep state recognition result of the initial audio segment can be more accurately obtained.

[0074] Specifically, the audio feature of the initial audio segment can be extracted, for example, F-Bank, MFCC, spectral energy, and the like; and then the classification model obtained by pre-training is used to classify and process the audio feature to obtain a classification result, which can include the recognition result of multiple sleep states. The classification model can be, for example, a random forest, an SVM, a decision tree, and the like.

[0075] The sleep state can include, for example, low ventilation, apnea, continuous breathing, continuous snoring, and the like. The classification model can process the audio feature of the initial audio segment to obtain the probability that the initial audio segment belongs to each sleep state, and then obtain the sleep state recognition result of the initial audio segment.

[0076] The preset sleep event includes a snoring event and a breathing event; and the preset sleep state includes low ventilation and apnea.

[0077] The sleep states of low ventilation and apnea are abnormal sleep states, and in the sleep segment labeled with the snoring event and the breathing event, the sleep states of low ventilation and apnea can exist.

[0078] Therefore, the audio segments in which the snoring events and the breathing events exist are determined in the initial audio segments, and then the audio segments in which the hypopnea and the apnea sleep state exist are determined in the audio segments, so that the abnormal sleep audio segments of hypopnea and apnea can be accurately identified.

[0079] In step 204, if the sleep state of the initial audio segment is the preset sleep state, the initial audio segment is determined as the target audio segment.

[0080] In actual application, if the sleep state of the initial audio segment is the preset sleep state, it can be considered that the initial audio segment may have an abnormal sleep state, and therefore, the initial audio segment can be determined as the target audio segment.

[0081] In step 205, in the initial audio segment, a first audio segment before the target audio segment and a second audio segment after the target audio segment are obtained.

[0082] The target audio segment determined may be an abnormal sleep audio segment, in order to more accurately determine whether it is an abnormal sleep audio segment, a first audio segment before the target audio segment and a second audio segment after the target audio segment can also be obtained.

[0083] Specifically, each initial audio segment obtained by the smart phone can have time information, for example, a segment 1, a segment 2 and a segment 3 are obtained, the time information of the three segments is continuous, if the segment 2 is the target audio segment, the segment 1 can be determined as the first audio segment before the target audio segment, and the segment 3 can be determined as the second audio segment after the target audio segment.

[0084] Further, the snoring information before and after the target audio segment can be determined according to other audio segments before and after the target audio segment, and then whether the target audio segment is an abnormal sleep audio segment can be more accurately determined.

[0085] In step 206, first snoring information before the target audio segment in the first audio segment is extracted, and second snoring information after the target audio segment in the second audio segment is extracted.

[0086] In actual application, the smart phone can extract the first snoring information in the first audio segment, for example, the snoring information at each time in the first audio segment can be processed to obtain the first snoring information, for example, the average value of the snoring intensity at each time can be taken as the first snoring information.

[0087] The smart phone can extract the second snoring information in the second audio segment. For example, the smart phone can process the snoring information at each time point in the second audio segment to obtain the second snoring information. For example, the average value of the snoring intensity at each time point can be taken as the second snoring information.

[0088] Specifically, the smart phone can take the snoring intensity at the last time point in the first audio segment as the first snoring information, and take the snoring intensity at the first time point in the second audio segment as the second snoring information.

[0089] Through this implementation manner, the first snoring information before the target audio segment and the second snoring information after the target audio segment can be determined, so that whether the target audio segment is an abnormal sleep audio segment can be determined according to the two snoring information. In this way, whether the target audio segment is an abnormal sleep audio segment can be determined in combination with the target audio segment itself and other audio segments before and after the target audio segment, so that the recognition result is more accurate.

[0090] In step 207, the abnormal snoring intensity is determined according to the first snoring information and the second snoring information.

[0091] Further, the first snoring information and the second snoring information can be snoring intensity, which can be specifically represented by a numerical value.

[0092] In actual application, the abnormal snoring intensity before and after the target audio segment can be determined according to the first snoring information and the second snoring information. The abnormal snoring intensity can be specifically a difference value of the snoring intensity before and after the target audio segment. For example, the difference value between the intensity value in the first snoring information and the intensity value in the second snoring information can be taken as the abnormal snoring intensity of the target audio segment.

[0093] In step 208, the confidence value of the target audio segment is determined according to the abnormal snoring intensity, the preset sleep state corresponding to the target audio segment, and the time length of the target audio segment.

[0094] The smart phone can determine the confidence value of the target audio segment according to the abnormal snoring intensity of the target audio segment, and the preset sleep state and the duration corresponding to the target audio segment.

[0095] For example, the abnormal snoring intensity of the target audio segment is p, the sleep state is low ventilation, and the duration of low ventilation is t. The confidence value of the target audio segment can be determined according to these three information. Values corresponding to different preset sleep states can be set in advance, or confidence determination methods corresponding to different preset sleep states can be set, so that the confidence value of the target audio segment can be determined.

[0096] The confidence value can be a value in the range of 0-1, and the greater the confidence value, the greater the possibility that the target audio segment is an abnormal sleep audio segment.

[0097] By this implementation, different preset sleep states, state duration, and other factors are fully considered, and these factors are combined with the abnormal snoring intensity of the target audio segment to determine whether the target audio segment is an abnormal sleep audio segment. Therefore, the abnormal sleep audio segment can be accurately determined.

[0098] For example, if the preset sleep state duration in the target audio segment is short, the target audio segment can not be an abnormal sleep audio segment. For another example, if the preset sleep state durations of two target audio segments are the same and the abnormal snoring intensities are the same, but the preset sleep states are different, the determination results can also be different, such as determining one of the target audio segments as an abnormal sleep audio segment and determining the other target audio segment as not an abnormal sleep audio segment.

[0099] In step 209, if the confidence value of the target audio segment is greater than a threshold value, the target audio segment is determined as an abnormal sleep audio segment.

[0100] The threshold value can be set in advance, and the confidence value of the target audio segment is compared with the threshold value to determine whether the target audio segment is an abnormal sleep audio segment.

[0101] Specifically, if the confidence value of the target audio segment is greater than a preset threshold value, the target audio segment can be determined as an abnormal sleep audio segment.

[0102] After the target audio segment that can be an abnormal sleep audio segment is determined, the confidence value of the target audio segment can be used to further determine whether the target audio segment is indeed an abnormal sleep audio segment, so that the abnormal sleep audio segment can be more accurately identified.

[0103] In step 210, a historical abnormal segment is obtained according to the sleep state of the abnormal sleep audio segment. The historical abnormal segment is an abnormal segment determined by user operation.

[0104] Specifically, the smart phone can also obtain the historical abnormal segment.

[0105] When the user uses the smart phone, the user can confirm which of the initial audio segments collected by the smart phone are abnormal segments, and specifically, the user can confirm which of the initial audio segments are low-airway abnormal segments and which of the initial audio segments are apnea abnormal segments.

[0106] Further, the smart phone can obtain the sleep state of the abnormal sleep audio segment, and obtain historical abnormal segments consistent with the sleep state. For example, if the sleep state of the currently determined abnormal sleep audio segment is a low ventilation segment, one or more historical abnormal segments belonging to the low ventilation segment can be obtained.

[0107] In step 211, it is determined whether the abnormal sleep audio segment is a real abnormal segment according to the historical abnormal segment.

[0108] In actual application, it can be determined again whether the abnormal sleep audio segment is a real abnormal segment according to the historical abnormal segment.

[0109] The historical abnormal segment is an abnormal sleep segment manually confirmed by the user, and thus can be considered as accurate. The historical abnormal segment can be used as standard data to determine again whether the historical abnormal segment determined by the smart phone is a real abnormal segment.

[0110] Specifically, the smart phone can compare the historical abnormal segment with the abnormal sleep audio segment. If the two are similar, it can be determined that the abnormal sleep audio segment is a real abnormal segment, otherwise, it is determined that the abnormal sleep audio segment is not a real abnormal segment.

[0111] In this way, the real abnormal segment can be determined more accurately by combining the historical abnormal segment of the user. In addition, the abnormal sleep segments of different users are not completely the same, and thus the personalized recognition can be performed according to the data of different users by the scheme of the present disclosure, and the recognition accuracy of the abnormal sleep audio segment is further improved.

[0112] Further, the smart phone can obtain the first audio feature of the historical abnormal segment, obtain the second audio feature of the abnormal sleep audio segment, and determine the similarity between the first audio feature and the second audio feature.

[0113] If the similarity between the first audio feature and the second audio feature meets a preset condition, it is determined that the abnormal sleep audio segment is a real abnormal segment. By comparing the features of the historical abnormal segment and the abnormal sleep audio segment, it can be determined whether the two are similar, and whether the abnormal sleep audio segment is similar to the historical abnormal segment of the user, and thus the abnormal sleep audio segment is determined to be consistent with the abnormal sleep condition of the user, and the purpose of personalized recognition is achieved.

[0114] In an optional embodiment, the corresponding auxiliary sleep device information and / or auxiliary medical resource information can also be obtained and displayed according to the abnormal sleep audio segment.

[0115] In actual application, the smart phone can also acquire information of an auxiliary sleep device for solving the abnormal sleep state according to the abnormal sleep audio segment, and acquire information of an auxiliary medical resource for solving the abnormal sleep state, and then display the information, so that the user can operate the smart phone to purchase the corresponding device or understand the corresponding auxiliary medical resource.

[0116] The auxiliary medical resource may be, for example, an online consultation resource. A corresponding function entry can be displayed on the smart phone, so that the user can use the function.

[0117] By the scheme provided in the embodiments of the present disclosure, the sleep problem can be automatically identified and a sleep problem solution can be provided, so that an integrated solution for solving the sleep problem can be provided for the user.

[0118] In an optional implementation, a sleep curve is generated and displayed according to the abnormal sleep audio segment, and the abnormal sleep audio segment is marked in the sleep curve. The sleep curve is used to represent the sleep state of the user at each time.

[0119] In the scheme provided in the present disclosure, the sleep curve can also be generated and displayed, and the abnormal sleep audio segment can be marked in the curve. For example, the sleep segment between the first time and the second time can be marked as the abnormal sleep audio segment for the user's reference.

[0120] In this way, the user can understand his / her own sleep condition in a more intuitive way, and the user can also operate the sleep curve to play the abnormal sleep audio segment, which can be confirmed or denied by the user, so that the historical abnormal segment can be updated according to the operation, and the sleep condition of the user can be more accurately identified.

[0121] Specifically, the data collected by the gyroscope, microphone and other sensors of the smart phone can be combined with artificial intelligence technology and normal sleep cycle rules, and the scheme can monitor the sleep cycle of the user, mainly monitoring three sleep cycles of wakefulness, light sleep and deep sleep.

[0122] The data collected by the gyroscope, microphone, smart wearable device and the like can be input into a preset deep learning model, and the model can estimate the sleep cycle of the feature data, and predict the sleep cycle of the user in combination with the sleep cycle rules of the human (light sleep and deep sleep appear alternately, and each cycle is 90-100 minutes).

[0123] The sleep curve can also be generated in combination with the sleep cycle, so that each sleep stage can also be marked in the sleep curve, so that the content of the sleep curve is more abundant, and the user can intuitively understand his / her own sleep condition.

[0124] In an alternative embodiment, a sleep report can also be outputted every day, every week, etc., in which the sleep condition of the user in this period of time is recorded.

[0125] Figure 3 A sleep curve diagram shown for an exemplary embodiment of the present disclosure.

[0126] As shown in Figure 3 , the smart phone can display a sleep curve as shown in Figure 3 . The user can also click on any position in the curve, and the smart phone can display the sleep state corresponding to the position.

[0127] Figure 4 A structure diagram of an abnormal sleep audio segment identification device shown for an exemplary embodiment of the present disclosure.

[0128] As shown in Figure 4 , the abnormal sleep audio segment identification device 400 provided by the present disclosure comprises:

[0129] An acquisition unit 410, configured to acquire a plurality of initial audio segments collected by a sensor;

[0130] A target determination unit 420, configured to determine a target audio segment conforming to a preset sleep state from the plurality of initial audio segments; the preset sleep state represents an abnormal sleep state;

[0131] A snoring sound determination unit 430, configured to determine first snoring sound information before the target audio segment and second snoring sound information after the target audio segment according to each of the initial audio segments;

[0132] A confidence determination unit 440, configured to determine a confidence value of the target audio segment according to the first snoring sound information and the second snoring sound information; the confidence value is used to represent the possibility that the target audio segment belongs to an abnormal sleep audio segment;

[0133] An abnormality determination unit 450, configured to determine whether the target audio segment is the abnormal sleep audio segment according to the confidence value of each of the target audio segments.

[0134] The abnormal sleep audio segment identification device provided by the present disclosure can preliminarily identify a target audio segment that may be an abnormal sleep audio segment from a plurality of initial audio segments, and then determine whether the target audio segment is indeed an abnormal sleep audio segment by using the snoring sound information before and after the target audio segment, so as to accurately determine the abnormal sleep audio segment from the plurality of initial audio segments.

[0135] Figure 5 A structure diagram of an abnormal sleep audio segment identification device shown for another exemplary embodiment of the present disclosure.

[0136] As Figure 5 shown, the abnormal sleep audio segment identification device 500 provided by the present disclosure includes an acquisition unit 510, a target determination unit 520, a snoring sound determination unit 530, a confidence determination unit 540, and an anomaly determination unit 550. Figure 4 The acquisition unit 510 is similar to the acquisition unit 410 shown in the figure, the target determination unit 520 is similar to the target determination unit 420 shown in the figure, the snoring sound determination unit 530 is similar to the snoring sound determination unit 430 shown in the figure, the confidence determination unit 540 is similar to the confidence determination unit 440 shown in the figure, and the anomaly determination unit 550 is similar to the anomaly determination unit 450 shown in the figure. Figure 4 Figure 4 Figure 4 Figure 4

[0137] Optionally, the target determination unit 520 includes:

[0138] an event identification module 521, configured to input the initial audio segment into a preset sleep event identification model to obtain a sleep event identification result corresponding to the initial audio segment;

[0139] a state identification module 522, configured to determine a sleep state identification result of the initial audio segment if the sleep event identification result of the initial audio segment meets a preset sleep event; the sleep state identification result includes the preset sleep state;

[0140] a target determination module 523, configured to determine that the initial audio segment is the target audio segment if the sleep state of the initial audio segment is the preset sleep state.

[0141] Optionally, the state identification module 522 is specifically configured to:

[0142] extract an audio feature of the initial audio segment;

[0143] determine the sleep state identification result of the initial audio segment according to the audio feature.

[0144] Optionally, the preset sleep event includes a snoring sound event and a breathing event.

[0145] The preset sleep state includes low ventilation and apnea.

[0146] Optionally, the snoring sound determination unit 530 includes:

[0147] a segment acquisition module 531, configured to acquire, in the initial audio segment, a first audio segment before the target audio segment and a second audio segment after the target audio segment;

[0148] ​​​​The snoring sound determination module 532 is configured to determine first snoring sound information before the target audio segment is extracted from the first audio segment and determine second snoring sound information after the target audio segment is extracted from the second audio segment.

[0149] Optionally, the confidence determination unit 540 includes:

[0150] The intensity determination module 541 is configured to determine abnormal snoring sound intensity according to the first snoring sound information and the second snoring sound information.

[0151] The confidence determination module 542 is configured to determine a confidence value of the target audio segment according to the abnormal snoring sound intensity, a preset sleep state corresponding to the target audio segment, and a duration of the preset sleep state.

[0152] Optionally, the anomaly determination unit 550 is configured to:

[0153] If the confidence value of the target audio segment is greater than a threshold value, the target audio segment is determined to be the abnormal sleep audio segment.

[0154] Optionally, if the target audio segment is determined to be the abnormal sleep audio segment, the device further includes a confirmation unit 560 configured to:

[0155] According to the sleep state of the abnormal sleep audio segment, a historical abnormal segment is obtained, wherein the historical abnormal segment is an abnormal segment determined through user operation.

[0156] According to the historical abnormal segment, it is determined whether the abnormal sleep audio segment is a real abnormal segment.

[0157] Optionally, the confirmation unit 560 includes:

[0158] The feature acquisition module 561 is configured to obtain a first audio feature of the historical abnormal segment and obtain a second audio feature of the abnormal sleep audio segment.

[0159] The confirmation module 562 is configured to determine that the abnormal sleep audio segment is a real abnormal segment if a similarity between the first audio feature and the second audio feature meets a preset condition.

[0160] Optionally, the device further includes an information acquisition unit 570 configured to:

[0161] According to the abnormal sleep audio segment, corresponding auxiliary sleep device information and / or auxiliary medical resource information is obtained and displayed.

[0162] Optionally, the device further includes a curve generation unit 580 configured to:

[0163] A sleep curve is generated and displayed based on the abnormal sleep audio segments, and the abnormal sleep audio segments are marked on the sleep curve; the sleep curve is used to characterize the user's sleep state at each time.

[0164] This disclosure provides a method, electronic device, and program product for identifying abnormal sleep audio segments, which utilizes deep learning technology in artificial intelligence to accurately identify abnormal sleep audio segments of users.

[0165] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0166] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0167] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.

[0168] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0169] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded into random access memory (RAM) 603 from storage unit 608. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0170] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0171] The computing unit 601 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the identification method of abnormal sleep audio segments. For example, in some embodiments, the identification method of abnormal sleep audio segments can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the identification method of abnormal sleep audio segments described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the identification method of abnormal sleep audio segments by any other appropriate means, such as by means of firmware.

[0172] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0173] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0174] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0175] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0176] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0177] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions typically taking place over a communication network. The relationship of client and server arises by interplay of both computers programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0178] It should be understood that the various forms of flow shown above can be reordered, steps added or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, in different orders, without limitation herein, so long as the desired results of the technology disclosed in the present disclosure are achieved.

[0179] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A method for identifying abnormal sleep audio segments, comprising: obtaining a plurality of initial audio segments collected by a sensor, and determining a target audio segment conforming to a preset sleep state from the plurality of initial audio segments; the preset sleep state represents an abnormal sleep state; determining first snoring information before the target audio segment and second snoring information after the target audio segment according to each initial audio segment; determining abnormal snoring intensity according to the first snoring information and the second snoring information; the abnormal snoring intensity is a difference between an intensity value in the first snoring information and an intensity value in the second snoring information; determining a confidence value of the target audio segment according to the abnormal snoring intensity, the preset sleep state corresponding to the target audio segment, and a duration of the preset sleep state; the confidence value is used to represent a possibility that the target audio segment belongs to an abnormal sleep audio segment; determining whether the target audio segment is the abnormal sleep audio segment according to the confidence value of each target audio segment; if the confidence value of the target audio segment is greater than a threshold value, it is determined that the target audio segment is the abnormal sleep audio segment.

2. The method of claim 1, wherein, the determination of the target audio segment conforming to the preset sleep state from the plurality of initial audio segments comprises: inputting the initial audio segment into a preset sleep event recognition model to obtain a sleep event recognition result corresponding to the initial audio segment; if the sleep event recognition result of the initial audio segment conforms to a preset sleep event, determining a sleep state recognition result of the initial audio segment; the sleep state recognition result includes the preset sleep state; if the sleep state of the initial audio segment is the preset sleep state, determining that the initial audio segment is the target audio segment.

3. The method of claim 2, wherein, the determination of the sleep state recognition result of the initial audio segment comprises: extracting an audio feature of the initial audio segment; determining the sleep state recognition result of the initial audio segment according to the audio feature.

4. The method of claim 2 or 3, wherein: the preset sleep event includes a snoring event and a breathing event; the preset sleep state includes low ventilation and apnea.

5. The method according to any one of claims 1-3, wherein, the determination of the first snoring information before the target audio segment and the second snoring information after the target audio segment according to the initial audio segment comprises: in the initial audio segment, obtaining a first audio segment before the target audio segment and a second audio segment after the target audio segment; extracting the first snoring information before the target audio segment in the first audio segment and the second snoring information after the target audio segment in the second audio segment.

6. The method of any one of claims 1-3, if it is determined that the target audio segment is the abnormal sleep audio segment, the method further comprises: obtaining a historical abnormal segment according to a sleep state of the abnormal sleep audio segment; wherein the historical abnormal segment is an abnormal segment determined by user operation. determine whether the abnormal sleep audio segment is a real abnormal sleep audio segment according to the historical abnormal sleep audio segment.

7. The method of claim 6, wherein, The determining whether the abnormal sleep audio segment is a real abnormal sleep audio segment according to the historical abnormal sleep audio segment comprises: obtaining a first audio feature of the historical abnormal sleep audio segment and a second audio feature of the abnormal sleep audio segment; if a similarity between the first audio feature and the second audio feature meets a preset condition, determining that the abnormal sleep audio segment is a real abnormal sleep audio segment.

8. The method according to any one of claims 1-3 and 7, further comprising: obtaining and displaying corresponding auxiliary sleep device information and / or auxiliary medical resource information according to the abnormal sleep audio segment.

9. The method according to any one of claims 1-3 and 7, further comprising: generating and displaying a sleep curve according to the abnormal sleep audio segment, and marking the abnormal sleep audio segment in the sleep curve; the sleep curve is used to represent a sleep state of a user at each time.

10. An abnormal sleep audio segment identification device, comprising: an obtaining unit, configured to obtain a plurality of initial audio segments collected by a sensor; a target determining unit, configured to determine a target audio segment meeting a preset sleep state from the plurality of initial audio segments; the preset sleep state represents an abnormal sleep state; a snoring sound determining unit, configured to determine first snoring sound information before the target audio segment and second snoring sound information after the target audio segment according to each of the initial audio segments; a confidence determining unit, configured to determine a confidence value of the target audio segment according to the first snoring sound information and the second snoring sound information; the confidence value is used to represent a possibility that the target audio segment belongs to an abnormal sleep audio segment; an abnormality determining unit, configured to determine whether the target audio segment is the abnormal sleep audio segment according to the confidence value of each of the target audio segments; the snoring sound determining unit comprises: a segment obtaining module, configured to obtain a first audio segment before the target audio segment and a second audio segment after the target audio segment from the initial audio segments; a snoring sound determining module, configured to extract the first snoring sound information before the target audio segment from the first audio segment and extract the second snoring sound information after the target audio segment from the second audio segment; the confidence determining unit comprises: an intensity determining module, configured to determine an abnormal snoring sound intensity according to the first snoring sound information and the second snoring sound information; the abnormal snoring sound intensity is a difference value between an intensity value in the first snoring sound information and an intensity value in the second snoring sound information; a confidence determining module, configured to determine the confidence value of the target audio segment according to the abnormal snoring sound intensity, a preset sleep state corresponding to the target audio segment, and a duration of the preset sleep state.

11. The apparatus of claim 10, wherein, the target determining unit comprises: an event recognition module, configured to input the initial audio segments into a preset sleep event recognition model to obtain a sleep event recognition result corresponding to the initial audio segments; The state recognition module is configured to: extract audio features of the initial audio segment; 12. The apparatus of claim 11, wherein, determine a sleep state recognition result of the initial audio segment according to the audio features.

13. The apparatus of claim 11 or 12, wherein, the preset sleep event includes snoring event and breathing event; the preset sleep state includes low ventilation and apnea. The anomaly determination unit is configured to: if the confidence value of the target audio segment is greater than a threshold value, determine that the target audio segment is the abnormal sleep audio segment.

14. The apparatus of any one of claims 10-12, wherein, 15. The apparatus of any one of claims 10-12, if it is determined that the target audio segment is the abnormal sleep audio segment, the apparatus further comprises a confirmation unit configured to: the historical abnormal segment is an abnormal segment determined by user operation; determine whether the abnormal sleep audio segment is a real abnormal segment according to the historical abnormal segment. According to the sleep state of the abnormal sleep audio segment, a history abnormal segment is obtained; wherein The confirmation unit comprises: a feature acquisition module configured to acquire first audio features of the historical abnormal segment and second audio features of the abnormal sleep audio segment; 16. The apparatus of claim 15, wherein, a confirmation module configured to determine that the abnormal sleep audio segment is a real abnormal segment if a similarity of the first audio features and the second audio features meets a preset condition.

17. The apparatus of any one of claims 10-12 and 16, further comprising an information acquisition unit configured to: acquire and display corresponding auxiliary sleep device information and / or auxiliary medical resource information according to the abnormal sleep audio segment.

18. The apparatus of any one of claims 10-12 and 16, further comprising a curve generation unit configured to: generate and display a sleep curve according to the abnormal sleep audio segment, and mark the abnormal sleep audio segment in the sleep curve; the sleep curve is used to represent a sleep state of a user at each time.

19. An electronic device comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9. The computer instructions are used to enable the computer to perform the method of any one of claims 1-9.

21. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of any one of claims 1-9.

20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, ​ ​

Citation Information

Patent Citations

  • Method for diagnosing obstructive sleep apnea hypopnea syndrome according to snore

    CN102579010A

  • Intelligent voice recognition wrist type sleep apnea monitoring system

    CN113397492A