Sleepiness prediction system
Patent Information
- Application Number
- JP2025022727
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-09-07
- Estimated Expiration
- 2045-02-14
Smart Images

Figure 2026141802000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a drowsiness prediction system that detects drowsiness of a target speaker by predicting a decrease in the heart rate of the target speaker from speech uttered by the target speaker to be predicted. [Background Art]
[0002] Conventionally, based on the relationship between a decrease in heart rate and drowsiness, as disclosed in Patent Document 1 below, there is known a system that detects drowsiness of a mobility passenger in a relatively short time based on the temporal change of the passenger's heartbeat interval and a predetermined criterion for the temporal change of the heartbeat interval with respect to the passenger's drowsiness.
[0003] As a system for predicting drowsiness several hours in advance, as disclosed in Patent Document 2 below, the system comprises: an acquisition device that acquires a biological signal BS of a passenger who has boarded a mobility such as an automobile when the passenger boards the mobility; an estimation unit that estimates a heart rate signal HS of the passenger based on the biological signal BS; a calculation unit that calculates drowsiness precursor data DD related to the passenger's drowsiness precursor based on the heart rate signal HS; and a state estimation unit that estimates whether or not the passenger's drowsiness has increased based on current drowsiness precursor data DDA of the passenger who is currently boarding the mobility and past drowsiness precursor data DDB detected when the passenger boarded the mobility in the past. There is known a drowsiness information service providing apparatus for a passenger including the above components.
[0004] According to such a drowsiness information service providing apparatus, when the passenger's drowsiness increases, it may be affected by the passenger's past living conditions (for example, it is assumed that the passenger has been in a sleep-deprived state for several days in the past or has accumulated physical fatigue). In such a case, even if no drowsiness has been detected, it is considered that there is a high possibility that the passenger's drowsiness will increase. As described above, the drowsiness information service providing apparatus estimates whether or not the passenger's drowsiness tends to increase based on the passenger's past data related to drowsiness. [Prior Art Documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2018-192128 [Patent Document 2] Japanese Patent Publication No. 2024-092783 [Overview of the project] [Problems that the invention aims to solve]
[0006] However, conventional sleepiness information service providers assume that all passengers wear a device to acquire heart rate signals beforehand and live their lives accordingly. This presents a problem in that it is difficult to predict sleepiness if biometric information is not collected by such a device or if the collection is insufficient.
[0007] In view of the above circumstances, the present invention aims to provide a sleepiness prediction system that can predict future sleepiness for any person based on their current state. [Means for solving the problem]
[0008] The sleepiness prediction system of the first invention is a sleepiness prediction system that predicts the decrease in the heart rate of a target speaker based on the voice emitted by the target speaker, and predicts the target speaker's sleepiness. A speech analysis model was constructed using binary logistic regression analysis with the following parameters: speech data of the voice emitted by the speaker and data on the correct heart rate reduction for that voice, with the explanatory variable representing heart rate reduction as 1 and non-heart rate reduction as 0, and using 12 frequency divisions as explanatory variables. A voice data acquisition unit that acquires voice data of the aforementioned speaker, The voice data acquired by the voice data acquisition unit is input to the voice analysis model, and the drowsiness output unit outputs the predicted drowsiness. It is characterized by being equipped with [the following features].
[0009] According to the sleepiness prediction system of the first invention, a speech analysis model is constructed by binary logistic regression analysis using speech data of speech emitted by a speaker and data of the correct heart rate decrease for said speech, with the explanatory variable being 1 for heart rate decrease and 0 for no heart rate decrease, and using 12 frequency divisions as explanatory variables. Using this speech analysis model, sleepiness can be predicted with high accuracy from the subsequent decrease in the speaker's heart rate.
[0010] Thus, according to the sleepiness prediction system of the first invention, it is possible to predict future sleepiness for any person based on their current state.
[0011] The drowsiness prediction system of the second invention is, in the first invention, The aforementioned speech analysis model is characterized by being reconstructed using only a predetermined proportion of data with high or low probabilities.
[0012] According to the drowsiness prediction system of the second invention, in the speech analysis model, by reconstructing the speech analysis model using only a predetermined proportion of data with high or low probabilities, if the speech analysis model is reconstructed using only the data with high probabilities, then this speech analysis model can predict drowsiness with higher accuracy from the subsequent decrease in the speaker's heart rate. On the other hand, if the speech analysis model is reconstructed using only the data with low probabilities, then this speech analysis model can predict with higher accuracy that the speaker's heart rate will not decrease afterward, for example, that they will remain awake.
[0013] Thus, according to the drowsiness prediction system of the second invention, it is possible to predict the subsequent drowsiness and wakefulness levels of any individual based on their current state with a higher probability.
[0014] The third invention's drowsiness prediction system is, in the first invention, The voice analysis model is characterized by being reconstructed using a predetermined proportion of data with higher probabilities and a predetermined proportion of data with lower probabilities.
[0015] According to the drowsiness prediction system of the third invention, in a speech analysis model, a typical model for cases where the heart rate decreases can be constructed by reconstructing the speech analysis model using data of a predetermined proportion from the upper probability range and the lower probability range.
[0016] As described above, according to the drowsiness prediction system of the third invention, it is possible to simply and reliably predict subsequent drowsiness from the current state for all people with a high probability based on the typical model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] [Figure 1] A system configuration diagram showing the overall configuration of the drowsiness prediction system according to an embodiment of the present invention. [Figure 2] An explanatory diagram showing processing contents of the speech analysis model in Fig. 1. [Figure 3] An explanatory diagram showing processing contents of the speech analysis model in Fig. 1. [Figure 4] An explanatory diagram showing processing contents of the speech analysis model in Fig. 1. [Figure 5A] An explanatory diagram showing processing contents of the speech analysis model in Fig. 1. [Figure 5B] An explanatory diagram showing processing contents of the speech analysis model in Fig. 1. [Figure 6] An explanatory diagram showing a modified example of the speech analysis model in Fig. 1. MODE FOR CARRYING OUT THE INVENTION
[0018] A drowsiness prediction system according to an embodiment of the present invention will be described below with reference to Fig. 1.
[0019] As shown in Fig. 1, the drowsiness prediction system is a system that predicts drowsiness by predicting a subsequent decrease in the heart rate of a target speaker to be estimated from speech uttered by the target speaker, and includes a speech analysis model (1), a speech data acquisition unit (2), and a drowsiness prediction output unit (3).
[0020] The speech analysis model 1 is constructed using binary logistic regression analysis with 12 frequency divisions as explanatory variables, using audio data of speech emitted by a speaker and data on the decrease in heart rate corresponding to the correct response to that speech, with heart rate decrease set to 1 and non-heart rate decrease set to 0. The training data, which consists of audio data of speech emitted by a speaker and the subsequent decrease in heart rate corresponding to the correct response to that speech, may be pre-stored in a speech database or similar, or they may be obtained as updated data via an external server or similar.
[0021] The voice data acquisition unit 2 acquires voice data of the target speaker to be estimated. For example, it may be configured with a microphone and convert the voice output of the target speaker's voice from the microphone into voice data (voice file data), or it may acquire recorded voice data (voice file data) from an external server or recording medium.
[0022] The drowsiness output unit 3 is a means for outputting a drowsiness prediction result from the estimated decrease in heart rate obtained by inputting the audio data (audio file data) acquired by the audio data acquisition unit 2 into the audio analysis model 1. The output means may be, for example, a display means such as a display that indicates that drowsiness is predicted (including displaying with color), or an audio output means such as a speaker that plays a warning or warning sound.
[0023] In the above configuration, the voice analysis model 1 is composed of hardware such as a CPU (Central Processing Unit), ROM (Read Only Memory), and RAM (Random Access Memory), and stores a program that executes the various processes described later in memory (not shown). By executing this program, it functions as a arithmetic unit (sequencer) for executing various processes. Alternatively, part or all of the voice analysis model 1 may be composed of other servers (external servers), and the drowsiness prediction system may be realized through distributed processing.
[0024] Next, with reference to Figures 2 to 6, we will explain the details of the voice analysis model 1, which is a characteristic of the drowsiness prediction system.
[0025] First, referring to Figure 2, the speech analysis model 1 extracted the frequencies of the speech speech data during periods of heart rate reduction (value < mean - σ) for each subject from the multiple speech data acquired by the speech data acquisition unit 2 (each speech data is labeled with a correct label indicating whether or not a decrease in heart rate occurred afterward). This was then compared with the frequencies of the speech speech data during periods of non-heart rate reduction. Specifically, 12 frequency divisions were used as explanatory variables, with heart rate reduction set to 1 and non-heart rate reduction to 0.
[0026] In this embodiment, heart rate data was collected from 16 subjects using a wearable device, and audio data was created from approximately 3 million records measured in beats per minute (bpm). The audio data was recorded once every hour via a mobile phone web application, and approximately 1,000 WAV files were created. Each audio recording contained approximately 3 seconds of data.
[0027] Next, the heart rate data from each subject was arranged in time series, and a dummy variable value of 1 was assigned to records that were more than one standard deviation below the mean, thus preparing dataset A.
[0028] Then, audio data was extracted from the audio data using Mel-frequency cepstrum coefficients (MFCCs), generating 12 data points for each record. These 12 data points were then arranged chronologically to form dataset B.
[0029] In practice, as shown in Figure 3, data augmentation was performed by adding variations to each audio record in dataset B. For example, variations included pitch changes, changes to the beginning of the audio, changes to the reverberation intensity, and the addition of white noise. The augmented dataset containing these variations was named dataset B'.
[0030] Then, using subject IDs and record generation times, unique keys were assigned to datasets A and B' as the datasets for analysis. Before integration, dataset A was advanced by 2 hours compared to dataset B' so that it could predict that heart rate would decrease within the next 2 hours based on the current voice data. The integrated dataset (dataset C) contained 233,800 records and 2,240 dummy variables with a value of 1.
[0031] As a result, as shown in Figure 4, differences were observed in the frequency characteristics of the two. It is thought that there is a relationship between spoken voice and heart rate. In other words, a model has been created that can determine the current decrease in heart rate (= possibility of sleepiness) based on the current voice.
[0032] Next, logistic regression was performed based on the following equation, with the dependent variable representing a decrease in heart rate as 1 and non-decreased heart rate as 0, and frequency as the independent variable.
[0033]
number
[0034] In the above equation, the left side is the odds ratio of becoming sleepy / not becoming sleepy, and the right side is the voice frequency at which the heart rate decreases (becoming sleepy) at the frequency (mfcc_1~12).
[0035] Here, the result of the logistic regression, as shown in Figure 5A, is an AUC of 0.72, demonstrating that a model has been created that can determine the current heart rate decrease (= sleepy state) based on the current voice. Note that AUC is the area under the curve of the logistic regression model, and the range of AUC is from 0 to 1, with a higher value indicating higher prediction accuracy.
[0036] Furthermore, in this model, by replacing the voice frequencies at the time of heart rate reduction (when sleepiness occurs) at the frequencies (mfcc_1~12) on the right-hand side of the above equation with voice frequencies within 2 hours prior to the time of heart rate reduction (when sleepiness occurs), the AUC is approximately 0.80, as shown in Figure 5B, and future heart rate reductions can be predicted with high accuracy.
[0037] Thus, the drowsiness prediction system of this embodiment can predict future drowsiness for any individual based on their current state with a higher probability. Therefore, it can contribute to reducing drowsiness-related accidents, for example, among drivers and high-risk workers.
[0038] In this embodiment, we have described a case where analysis is performed using all the data in the integrated dataset, but this is not the only case.
[0039] For example, as shown in Figure 6, the speech analysis model may be reconstructed using data with a predetermined ratio of high and low probabilities (in Figure 6, data with a ratio of 10% for both high and low probabilities). This makes it possible to construct a typical model for cases where the heart rate decreases.
[0040] Furthermore, in the speech analysis model, the model may be reconstructed using only a predetermined proportion of data with high or low probabilities. In this case, if the speech analysis model is reconstructed using only data with high probabilities, such a model can predict drowsiness with higher accuracy from the subsequent decrease in the speaker's heart rate. On the other hand, if the speech analysis model is reconstructed using only data with low probabilities, such a model can predict with higher accuracy that the speaker's heart rate will not decrease afterward, for example, that they will remain awake. [Explanation of symbols]
[0041] 1...Speech analysis model, 2...Speech data acquisition unit, 3...Drowsiness output unit.
Claims
1. A sleepiness prediction system that predicts a decrease in the heart rate of a target speaker based on the voice emitted by the target speaker, thereby predicting the target speaker's sleepiness, A speech analysis model was constructed using binary logistic regression analysis with the following parameters: speech data of the voice emitted by the speaker and data on the correct heart rate reduction for the said voice, with the explanatory variable representing heart rate reduction as 1 and non-heart rate reduction as 0, and using 12 frequency divisions as explanatory variables. A voice data acquisition unit that acquires voice data of the aforementioned speaker, The voice data acquired by the voice data acquisition unit is input to the voice analysis model, and the drowsiness output unit outputs the predicted drowsiness. A sleepiness prediction system characterized by having the following features.
2. In the sleepiness prediction system according to claim 1, A sleepiness prediction system characterized in that, in the aforementioned voice analysis model, the voice analysis model is reconstructed using only a predetermined proportion of data with high or low probabilities.
3. In the sleepiness prediction system according to claim 1, A sleepiness prediction system characterized in that the voice analysis model is reconstructed using a predetermined proportion of data with higher probabilities and a predetermined proportion of data with lower probabilities.
Citation Information
Patent Citations
Outrigger device of track traveling crane
JP1997002783A
Drowsiness determination device and program
JP2018192128A