Sleepiness prediction system

The system predicts sleepiness from voice data using a machine learning-based speech analysis model that compresses and restores frequency bands to extract features, enhancing prediction accuracy and suitability for mobile devices.

JP2026136892APending Publication Date: 2026-08-26RISK MEASUREMENT TECHNOLOGIES CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025022726
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2026-08-26
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Conventional sleepiness prediction systems require passengers to wear devices continuously to collect heart rate signals, making it difficult to predict sleepiness if biometric information is not collected or insufficiently collected.

Method used

A sleepiness prediction system that predicts heart rate decrease from voice data using a speech analysis model constructed through machine learning, where speech data is compressed and restored in stages to extract features, and performs machine learning on fully connected data to enhance prediction accuracy.

Benefits of technology

Enables accurate prediction of future sleepiness based on current voice data, suitable for mobile devices and reducing sleepiness-related accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136892000001_ABST
    Figure 2026136892000001_ABST
Patent Text Reader

Abstract

We provide a sleepiness prediction system that can predict future sleepiness levels for any individual based on their current state. [Solution] A system for predicting drowsiness by predicting the subsequent decrease in the heart rate of a target speaker based on the voice emitted by the target speaker, comprising a voice analysis model 1, a voice data acquisition unit 2, and a drowsiness prediction output unit 3.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a drowsiness prediction system that predicts the drowsiness of a target speaker by predicting a decrease in the heart rate of the target speaker from the voice emitted by the target speaker to be predicted.

Background Art

[0002] Conventionally, regarding the drowsiness of a passenger in a mobility, as shown in Patent Document 1 below, based on the time change of the passenger's heart rate interval and the determination criteria for the time change of the predetermined heart rate interval, a system for detecting the drowsiness of a passenger in a relatively short time is known from the relationship between the decrease in heart rate and drowsiness.

[0003] And, as a system for predicting drowsiness several hours in advance, as shown in Patent Document 2 below, an acquisition device that acquires a biological signal BS at the time of boarding of a passenger who has boarded a mobility such as a vehicle, an estimation unit that estimates the heart rate signal HS of the passenger based on the biological signal BS, a calculation unit that calculates drowsiness prediction data DD regarding the drowsiness omen of the passenger based on the heart rate signal HS, and the current drowsiness prediction data DDA of the passenger in the state of currently boarding the mobility and the past drowsiness prediction data DDB detected when the passenger boarded the mobility in the past. A passenger drowsiness information service providing device including a state estimation unit that estimates whether or not the drowsiness of the passenger has increased is known.

[0004] According to such a drowsiness information service providing device, when the drowsiness of a passenger increases, it may be affected by the passenger's past living conditions (for example, it is assumed that the sleep deprivation state has continued or physical fatigue has accumulated over the past few days). In such a case, even when drowsiness has not been detected, it is considered highly likely that the drowsiness of the passenger will increase. Thus, the drowsiness information service providing device estimates whether or not the drowsiness of the passenger tends to increase based on the past data of the passenger regarding drowsiness.

Prior Art Documents

Patent Documents

[0005] [Patent Document 1] Japanese Patent Publication No. 2018-192128 [Patent Document 2] Japanese Patent Publication No. 2024-092783 [Overview of the project] [Problems that the invention aims to solve]

[0006] However, conventional sleepiness information service devices assume that all passengers wear a device to acquire heart rate signals beforehand and go about their daily lives with that device. This presents a problem in that it is difficult to predict sleepiness if biometric information is not collected by such a device or if the collection is insufficient.

[0007] In view of the above circumstances, the present invention aims to provide a sleepiness prediction system that can predict future sleepiness for any person based on their current state. [Means for solving the problem]

[0008] The sleepiness prediction system of the first invention is A sleepiness prediction system that predicts a decrease in the heart rate of a target speaker based on the voice emitted by the target speaker, thereby predicting the target speaker's sleepiness, A speech analysis model constructed by machine learning processing using speech data of the voice emitted by the speaker and the correct heart rate decrease for the said voice as training data, A voice data acquisition unit that acquires voice data of the aforementioned speaker, A sleepiness output unit outputs predicted sleepiness by inputting the voice data acquired by the voice data acquisition unit into the voice analysis model. Equipped with, The aforementioned speech analysis model is characterized by performing the machine learning process using feature audio data obtained by compressing the speech data by dividing the frequency band so that the number of divisions decreases in stages, and then restoring the compressed speech data so that the number of divisions increases in stages, thereby extracting features from the speech data.

[0009] According to the sleepiness prediction system of the first invention, when constructing a speech analysis model by machine learning processing using speech data of speech emitted by a speaker and the correct heart rate decrease of the speech indicating whether or not the speaker's heart rate subsequently decreased as training data, the speech data is compressed by dividing the frequency band so that the number of divisions decreases in stages, and then the compressed speech data is restored so that the number of divisions increases in stages. By using this feature-extracted speech data, it is possible to extract the features necessary for sleepiness prediction.

[0010] Thus, according to the sleepiness prediction system of the first invention, it is possible to predict future sleepiness for any person based on their current state.

[0011] The drowsiness prediction system of the second invention is, in the first invention, The aforementioned speech analysis model is characterized in that, when performing the machine learning process using the feature speech data obtained by repeatedly performing the feature extraction process, it excludes a portion of the compressed speech data or the speech data to be restored, and then performs the machine learning process using the fully concatenated speech data.

[0012] According to the drowsiness prediction system of the second invention, when performing machine learning processing on feature audio data obtained by repeatedly performing a feature extraction process on audio data that involves compression and restoration by dividing the frequency band, it is possible to intentionally remove some of the data and then train the system again by performing the machine learning processing on fully connected audio data after excluding a portion of the compressed audio data or the audio data to be restored. This makes it possible to further increase the probability of predicting drowsiness.

[0013] Thus, according to the drowsiness prediction system of the second invention, it is possible to predict the subsequent drowsiness of any person with a higher probability from the current state.

Brief Description of the Drawings

[0014] [Figure 1] System configuration diagram showing the overall configuration of the drowsiness prediction system according to an embodiment of the present invention. [Figure 2] Explanatory diagram showing the processing content of the voice analysis model in FIG. 1. [Figure 3] Explanatory diagram showing the processing content of the voice analysis model in FIG. 1. [Figure 4] Explanatory diagram showing the processing content of the voice analysis model in FIG. 1.

Mode for Carrying Out the Invention

[0015] A drowsiness prediction system according to an embodiment of the present invention will be described below with reference to FIG. 1.

[0016] As shown in FIG. 1, the drowsiness prediction system is a system that predicts drowsiness by predicting a subsequent decrease in the heart rate of a target speaker from the voice emitted from the target speaker to be estimated, and includes a voice analysis model 1, a voice data acquisition unit 2, and a drowsiness prediction output unit 3.

[0017] The voice analysis model 1 is constructed by machine learning processing using the voice data of the voice emitted from the speaker and the correct decrease in the heart rate of the voice as teacher data. Note that the voice data of the voice emitted from the speaker serving as teacher data and the subsequent correct decrease in the heart rate may be those stored in advance in a voice database or the like, or those obtained by being updated each time via an external server or the like.

[0018] The voice data acquisition unit 2 acquires voice data of the target speaker to be estimated. For example, it is composed of a microphone, and in addition to converting the voice output of the target speaker output from the microphone into voice data (voice file data), it may also acquire recorded voice data (voice file data) from an external server, a recording medium, or the like.

[0019] The drowsiness output unit 3 is a means for outputting a drowsiness prediction result from the decrease in heart rate estimated by inputting the voice data (voice file data) acquired by the voice data acquisition unit 2 into the voice analysis model 1. The output means may be, for example, display means such as a display indicating that drowsiness is predicted (including the case of displaying in color), or voice output means such as a speaker for emitting a warning or a warning sound.

[0020] In the above configuration, the voice analysis model 1 is composed of hardware such as a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory), for example, and stores and holds a program for executing various processes described later in a memory (not shown). By executing the program, it functions as an arithmetic unit (sequencer) for executing various processes. Also, part or all of the voice analysis model 1 may be composed of another server (external server) to realize the drowsiness prediction system by distributed processing.

[0021] Next, referring to FIG. 2, the details of the voice analysis model 1, which is a feature of the drowsiness prediction system, will be described.

[0022] The voice analysis model 1 performs feature extraction processing on a plurality of voice data (each of the plurality of voice data is labeled with the presence or absence of a correct heart rate decrease for which it has been determined whether a subsequent heart rate decrease has occurred) acquired by the voice data acquisition unit 2, and then performs machine learning processing using this as teacher data.

[0023] In this embodiment, the machine learning process is performed by training a neural network for emotion estimation using deep learning, and comparing the outputs of multiple channels of the neural network with the correct heart rate decrease (labels indicating whether or not there is a decrease in heart rate). However, the embodiment is not limited to this, and machine learning methods other than deep learning may be employed.

[0024] Furthermore, regarding the correct answer for decreased heart rate (label indicating whether or not there was a decrease in heart rate), while most people feel sleepy at night and their heart rate decreases, here we focus on the fact that heart rate also decreases during the day, and by focusing on the mean and standard deviation (SD), we define decreased heart rate as a heart rate value that is 1 SD or more lower than the mean.

[0025] Here, the feature extraction process involves dividing the audio data into frequency bands so that the number of divisions decreases in stages (as shown in the figure, 128 → 96 → 64 divisions) and compressing it, and then restoring the compressed audio data so that the number of divisions increases in stages (as shown in the figure, 64 → 96 → 128 divisions) to extract the features of the audio data.

[0026] In this way, by compressing audio data by dividing the frequency band so that the number of divisions decreases in stages, and then restoring the compressed audio data so that the number of divisions increases in stages, it is possible to extract features necessary for emotion estimation by using feature audio data that has undergone such feature extraction processing. This differs from noise reduction by filtering.

[0027] Here, it is preferable to perform the feature extraction process multiple times before performing the machine learning process. In these multiple feature extraction processes, a portion of the compressed audio data or the audio data to be restored is excluded (Dropout in the figure), and the machine learning process is performed on the fully connected audio data (Fully connected in the figure).

[0028] In this way, by repeatedly performing feature extraction on audio data, which involves compression and restoration by dividing the audio data into frequency bands, and then applying machine learning processing to this feature audio data, the accuracy of predicting heart rate reduction can be improved. Furthermore, by excluding some of the compressed or reconstructed audio data and then performing machine learning processing on the fully connected audio data, it is possible to intentionally remove some data and then train the model again, thereby increasing the probability of emotion estimation.

[0029] In this embodiment, heart rate (HR, beats per minute (bpm)) data was collected from 12 subjects using a wearable device. Approximately 3 million data points were collected. Audio data was collected once every hour via a mobile phone web application, creating approximately 1,000 files. Each audio data point lasted 3 seconds.

[0030] Next, the HR data for each subject was arranged in time series, and a dummy variable (0 or 1) was assigned depending on whether the HR value was less than or equal to one standard deviation.

[0031] Audio features were extracted from audio data using a logarithmic Mel spectrogram (LMS), creating 128 data points per record, which were then arranged in a time series. This was used as a deep learning model to construct speech analysis model 1.

[0032] Furthermore, in this embodiment, as shown in Figure 3, data augmentation is applied to the audio data to construct the model. Specifically, "pitch (voice tone)", "start shift", "reverberation strength", and "white noise" were used as data augmentation.

[0033] Furthermore, in this embodiment, as shown in Figure 4, the audio data (audio features) was integrated with dummy variables as the analysis dataset. There was a two-hour difference between the audio features and the dummy variables, and the dummy variables were advanced relative to the audio features to predict the likelihood of becoming sleepy within two hours. Finally, an analysis dataset with 267,900 records containing 2,400 dummy variables was obtained and used.

[0034] As a result, voice analysis model 1 was successfully constructed as a deep learning model that predicts drowsiness with 99.6% accuracy using voice data. This model can predict the decrease in heart rate associated with the onset of drowsiness (PNS dominance) and has stable training accuracy (99.05%~99.06%) and validation accuracy (98.70%~98.74%). This model also works with low-quality audio, making it suitable for use on mobile devices.

[0035] Thus, the drowsiness prediction system of this embodiment can predict future drowsiness for any individual based on their current state with a higher probability. Therefore, it can contribute to reducing drowsiness-related accidents, for example, among drivers and high-risk workers. [Explanation of Symbols]

[0036] 1...Speech analysis model, 2...Speech data acquisition unit, 3...Sleepiness output unit.

Claims

1. A sleepiness prediction system that predicts a decrease in the heart rate of a target speaker based on the voice emitted by the target speaker, thereby predicting the target speaker's sleepiness, A speech analysis model constructed by machine learning processing using speech data of the voice emitted by the speaker and the correct heart rate decrease for the said voice as training data, A voice data acquisition unit that acquires voice data of the aforementioned speaker, The voice data acquired by the voice data acquisition unit is input to the voice analysis model, and the drowsiness output unit outputs the predicted drowsiness. Equipped with, The drowsiness prediction system is characterized in that the voice analysis model performs a feature extraction process to extract features from the voice data by compressing the voice data by dividing the frequency band so that the number of divisions decreases in stages, and then restoring the compressed voice data so that the number of divisions increases in stages, thereby performing the machine learning process on the feature voice data.

2. In the sleepiness prediction system according to claim 1, The drowsiness prediction system is characterized in that, when performing the machine learning process on the feature audio data obtained by repeatedly performing the feature extraction process on the voice analysis model, it excludes a portion of the compressed audio data or the audio data to be restored, and then performs the machine learning process on the fully concatenated audio data.

Citation Information

Patent Citations

  • Outrigger device of track traveling crane

    JP1997002783A

  • Drowsiness determination device and program

    JP2018192128A