Information processing device, information processing method, information processing system, and information processing program

JPWO2024116254A5Active Publication Date: 2025-11-10PST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024560999
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-28
Filing Date
2022-11-28
Publication Date
2025-11-10
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

Existing devices for estimating psychiatric or nervous system diseases from voice data have limitations in accuracy due to the reliance on static acoustic parameters, lacking dynamic analysis of voice patterns.

Method used

The implementation of a dynamic time warping method on preprocessed voice data, which extracts stable central portions, aligns and samples data to match reference patterns, and applies expansion/contraction processing to enhance feature extraction, thereby improving disease symptom estimation accuracy.

Benefits of technology

This approach allows for precise estimation of disease symptoms by calculating a score based on dynamic time warping results, significantly enhancing accuracy and robustness against variations in recording conditions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided is an information processing device that acquires speech data, which is time-series data of speech uttered by a user, and generates preprocessed speech data representing the data that, from out of the acquired speech data, is no earlier than a first time from the start point of the speech data and no later than a second time from the end of the speech data. Additionally, this information processing device generates processing result data by applying dynamic time warping to the preprocessed speech data that is generated. The information processing device calculates a score representing the degree to which the user has a predetermined disease or symptom on the basis of the generated processing result data, and estimates whether or not the user has the predetermined disease or symptom on the basis of the calculated score.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, information processing system, and information processing program

[0001] The disclosed technology relates to an information processing device, an information processing method, an information processing system, and an information processing program.

[0002] International Publication No. 2020 / 013296 discloses a device for predicting whether a user has a psychiatric or neurological disorder. This device calculates various acoustic parameters from a user's voice data and uses the acoustic parameters to predict whether the user has a psychiatric or neurological disorder.

[0003] Although the device disclosed in the above-mentioned International Publication No. 2020 / 013296 estimates diseases using acoustic parameters calculated from voice data, there is room for improvement in terms of accuracy.

[0004] The disclosed technology has been developed in consideration of the above circumstances, and provides an information processing device, an information processing method, an information processing system, and an information processing program that can accurately estimate whether a user has a specified disease or symptom by applying a dynamic time warping method to voice data, which is time-series data of voice spoken by a user.

[0005] In order to achieve the above object, a first aspect of the present disclosure is an information processing device including: an acquisition unit that acquires voice data, which is time-series data of voice uttered by a user; a processing unit that generates preprocessed voice data representing data from the voice data acquired by the acquisition unit that is data that is a first time or later from the start point of the voice data and that is a second time or earlier than the end point of the voice data; a generation unit that generates processing result data by applying dynamic time warping to the preprocessed voice data generated by the processing unit; a calculation unit that calculates a score that represents the degree to which the user has a specified disease or symptom based on the processing result data generated by the generation unit; and an estimation unit that estimates whether the user has a specified disease or symptom based on the score calculated by the calculation unit.

[0006] A second aspect of the present disclosure is an information processing method that causes a computer to perform the following processes: acquire voice data, which is time-series data of voice uttered by a user; generate pre-processed voice data representing data from the acquired voice data that is a first time or later from the start point of the voice data and a second time or earlier than the end point of the voice data; generate processing result data by applying dynamic time warping to the generated pre-processed voice data; calculate a score representing the degree to which the user has a specified disease or symptom based on the generated processing result data; and estimate whether the user has the specified disease or symptom based on the calculated score.

[0007] A third aspect of the present disclosure is an information processing program for causing a computer to execute a process of acquiring voice data, which is time-series data of voice uttered by a user, generating preprocessed voice data representing data from the acquired voice data that is a first time or later from the start point of the voice data and a second time or earlier than the end point of the voice data, generating processing result data by applying dynamic time warping to the generated preprocessed voice data, calculating a score representing the degree to which the user has a specified disease or symptom based on the generated processing result data, and estimating whether or not the user has the specified disease or symptom based on the calculated score.

[0008] According to the disclosed technology, by applying a dynamic time warping method to voice data, which is time series data of voice uttered by a user, it is possible to accurately estimate whether a user has a specified disease or a specified symptom.

[0009] FIG. 1 is a diagram illustrating an example of a schematic configuration of an information processing system according to a first embodiment. FIG. 2 is a diagram for explaining an overview of the first embodiment. FIG. 3 is a diagram illustrating audio data for a predetermined period. FIG. 4 is a diagram illustrating shift processing for audio data. FIG. 5 is a diagram illustrating sampling processing for audio data. FIG. 6 is a diagram illustrating an example of a usage form of the information processing system according to the first embodiment. FIG. 7 is a diagram illustrating an example of a computer constituting an information processing device. FIG. 8 is a diagram illustrating an example of processing executed by the information processing device according to the first embodiment. FIG. 9 is a diagram illustrating an overview of a second embodiment. FIG. 10 is a diagram illustrating an example of a usage form of the information processing system according to the second embodiment. FIG. 11 is a diagram illustrating an example of an experimental result of an example. FIG. 12 is a diagram illustrating an experimental result of an example. FIG. 13 is a diagram illustrating an experimental result of an example. FIG. 14 is a diagram illustrating an experimental result of an example. FIG. 15 is a diagram illustrating an experimental result of an example.

[0010] Hereinafter, embodiments of the disclosed technology will be described in detail with reference to the drawings.

[0011] <Information Processing System of First Embodiment>

[0012] 1 shows an information processing system 10 according to the first embodiment. As shown in FIG. 1, the information processing system 10 according to the first embodiment includes a microphone 12, an information processing device 14, and a display device 16.

[0013] The information processing system 10 estimates whether the user has a predetermined disease or a predetermined symptom (hereinafter simply referred to as "disease, etc.") based on the user's voice collected by the microphone 12. Note that the information processing system 10 of this embodiment estimates whether the user has a psychiatric disease or a neurological disease, or a mental disorder symptom or a cognitive dysfunction symptom, as examples of the predetermined disease or the predetermined symptom.

[0014] The information processing device 14 of the information processing system 10 of the first embodiment performs predetermined preprocessing on voice data, which is time-series data of voice uttered by a user, to generate preprocessed data. Then, the information processing device 14 determines whether the user has a disease or the like based on the result of applying dynamic time warping to the preprocessed data.

[0015] In the dynamic time warping method, the distance between one time series data and another time series data is calculated. In this embodiment, the processing result data obtained by the dynamic time warping method is used to estimate whether or not the user has a disease.

[0016] The specific details will be explained below.

[0017] 1, the information processing device 14 functionally includes an acquisition unit 20, a voice data storage unit 22, a reference data storage unit 24, a processing unit 26, a generation unit 28, a calculation unit 30, an estimation unit 32, and an output unit 34. The information processing device 14 is realized by a computer as described below.

[0018] The acquisition unit 20 acquires voice data, which is time-series data of voices uttered by the user, and stores the voice data in the voice data storage unit 22.

[0019] The voice data storage unit 22 stores the voice data acquired by the acquisition unit 20 .

[0020] The reference data storage unit 24 stores voice data of reference users who are known to have or not have a disease or the like.

[0021] The processing unit 26 reads out the audio data stored in the audio data storage unit 22. The processing unit 26 then performs predetermined preprocessing on the audio data to generate preprocessed audio data. The method for generating preprocessed audio data will be described in detail below. Figure 2 shows a diagram for explaining preprocessed audio data.

[0022] (Extraction of the central part of the audio data)

[0023] When estimating whether a user has a disease or the like based on voice data uttered by the user, it is preferable to use voice data in which the user's voice is stable.

[0024] In this regard, since the initial part of the time-series data represented by the voice data is data at the time when the user starts speaking, it is often undesirable to use the data at that time for inferring diseases, etc. For example, if a user suddenly starts speaking after being silent, it is expected that the user's voice will become unstable, causing the voice to become hoarse or the volume to become low. Even if such data is used for inferring diseases, etc., it is expected that accurate results will not be obtained.

[0025] Furthermore, it is often not desirable to use the portion of the time series data represented by the voice data near the end point for estimating diseases, etc. For example, when a user utters a long pronunciation, it is expected that the user may run out of breath and be unable to continue speaking, or the pronunciation of the end of the word may become unclear.

[0026] Therefore, the processing unit 26 of the information processing device 14 of this embodiment extracts central data from the audio data, which is time-series data.

[0027] 2, the processing unit 26 generates data D2 representing data from the voice data D1 that is a first time T1 or later from the start point of the voice data D1 and that is a second time T2 or earlier than the end point of the voice data D1. The data D2 corresponds to a time interval T3 of the voice data D1. This generates data from the central portion where the user's voice is stable.

[0028] (Extraction of Data for a Predetermined Cycle) Furthermore, the processing unit 26 extracts data for a predetermined cycle from the extracted central portion of data. Fig. 3 shows a diagram for explaining data for a predetermined cycle. As shown in Fig. 3, the audio data Df is time-series data, and a predetermined signal is repeated. For example, in the example shown in Fig. 3, a similar signal waveform is repeated for each time interval T.

[0029] As described below, when estimating whether a user has a disease or the like, voice data uttered by the user whose disease or the like is to be estimated may be compared with voice data of a reference user whose disease or the like is known to be present. Therefore, it is preferable that a predetermined period of data extracted from the voice data uttered by the user whose disease or the like is to be estimated is aligned with a predetermined period of data in the voice data of the reference user. For this reason, for example, the processing unit 26 extracts data of a predetermined period that is the same as the period of the voice data of the reference user from the extracted central portion of the data. This predetermined period may be set, for example, in advance. Alternatively, the predetermined period may be varied depending on the type of data, for example.

[0030] (Extraction of Data by Shifting Along the Time Axis) Next, the processing unit 26 shifts the extracted data for a predetermined period along the time axis. FIG. 4 is a diagram for explaining the shifting of data along the time axis. As shown in FIG. 4, the voice data Ds contains a repetition of a signal with a period Ts, and the reference user's voice data D Ref has a period T Ref In this case, as shown in FIG. 4, the start portion P1 of the extracted voice data Ds and the reference user's voice data D Ref If the voice data Ds and the reference user's voice data D Ref Even if the audio data D and the audio data D are similar to each other, the audio data D and the audio data D are calculated by the dynamic time warping method. Ref It is possible that the value representing the distance between and may become large.

[0031] Therefore, the processing unit 26 shifts the extracted data for a predetermined period along the time axis. For example, the processing unit 26 shifts the data for a predetermined period shown in FIG. 4 along the time axis indicated by the arrow S by a predetermined time. Note that the amount of shift for this predetermined time is set in advance, for example. Alternatively, the amount of shift may be changed depending on the type of data, for example.

[0032] (Extraction of data by sampling at a predetermined sampling rate) Next, the processing unit 26 extracts sampled data obtained by sampling from the data shifted along the time axis. As described above, in this embodiment, when estimating whether a user has a disease or the like, voice data uttered by the user whose disease or the like is to be estimated may be compared with voice data of a reference user whose disease or the like is known to be present. Therefore, it is preferable that the sampling rate for the voice data uttered by the user whose disease or the like is to be estimated and the sampling rate for the voice data of the reference user are the same.

[0033] For example, as shown in FIG. 5, sampled data D obtained by sampling audio data D is A , D B In this case, the sampled data D generated by the sampling rate A is A and sampling data D generated at sampling rate B. B When the distance between is calculated using the dynamic time warping method, a value representing a predetermined distance is calculated even though the original audio data D is the same.

[0034] For example, the processing unit 26 generates sampled data extracted at the same sampling rate as the sampling rate of the reference user's voice data. This sampling rate is set in advance. Alternatively, the sampling rate may be changed depending on the type of data. For example, 200 sampling points per cycle of data are extracted from a predetermined cycle of data.

[0035] (Expansion and contraction of data in the time axis direction) Next, the processing unit 26 performs expansion and contraction processing in the time axis direction on the sampling data obtained by sampling from the voice data. As described above, in the present embodiment, when estimating whether or not a user has a disease, etc., the voice data uttered by the user whose disease, etc. is to be estimated may be compared with the voice data of a reference user whose disease, etc. is known to be present or absent.

[0036] Therefore, it is preferable that the intervals along the time axis of the voice data uttered by the user whose disease or the like is to be estimated and the intervals along the time axis of the voice data of the reference user are aligned. For this reason, for example, the processing unit 26 performs a predetermined expansion / contraction process along the time axis on the data D3 shown in FIG. 2. The predetermined expansion / contraction process method is set in advance. Alternatively, for example, the expansion / contraction process method may be changed depending on the type of data.

[0037] (Expansion and contraction of data in the amplitude direction) Next, the processing unit 26 performs expansion and contraction processing in the amplitude direction on the data that has been subjected to expansion and contraction processing in the time axis direction. As described above, in this embodiment, when estimating whether or not a user has a disease, etc., voice data uttered by the user whose disease, etc. is to be estimated may be compared with voice data of a reference user whose disease, etc. is known to be present.

[0038] Therefore, it is preferable that the amplitude of the voice data uttered by the user whose disease or the like is to be estimated and the amplitude of the voice data of the reference user are aligned. For this reason, for example, the processing unit 26 performs a predetermined expansion / contraction process in the amplitude direction on the data D4 shown in FIG. 2. The predetermined expansion / contraction process method is set in advance. Alternatively, for example, the expansion / contraction process method may be changed depending on the type of data.

[0039] The processing unit 26 performs the above-described multiple preprocessing processes on the audio data to generate preprocessed audio data.

[0040] The generation unit 28 generates processing result data by applying dynamic time warping to the preprocessed audio data generated by the processing unit 26. The processing result data obtained by applying dynamic time warping is calculated as a distance matrix representing the distance between each point of one time series data and each point of another time series data.

[0041] Specifically, the generation unit 28 reads out the reference user's voice data stored in the reference data storage unit 24. Then, as shown in FIG. 2, the generation unit 28 generates the preprocessed voice data D5 and the reference user's voice data D Ref By applying the dynamic time warping method to the preprocessed voice data D5 and the reference user voice data D Ref The processing result data representing the distance between the user and the reference user may also be subjected to the above-described preprocessing.

[0042] The generation unit 28 may generate the processing result data using only the preprocessed audio data. For example, the generation unit 28 may apply dynamic time warping to first audio data representing data in a first time interval within the preprocessed audio data and second audio data representing data in a second time interval within the preprocessed audio data, thereby generating processing result data representing the distance between the first audio data and the second audio data.

[0043] More specifically, for example, as shown in FIG. 2, the generation unit 28 generates processing result data representing the distance between the first audio data D5-1 and the second audio data D5-2 by applying dynamic time warping to first audio data D5-1 representing data in a first time interval within the preprocessed audio data D5 and second audio data D5-1 representing data in a second time interval within the preprocessed audio data D5.

[0044] Next, as shown in Figure 2, the generation unit 28 applies dynamic time warping to second audio data D5-2 representing data in a second time interval within the preprocessed audio data D5, and third audio data D5-3 representing data in a third time interval within the preprocessed audio data D5, thereby generating processing result data representing the distance between the second audio data D5-2 and the third audio data D5-3.

[0045] Furthermore, the generation unit 28 generates processing result data representing the distance between the first audio data D5-1 and the third audio data D5-3 by applying dynamic time warping to the first audio data D5-1 and the third audio data D5-3. In this way, the generation unit 28 generates processing result data for each pair of audio data D5-1 to D5-9 within a predetermined time period.

[0046] The calculation unit 30, which will be described later, may calculate a score representing the degree to which the user has a disease or the like based on the processing result data generated in this manner using only the preprocessed voice data D5.

[0047] The calculation unit 30 calculates a score representing the degree to which the user has a disease or the like, based on the processing result data generated by the generation unit 28. For example, the calculation unit 30 uses the average value, maximum value, minimum value, standard deviation, and median value of each element of the distance matrix generated by the generation unit 28 to calculate a score representing the degree to which the user has a predetermined disease or symptom, using a known method.

[0048] The estimation unit 32 estimates whether or not the user has a disease, etc., based on the score calculated by the calculation unit 30. For example, if the score is equal to or greater than a predetermined threshold, the estimation unit 32 estimates that the user has a disease, etc., and if the score is less than the predetermined threshold, the estimation unit 32 estimates that the user does not have a disease, etc.

[0049] The output unit 34 outputs the estimation result estimated by the estimation unit 32. Note that the output unit 34 may output the score itself as the estimation result.

[0050] The display device 16 displays the estimation result output from the estimation unit 32 .

[0051] A medical professional or a user operating the information processing device 14 checks the estimation results output from the display device 16 and confirms what disease or symptom the user is likely to have.

[0052] The information processing system 10 of this embodiment is assumed to be used, for example, under the circumstances shown in FIG.

[0053] In the example of FIG. 6 , a medical professional H, such as a doctor, holds a tablet terminal, which is an example of the information processing system 10. The medical professional H uses a microphone (not shown) provided on the tablet terminal to collect voice data from a user U, who is a subject. The tablet terminal then estimates whether or not the user U has a disease or symptom based on the voice data of the user U, and outputs the estimation result to a display unit (not shown). The medical professional H determines whether or not the user U has a disease or symptom by referring to the estimation result displayed on the display unit (not shown) of the tablet terminal.

[0054] The information processing device 14 can be realized, for example, by a computer 50 shown in FIG. 7 . The computer 50 includes a CPU 51, a memory 52 serving as a temporary storage area, and a non-volatile storage unit 53. The computer 50 also includes an input / output interface (I / F) 54 to which external devices, output devices, etc. are connected, and a read / write (R / W) unit 55 that controls reading and writing of data from and to a recording medium. The computer 50 also includes a network I / F 56 that is connected to a network such as the Internet. The CPU 51, memory 52, storage unit 53, input / output I / F 54, R / W unit 55, and network I / F 56 are connected to one another via a bus 57.

[0055] The storage unit 53 can be realized by a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage unit 53 as a storage medium stores a program for operating the computer 50. The CPU 51 reads the program from the storage unit 53, loads it into the memory 52, and sequentially executes the processes contained in the program.

[0056] [Operation of the information processing system of the first embodiment]

[0057] Next, a specific operation of the information processing system 10 according to the first embodiment will be described. The information processing device 14 of the information processing system 10 executes the processes shown in FIG.

[0058] First, in step S100, the acquisition unit 20 acquires voice data of the user collected by the microphone 12. Then, the acquisition unit 20 stores the voice data in the voice data storage unit 22.

[0059] Next, in step S102, the processing unit 26 reads out the audio data stored in the audio data storage unit 22. Then, the processing unit 26 extracts audio data of a central portion, which is data within a predetermined time period, from the audio data.

[0060] In step S104, the processing unit 26 extracts data for a predetermined period from the central portion of the audio data acquired in step S102.

[0061] In step S105, the processing unit 26 executes a shift process on the audio data for the predetermined period acquired in step S104.

[0062] In step S106, the processing unit 26 generates sampling data by performing a predetermined sampling process on the data for a predetermined period that has been subjected to the shift process and obtained in step S105.

[0063] In step S108, the processing unit 26 performs expansion / contraction processing in the amplitude direction on the sampling data generated in step S106.

[0064] In step S110, the processing unit 26 performs expansion / compression processing in the time axis direction on the sampling data that has been expanded / compressed in the amplitude direction and obtained in step S108.

[0065] By performing the processes in steps S102 to S110, preprocessed audio data is generated by performing preprocessing on the audio data.

[0066] In step S112, the estimation unit 32 generates processing result data representing the distance between the preprocessed voice data and the reference user's voice data stored in the reference data storage unit 24 by applying dynamic time warping to the preprocessed voice data and the reference user's voice data stored in the reference data storage unit 24.

[0067] As the reference users, reference users who have a predetermined disease or the like and reference users who do not have the predetermined disease are set.

[0068] For this reason, for example, the estimation unit 32 generates processing result data between the preprocessed voice data and the voice data of a reference user who has a disease or the like, which is stored in the reference data storage unit 24. Alternatively, for example, the estimation unit 32 generates processing result data between the preprocessed voice data and the voice data of a reference user who does not have a disease or the like, which is stored in the reference data storage unit 24.

[0069] In step S114, the calculation unit 30 calculates a score representing the degree to which the user has a disease or the like, based on the processing result data generated in step S112. Note that the score may take a larger value, for example, the higher the degree to which the user has a disease or the like. Alternatively, the score may take a smaller value, for example, the higher the degree to which the user has a disease or the like.

[0070] For example, consider a case where the score increases as the degree to which the user has a disease, etc. In this case, when the distance between the preprocessed voice data and the voice data of the reference user who has a disease, etc. is small, the calculation unit 30 calculates the score so that the degree to which the user has a disease, etc. is high. On the other hand, when the distance between the preprocessed voice data and the voice data of the reference user who has a disease, etc. is large, the calculation unit 30 calculates the score so that the degree to which the user has a disease, etc. is low.

[0071] Furthermore, for example, when the distance between the preprocessed voice data and the voice data of the reference user who does not have a disease, etc. is small, the calculation unit 30 calculates the score so that the degree to which the user has a disease, etc. is low. On the other hand, when the distance between the preprocessed voice data and the voice data of the reference user who does not have a disease, etc. is large, the calculation unit 30 calculates the score so that the degree to which the user has a disease, etc. is high.

[0072] In step S116, the estimation unit 32 estimates whether or not the user has a disease, etc., based on the score calculated in step S114. For example, if the score is equal to or greater than a predetermined threshold, the estimation unit 32 estimates that the user has a disease, etc., and if the score is less than the predetermined threshold, the estimation unit 32 estimates that the user does not have a disease, etc.

[0073] In addition, the estimation unit 32 may estimate which disease the user has based on the processing result data for each of the voice data of the reference user having disease A, the voice data of the reference user having disease B, and the voice data of the reference user having disease C.

[0074] In step S118, the output unit 34 outputs the estimation result obtained in step S116.

[0075] The display device 16 displays the estimation results output from the output unit 34. A medical professional or a user operating the information processing device 14 checks the estimation results output from the display device 16 and confirms what disease or symptom the user may have.

[0076] As described above, the information processing system 10 of the first embodiment acquires voice data, which is time-series data of voice uttered by a user, and generates preprocessed data. The information processing device 14 then applies dynamic time warping to the generated preprocessed voice data to generate processed result data, and calculates a score representing the degree to which the user has a predetermined disease or symptom based on the generated processed result data. The information processing device 14 then estimates whether the user has a predetermined disease or symptom based on the calculated score. By applying dynamic time warping to the voice data, which is time-series data of voice uttered by a user, it is possible to accurately estimate whether the user has a predetermined disease or symptom.

[0077] The preprocessed voice data is a central portion of the acquired voice data that is a first hour or later from the start point of the voice data and a second hour or earlier from the end point of the voice data. By using the central portion of the voice data as the preprocessed voice data, it is possible to use the stable central portion of the voice uttered by the user to accurately estimate whether the user has a predetermined disease or a predetermined symptom.

[0078] The preprocessed voice data is also data for a predetermined period. The preprocessed voice data is also data obtained by shifting data in the time axis direction. The preprocessed voice data is also data obtained by performing a predetermined sampling process. The preprocessed voice data is also data obtained by performing a process to expand or contract the voice data in the time axis direction. The preprocessed voice data is also data obtained by performing a process to expand or contract the voice data in the amplitude direction. By performing these preprocessing processes on the voice data, it is possible to format the voice data in a manner suitable for estimating a disease, etc., and it is possible to accurately estimate whether or not the user has a disease, etc.

[0079] <Information Processing System of Second Embodiment>

[0080] Next, a second embodiment will be described. Note that, among the configuration of the information processing system of the second embodiment, the parts having the same configuration as those of the first embodiment will be assigned the same reference numerals and the description thereof will be omitted.

[0081] Fig. 9 shows an information processing system 310 according to the second embodiment. As shown in Fig. 9, the information processing system 310 includes a user terminal 18 and an information processing device 314. The information processing device 314 further includes a communication unit 36.

[0082] The information processing device 314 of the information processing system 310 estimates whether the user has a disease or the like based on the user's voice collected by the microphone 12 provided in the user terminal 18.

[0083] The information processing system 310 of the second embodiment is expected to be used under the conditions shown in FIGS. 10 and 11, for example.

[0084] 10 , a medical professional H such as a doctor operates an information processing device 314, and a user U, who is a subject, operates a user terminal 18. The user U collects his or her own voice data "XXXX" using the microphone 12 of the user terminal 18 that he or she operates. The user terminal 18 then transmits the voice data to the information processing device 314 via a network 19 such as the Internet.

[0085] The information processing device 314 receives the voice data "XXX" of the user U transmitted from the user terminal 18. Then, the information processing device 314 estimates whether or not the user U has any disease or symptom based on the received voice data, and outputs the estimation result to the display unit 315 of the information processing device 314. The medical worker H refers to the estimation result displayed on the display unit 315 of the information processing device 314 and determines whether or not the user U has any disease or symptom.

[0086] On the other hand, in the example of FIG. 11 , a user U, who is a subject, collects his or her own voice data using the microphone 12 of a user terminal 18 that the user operates. The user terminal 18 then transmits the voice data to an information processing device 314 via a network 19, such as the Internet. The information processing device 314 receives the user U's voice data transmitted from the user terminal 18. The information processing device 314 then estimates whether or not the user U has any disease or symptom based on the received voice data, and transmits the estimation result to the user terminal 18. The user terminal 18 receives the estimation result transmitted from the information processing device 14 and displays the estimation result on a display unit (not shown). The user checks the estimation result and confirms what disease or symptom the user is likely to have.

[0087] The information processing device 314 executes the same information processing routine as that shown in FIG.

[0088] As described above, the information processing system of the second embodiment can estimate whether a user has a psychiatric disorder, a neurological disorder, or symptoms thereof using an information processing device 214 installed on the cloud.

[0089] Next, an example will be described. In this example, experimental results regarding the effect of the pretreatment described in this embodiment will be shown.

[0090] FIG. 12 is a graph plotting speech data obtained from subjects evaluated as depressed patients (represented by squares in FIG. 12 ) and speech data obtained from subjects evaluated as healthy individuals (represented by circles in FIG. 12 ). FIG. 12 shows data obtained using the preprocessing and DTW of this embodiment. The horizontal axis dist2 of the graph in FIG. 12 represents the distance from the average reference for healthy individuals, and the vertical axis dist3 represents the distance from the average reference for depressed individuals. As shown in FIG. 12 , the speech data of depressed individuals, represented by squares, tends to be farther from the average reference for healthy individuals and farther from the average reference for depressed individuals. Furthermore, the speech data of healthy individuals, represented by circles, tends to be farther from the average reference for healthy individuals and farther from the average reference for depressed individuals. Furthermore, FIG. 12 also shows the ROC curve and AUC values. As shown in FIG. 12, the AUC value is 1.0 when depression is determined by combining the distance dist2 from the average reference of healthy subjects and the distance dist3 from the average reference of depressed patients.

[0091] FIG. 13 is a table of data showing the data shown in FIG. 12 and other experimental results. In FIG. 13, HAMD represents a depression assessment score. HAMD≧7 indicates that only those with a score of 7 or higher were evaluated. MDD represents depressed patients, PD represents patients with Parkinson's disease, AD represents patients with Alzheimer's disease, and HE represents healthy individuals. MDD=20 indicates that data from 20 depressed patients was used, and HE=14 indicates that data from 14 healthy individuals was used. Intra-person DTW refers to the case where DTW is performed by generating pairs between one section and another section in the speech data of a single subject, generating features, and then performing DTW. FIG. 13 shows the performance without the period adjustment preprocessing of this embodiment, and the performance value is AUC=0.7893, which is lower than the performance value when preprocessing is performed. As shown in the results in the top row of Figure 13, the AUC value when distinguishing between HE and MDD is 0.9643, and as shown in the results in the bottom row, the AUC value when distinguishing between HE and PD is 0.9173.

[0092] Figure 14 shows experimental results showing the effects of various preprocessing methods used in this embodiment. The results in the top row of Figure 14 are the baseline results. The results from the second row onwards show the effects of applying each preprocessing method to the data, and it can be seen that each preprocessing method contributes to the depression assessment performance. Note that in the data in the bottom row of the table in Figure 14, the performance evaluation (AUC) values ​​are reversed, which will be explained below.

[0093] Fig. 15 shows the difference in distance calculated by DTW when no amplitude adjustment is performed (denoted as "without amplitude normalization" in Fig. 15) and when amplitude adjustment is performed (denoted as "with amplitude normalization" in Fig. 15). The vertical axis of the graph shown in Fig. 15 is the distance calculated by DTW.

[0094] FIG. 15 shows the distance values ​​calculated by DTW for HE_HospitalA, which represents data obtained from multiple healthy individuals at Hospital A, and HE_HospitalB, which represents data obtained from multiple depressed patients MDD and multiple healthy individuals at Hospital B.

[0095] In the "without amplitude normalization" section of Figure 15, it can be seen that there is a large difference between HE_HospitalA, which represents data obtained from multiple healthy subjects at Hospital A, and HE_HospitalB, which represents data obtained from multiple healthy subjects at Hospital B. This is thought to be due to differences in the recording environment and recording settings of the audio data. When the sound pressure of the recorded audio data is low, the distance value calculated by DTW is small, while when the sound pressure of the audio data is high, the distance value calculated by DTW is large. As a result, as shown in the "without amplitude normalization" section of Figure 15, there is a difference between the data distributions of HE_HospitalA and HE_HospitalB, which were recorded in different environments (the difference in mean values ​​was significant by t-test: p<0.01).

[0096] In contrast, in the case of "with amplitude normalization" in Figure 15, the difference between the recording conditions at HE_HospitalA and HE_HospitalB is corrected, and no difference is observed between the data distributions of HE_HospitalA and HE_HospitalB, which were recorded in different environments (t-test showed no significant difference in the mean values: p>0.1). In this way, by incorporating amplitude normalization into the preprocessing, it is possible to correctly classify whether a user has a psychiatric disorder, a neurological disorder, or their symptoms, without being affected by differences in recording conditions.

[0097] Note that similar diseases can be predicted for Alzheimer's disease and Parkinson's disease. FIG. 16 shows various conditions and the results of discrimination between healthy individuals (HE) and patients (Sick) suffering from major depressive disorder (MDD), Alzheimer's disease (AD), and Parkinson's disease (PD) using the Intra-Person DTW of this embodiment. FIG. 17 shows an ROC curve corresponding to the performance evaluation AUC shown in FIG. 16 . FIG. 18 shows DTW values ​​calculated under the conditions shown in FIG. 16 . FIG. 19 shows actual symptoms (denoted "Actual" in FIG. 19 ) and prediction results (denoted "Prediction" in FIG. 19 ) using the method of this embodiment. FIG. 20 shows the results of a multiple comparison test.

[0098] As shown in Figure 16, the AUC value when distinguishing between healthy individuals (HE) and patients with some disease (Sick) is 0.8486. Furthermore, as shown in Figure 20, a multiple comparison test of the mean DTW values ​​showed that the distributions of healthy individuals (HE) and those with each disease (MDD, AD, PD) differed (significant difference in the mean values: p<0.01). Note that "E" in the table represents "×10," and the number next to it represents the exponent.

[0099] The results shown in Figures 16 to 20 also indicate that the method of this embodiment makes it possible to accurately estimate whether a user has a psychiatric disorder, a neurological disorder, or symptoms thereof.

[0100] The technology of the present disclosure is not limited to the above-described embodiment, and various modifications and applications are possible without departing from the spirit of the present invention.

[0101] For example, although the present specification has been described as an embodiment in which a program is pre-installed, the program may also be provided by being stored on a computer-readable recording medium.

[0102] In the above embodiment, the processing performed by the CPU after reading the software (program) may be performed by various processors other than the CPU. Examples of processors in this case include a programmable logic device (PLD) (such as a field-programmable gate array (FPGA) whose circuit configuration can be changed after manufacture), and a dedicated electrical circuit such as an application-specific integrated circuit (ASIC) that is a processor having a circuit configuration designed specifically to perform specific processing. Alternatively, a general-purpose graphics processing unit (GPGPU) may be used as the processor. Each process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, a plurality of FPGAs, or a combination of a CPU and an FPGA, etc.) The hardware structure of these various processors is, more specifically, an electric circuit that combines circuit elements such as semiconductor elements.

[0103] In addition, although the above embodiments have described the case where the program is pre-stored (installed) in storage, the present invention is not limited to this. The program may be provided in a form stored on a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The program may also be downloaded from an external device via a network.

[0104] Furthermore, each process of this embodiment may be implemented by a computer or server equipped with a general-purpose processor and a storage device, and each process may be executed by a program. This program is stored in a storage device, and may be recorded on a recording medium such as a magnetic disk, optical disk, or semiconductor memory, or may be provided via a network. Of course, any other components do not have to be implemented by a single computer or server, and may be distributed across multiple computers connected via a network.

[0105] Furthermore, in the above-described embodiments, examples of the predetermined disease or predetermined symptom have been described in which it is estimated whether or not a person has a psychiatric disease or a neurological disease, or a mental disorder symptom or a cognitive impairment symptom, but this is not limited thereto. The predetermined disease or predetermined symptom may be of any type. It is expected that various diseases or symptoms will be reflected in the voice data. For example, not only respiratory diseases and symptoms but also psychiatric diseases and the like will have an effect on the voice data. Therefore, in the above-described embodiments, examples of the predetermined disease or predetermined symptom have been described in which it is estimated whether or not a person has a psychiatric disease or a neurological disease, or a mental disorder symptom or a cognitive impairment symptom, but this is not limited thereto. Any disease or the like may be estimated as long as the effect of the disease or the like is reflected in the voice data.

[0106] In addition, in the above-described embodiments, the case where all of the preprocessing processes are performed when generating preprocessed audio data has been described as an example, but the present invention is not limited to this. Preprocessed audio data may be generated using at least one of the preprocessing processes described above.

[0107] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. an acquisition unit that acquires voice data that is time-series data of voice uttered by a user; a processing unit that generates preprocessed audio data representing data from the audio data acquired by the acquisition unit that is a first time after the start point of the audio data and a second time before the end point of the audio data; and a generating unit that generates processing result data by applying dynamic time warping to the pre-processed audio data generated by the processing unit; a calculation unit that calculates a score representing the degree to which the user has a predetermined disease or symptom based on the processing result data generated by the generation unit; an estimation unit that estimates whether the user has a predetermined disease or symptom based on the score calculated by the calculation unit; An information processing device comprising:

2. the processing unit generates, as the preprocessed audio data, data for a predetermined period of data that is a first time or later from the start point of the audio data and a second time or earlier than the end point of the audio data. The information processing device according to claim 1 .

3. the processing unit generates, as the preprocessed audio data, data obtained by performing a predetermined sampling process on data that is a first time after the start point of the audio data and that is a second time before the end point of the audio data.

3. The information processing device according to claim 1.

4. the processing unit generates the preprocessed audio data by performing a process of expanding or contracting, in a time axis direction, data that is a first time or later from the start point of the audio data and that is a second time or earlier than the end point of the audio data.

3. The information processing device according to claim 1.

5. the processing unit generates the preprocessed audio data by performing a process of expanding or contracting in an amplitude direction data that is a first time or later from the start point of the audio data and that is a second time or earlier than the end point of the audio data.

3. The information processing device according to claim 1.

6. the processing unit generates the preprocessed audio data by shifting data in a time axis direction, the data being a first time or later from the start point of the audio data and being a second time or earlier than the end point of the audio data.

3. The information processing device according to claim 1.

7. the generation unit applies the dynamic time warping method to the preprocessed voice data and the voice data of a reference user who is known to have the predetermined disease or symptom, thereby generating the processing result data representing a distance between the preprocessed voice data and the voice data of the reference user.

3. The information processing device according to claim 1.

8. the generation unit applies the dynamic time warping method to first audio data representing data in a first time interval within the preprocessed audio data and second audio data representing data in a second time interval within the preprocessed audio data, thereby generating the processing result data representing a distance between the first audio data and the second audio data.

3. The information processing device according to claim 1.

9. 3. An information processing system including a user terminal equipped with a microphone and the information processing device according to claim 1, the user terminal transmits the voice data acquired by the microphone to the information processing device; the acquisition unit of the information processing device acquires the voice data transmitted from the user terminal, a communication unit of the information processing device transmitting an estimation result estimated by the estimation unit to a user terminal; the user terminal receives the estimation result transmitted from the information processing device. Information processing system.

10. Acquires voice data, which is time-series data of the voice uttered by the user, generating preprocessed audio data representing data from the acquired audio data that is a first time after the start point of the audio data and a second time before the end point of the audio data; generating processing result data by applying dynamic time warping to the generated pre-processed audio data; Calculating a score representing the degree to which the user has a predetermined disease or symptom based on the generated processing result data; Inferring whether the user has a predetermined disease or symptom based on the calculated score. An information processing method that causes a computer to execute a process.

11. Acquires voice data, which is time-series data of the voice uttered by the user, generating preprocessed audio data representing data from the acquired audio data that is a first time after the start point of the audio data and a second time before the end point of the audio data; generating processing result data by applying dynamic time warping to the generated pre-processed audio data; Calculating a score representing the degree to which the user has a predetermined disease or symptom based on the generated processing result data; Inferring whether the user has a predetermined disease or symptom based on the calculated score. An information processing program that causes a computer to execute a process.