Speech processing device, speech processing method, program, and speech authentication system
The voice processing device uses machine learning to analyze speech features for assessing driver health, overcoming the cost and reluctance issues of biometric sensor-based systems by providing efficient and immediate health evaluations.
Patent Information
- Application Number
- JP2022539897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-07-30
AI Technical Summary
Existing technologies for assessing driver health in commercial vehicles require biometric sensors and cameras, leading to high costs and reluctance from companies to adopt such systems.
A voice processing device that extracts features from speech data using machine learning to determine a person's state without the need for a user interview or biometric sensors, utilizing a classifier trained on normal speech data to calculate an index value representing similarity with registered data.
Enables easy determination of a person's condition without interviews or biometric sensors, providing immediate assessment results based on speech analysis.
Smart Images

Figure 0007718417000001 
Figure 0007718417000002 
Figure 0007718417000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice processing device, a voice processing method, a recording medium, and a voice authentication system, and more particularly to a voice processing device, a voice processing method, a recording medium, and a voice authentication system that verify speakers based on voice data. [Background technology]
[0002] Taxi and bus companies hold "roll calls" in which all drivers participate. Operations managers check the health of drivers by conducting brief interviews with them. However, health checks conducted through interviews can lead to drivers consciously or unconsciously lying, or overconfident or misperceived beliefs about their own health. Therefore, related technologies have been developed to reliably check the health of drivers. For example, Patent Document 1 describes a technology that uses biosensors and cameras installed in commercial vehicles carrying drivers to detect electrocardiograms, electromyograms, eye movements, electroencephalograms, respiration, blood pressure, sweating, and other factors to comprehensively assess the driver's physical and mental health. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2020 / 003392 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-201014 [Patent Document 3] Japanese Patent Application Laid-Open No. 2015-069255 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the related technology described in Patent Document 1 requires the installation of a biometric sensor and a camera in each commercial vehicle owned by a company, which can lead to high costs and lead to some companies being reluctant to adopt such technology.
[0005] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to provide a technology that can easily determine the condition of a person being evaluated without the need for a user to meet with the person being evaluated or for a biometric sensor. [Means for solving the problem]
[0006] A voice processing device according to one aspect of the present invention includes a feature extraction means for extracting features of input data from input data based on the speech of the person to be judged, using a classifier that has undergone machine learning using speech data based on the speech of the person to be judged when in a normal state as training data; an index value calculation means for calculating an index value representing the degree of similarity between the features of the input data and the features of speech data based on the speech of the person to be judged when in a normal state; and a state judgment means for judging whether the person to be judged is in a normal state or an abnormal state based on the index value.
[0007] A voice processing method according to one aspect of the present invention includes: extracting features of input data from input data based on the speech of the person to be determined when the person is in a normal state, using a classifier that has undergone machine learning using speech data based on the speech of the person to be determined when the person is in a normal state as training data; calculating an index value that represents the degree of similarity between the features of the input data and the features of speech data based on the speech of the person to be determined when the person is in a normal state; and determining whether the person to be determined is in a normal state or an abnormal state based on the index value.
[0008] A recording medium according to one aspect of the present invention stores a program for causing a computer to execute the following steps: extracting features of input data from input data based on the speech of the person being judged using a classifier that has undergone machine learning using voice data based on the speech of the person being judged when in a normal state as training data; calculating an index value representing the degree of similarity between the features of the input data and the features of voice data based on the speech of the person being judged when in a normal state; and determining whether the person being judged is in a normal state or an abnormal state based on the index value.
[0009] A voice authentication system according to one aspect of the present invention includes the voice processing device according to the above-described aspect, and a learning device that trains the classifier using voice data based on the speech of the person to be judged when in a normal state as the training data. [Effects of the Invention]
[0010] According to one aspect of the present invention, the state of a person to be assessed can be easily determined without the need for a user to interview the person to be assessed or for a biometric sensor. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram for explaining the configuration and operation of a voice processing device according to a first embodiment. [Figure 2] FIG. 10 is a block diagram showing the configuration of a voice processing device according to a second embodiment. [Figure 3] 10 is a flowchart showing the operation of the voice processing device according to the second embodiment. [Figure 4] FIG. 10 is a block diagram showing the configuration of a voice processing device according to a third embodiment. [Figure 5] 10 is a flowchart showing the operation of the voice processing device according to the third embodiment. [Figure 6] FIG. 10 is a diagram illustrating a hardware configuration of a voice processing device according to a second or third embodiment. [Figure 7]FIG. 10 is a block diagram showing the configuration of a voice authentication system including a voice processing device according to a second or third embodiment and a learning device. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, several embodiments will be described in detail with reference to the drawings.
[0013] [Embodiment 1] (Configuration and Operation of the Audio Processing Device X00 According to the First Embodiment) FIG. 1 is a diagram illustrating an outline of the configuration and operation of a voice processing device X00 according to a first embodiment. As shown in FIG. 1, the voice processing device X00 receives a voice signal (input data in FIG. 1) input by a person to be determined using an input device such as a microphone. An example of a person to be determined is a person whose state is to be determined by the voice processing device X00. The configuration and operation of the voice processing device X00 described in the first embodiment can also be realized in a voice processing device 100 according to a second embodiment and a voice processing device 200 according to a third embodiment, which will be described later.
[0014] For example, the voice processing device X00 supports a crew member (e.g., a driver) in performing his / her duties normally at a company that provides bus operation services. In this case, the subject of the judgment is the bus crew member. Specifically, the voice processing device X00 judges the condition of the crew member using the method described below, and determines whether or not the crew member is allowed to drive based on the judgment result.
[0015] The voice processing device X00 communicates with a microphone installed in a specific location (for example, a bus office) via a wireless network, and receives, as input data, a voice signal input to the microphone when the person to be determined speaks into the microphone. Alternatively, the voice processing device X00 may receive, as input data, a voice signal input to a microphone worn by the person to be determined at any timing. For example, the voice processing device X00 receives, as input data, a voice signal input to a microphone worn by the person to be determined immediately before the driver who is the person to be determined leaves the bus depot.
[0016] Furthermore, the voice processing device X00 may receive a voice signal (registered data in FIG. 1) that has been registered in advance in a database (DB). The registered data is a voice signal input by the person to be determined when it has been confirmed that the person to be determined is in a normal state through a medical examination or an analysis of biometric data. The registered data is stored in the DB in association with identification information of the person to be determined and identification information of the microphone used by the person to be determined.
[0017] The voice processing device X00 determines whether the person to be determined is in a normal state or an abnormal state based on input data based on the speech of the person and registered data.
[0018] In a more detailed example, the voice processing device X00 compares input data based on the speech of the person to be determined with registered data, and determines the state of the person to be determined based on an index value representing the degree of similarity between the input data and registered data. The state of the person to be determined here refers to an evaluation of the person's physical and mental state.
[0019] In one example, the state of the person to be assessed represents the physical condition or emotions of the person to be assessed. In this case, the abnormal state of the person to be assessed means that the person to be assessed is in poor physical condition due to fever or lack of sleep, has an illness such as a cold, or has a psychological problem (such as anxiety). On the other hand, the normal state of the person to be assessed means that the person to be assessed does not have any of the problems exemplified above. More specifically, the normal state of the person to be assessed means that the person to be assessed does not have any physical or mental problems that may hinder the person to be assessed from performing their work or associated tasks.
[0020] In the following description, it is assumed that the person to be judged is the person whose identification information is registered along with the registration data, and that this person has been confirmed by the operations manager through visual inspection or other means, such as facial recognition, iris recognition, fingerprint recognition, or other biometric authentication.
[0021] [Embodiment 2] A second embodiment will be described with reference to FIGS.
[0022] (Speech processing device 100) The configuration of a voice processing device 100 according to the second embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the configuration of the voice processing device 100.
[0023] As shown in FIG. 2, the voice processing device 100 includes a feature extraction unit 110, an index value calculation unit 120, and a state determination unit .
[0024] The feature extraction unit 110 extracts features of input data from input data based on the speech of the person to be determined, using a classifier (FIG. 1 or FIG. 7) that has undergone machine learning using speech data based on the speech of the person to be determined when in a normal state as training data. The feature extraction unit 110 is an example of a feature extraction means. The training data is speech data based on the speech of the person to be determined when in a normal state.
[0025] In one example, the feature extraction unit 110 receives input data ( FIG. 1 ) input using an input device such as a microphone. The feature extraction unit 110 also receives enrollment data ( FIG. 1 ) from a database (not shown). The feature extraction unit 110 inputs the input data to a trained classifier (hereinafter simply referred to as a classifier) and extracts features of the input data from the classifier. The feature extraction unit 110 also inputs enrollment data to the classifier and extracts features of the enrollment data from the feature extraction unit 110.
[0026] The feature extraction unit 110 may use any machine learning technique to extract the features of the input data and the registered data. An example of machine learning here is deep learning, and an example of a classifier is a deep neural network (DNN). In this case, the feature extraction unit 110 inputs the input data to the DNN and extracts the features of the input data from the intermediate layer of the DNN. In one example, the features extracted from the input data may be mel-frequency cepstrum coefficients (MFCC) or linear predictive coding (LPC) coefficients, or may be a power spectrum or a spectral envelope. Alternatively, the features of the input data may be a feature vector of any dimension (hereinafter referred to as an acoustic vector) composed of feature quantities obtained by frequency analysis of audio data.
[0027] The feature extraction unit 110 outputs the data on the features of the registered data and the data on the features of the input data to the index value calculation unit 120.
[0028] The index value calculation unit 120 calculates an index value that represents the degree of similarity between the features of the input data and the features of the voice data based on the utterance of the subject of judgment when the subject was in a normal state. The index value calculation unit 120 is an example of an index value calculation means. The voice data based on the utterance of the subject of judgment when the subject was in a normal state corresponds to the above-mentioned registered data.
[0029] In one example, the index value calculation unit 120 receives feature data of the input data from the feature extraction unit 110. The index value calculation unit 120 also receives feature data of the enrollment data from the feature extraction unit 110. The index value calculation unit 120 identifies the phonemes included in the input data and the phonemes included in the enrollment data. The index value calculation unit 120 associates the phonemes included in the input data with the same phonemes included in the enrollment data.
[0030] Next, in one example, the index value calculation unit 120 calculates a score representing the degree of similarity between the feature of a phoneme included in the input data and the feature of the same phoneme included in the enrollment data, and calculates the sum of the scores calculated for all phonemes as the index value. The feature of a phoneme included in the input data and the feature of a phoneme included in the enrollment data may be feature vectors of the same dimension. Furthermore, the score representing the degree of similarity may be the inverse of the distance between the feature vector of a phoneme included in the input data and the feature vector of the same phoneme included in the enrollment data, or may be "(upper limit of the distance) - distance." In the following description, "score" refers to the sum of the above-mentioned scores. Furthermore, "features of the input data" and "features of the enrollment data" refer to "features of a phoneme included in the input data" and "features of the same phoneme included in the enrollment data," respectively.
[0031] The index value calculation unit 120 outputs data of the calculated index value (for example, a score) to the state determination unit 130.
[0032] The state determination unit 130 determines whether the person being determined is in a normal state or an abnormal state based on the index value. The state determination unit 130 is an example of a state determination means. In one example, the state determination unit 130 receives index value data representing the degree of similarity between the features of the input data and the features of the registered data from the index value calculation unit 120.
[0033] Next, in one example, the state determination unit 130 compares the index value with a predetermined threshold. If the index value is greater than the threshold, the state determination unit 130 determines that the person being determined is in a normal state. On the other hand, if the index value is equal to or less than the threshold, the state determination unit 130 determines that the person being determined is in an abnormal state. The state determination unit 130 outputs the result of the determination.
[0034] In addition, the state determination unit 130 may limit the authority of the person to be determined to operate the object. For example, the object is a commercial vehicle that the person to be determined is attempting to operate. In this case, the state determination unit 130 may control a computer of the commercial vehicle so that the engine of the commercial vehicle cannot be started.
[0035] (Operation of the voice processing device 100) An example of the operation of the voice processing device 100 according to the second embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of processing executed by each unit (Fig. 2) of the voice processing device 100 in this example.
[0036] 3, the feature extraction unit 110 extracts features of input data from the input data (FIG. 1) (S101). The feature extraction unit 110 also extracts features of registered data from the registered data (FIG. 1). Then, the feature extraction unit 110 outputs data on the features of the input data and the registered data to the index value calculation unit 120.
[0037] The index value calculation unit 120 receives data on the features of the input data and data on the features of the registered data from the feature extraction unit 110. The index value calculation unit 120 calculates an index value that represents the degree of similarity between the features of the input data and the features of the registered data (S102). In one example, the index value calculation unit 120 calculates a score that represents the distance between a feature vector that represents the features of the input data and a feature vector that represents the features of the registered data as the index value. The index value calculation unit 120 outputs data on the calculated index value (score) to the state determination unit 130.
[0038] The state determination unit 130 receives score data representing the degree of similarity between the features of the input data and the features of the registered data from the index value calculation unit 120. The state determination unit 130 compares the score with a predetermined threshold value (S103).
[0039] If the score is greater than the threshold value (Yes in S103), the state determination unit 130 determines that the person being determined is in a normal state (S104A).
[0040] On the other hand, if the score is equal to or less than the threshold value (No in S103), the state determination unit 130 determines that the person being determined is in an abnormal state (S104B). Thereafter, the state determination unit 130 may output the determination result (step S104A or S104B).
[0041] This completes the operation of the voice processing device 100 according to the second embodiment.
[0042] (Effects of this embodiment) According to the configuration of this embodiment, the feature extraction unit 110 extracts features of input data from input data based on the speech of the subject of assessment using a classifier that has undergone machine learning using speech data based on the speech of the subject of assessment when in a normal state as training data. The index value calculation unit 120 calculates an index value representing the degree of similarity between the features of the input data and the features of speech data based on the speech of the subject of assessment when in a normal state. The state determination unit 130 determines whether the subject of assessment is in a normal state or an abnormal state based on the index value. The voice processing device 100 can use the classifier to obtain an index value indicating the likelihood that the person is in a normal state. The determination result based on this index value indicates how similar the speech of the subject of assessment is to the speech of the person in a normal state. Therefore, the voice processing device 100 can easily determine the state of the subject of assessment (whether normal or abnormal) without the user needing to meet with the subject of assessment or using a biometric sensor. Furthermore, when the determination result by the voice processing device 200 is output, the user can immediately check the state of the subject of assessment.
[0043] [Embodiment 3] A third embodiment will be described with reference to FIGS.
[0044] (Speech processing device 200) The outline of the operation of the audio processing device 200 according to the present embodiment 3 is the same as the operation of the audio processing device 100 described in the above-described embodiment 2. Basically, the audio processing device 200 performs the same operation as the audio processing device X00 described in the above-described embodiment 1 with reference to Fig. 1, but as will be described below, the audio processing device 200 also performs operations that are partially different from those of the audio processing device X00.
[0045] FIG. 4 is a block diagram showing the configuration of a voice processing device 200 according to the third embodiment. As shown in FIG. 4, the voice processing device 200 includes a feature extraction unit 110, an index value calculation unit 120, and a state determination unit 130. The voice processing device 200 also includes a presentation unit 240. That is, the configuration of the voice processing device 200 according to the third embodiment differs from that of the voice processing device 100 according to the second embodiment in that it includes the presentation unit 240. In the third embodiment, the processes performed by the components with the same reference numerals as those in the second embodiment are also the same. Therefore, in the third embodiment, only the processes performed by the presentation unit 240 will be described.
[0046] The presentation unit 240 presents information indicating whether the person to be determined is in a normal state or an abnormal state, based on the result of the determination by the state determination unit 130 of the voice processing device 200. The presentation unit 240 is an example of a presentation means.
[0047] In one example, the presenting unit 240 obtains data on the determination result indicating whether the person being determined is in a normal state or an abnormal state from the state determining unit 130. The presenting unit 240 may present different information depending on the data on the determination result.
[0048] For example, when the state determination unit 130 determines that the person to be determined is in a normal state, the presentation unit 240 acquires data on the index value (score) calculated by the index value calculation unit 120 and presents information indicating the likelihood of the determination result based on the index value (score). Specifically, the presentation unit 240 displays on the screen that the person to be determined is in a normal state by using text, symbols, or a light on the screen. On the other hand, when the state determination unit 130 determines that the person to be determined is in an abnormal state, the presentation unit 240 issues an alarm. In addition, the presentation unit 240 may acquire data on the index value (score) calculated by the index value calculation unit 120 and output the acquired data on the index value (score) to a display device (not shown), thereby displaying the index value (score) on the screen of the display device.
[0049] (Operation of the voice processing device 200) The operation of the voice processing device 200 according to the third embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the processes executed by each unit (Fig. 4) of the voice processing device 200.
[0050] As shown in Fig. 5, the presentation unit 240 outputs data of a message encouraging the subject of judgment to make a long utterance to a display device (not shown), thereby displaying the message on the screen of the display device (S201). The meaning of a long utterance (or the definition of the length of an utterance) may be determined appropriately by the user of the speech processing device 200. In one example, a long utterance is an utterance including N words or more (N is a number set by the user). The reason for requesting the subject of judgment to make a long utterance is to accurately calculate an index value representing the degree of similarity between the features of the input data and the features of the registered data.
[0051] The feature extraction unit 110 receives an audio signal (input data in FIG. 1) obtained by collecting the speech of the person to be determined from an input device such as a microphone (S202). The feature extraction unit 110 also receives an audio signal (registered data in FIG. 1) recorded when the person to be determined is in a normal state from the DB.
[0052] The feature extraction unit 110 extracts features of the input data from the input data (S203). Also, the feature extraction unit 110 extracts features of the registered data from the registered data.
[0053] Then, the index value calculation unit 120 calculates an index value (score) that represents the degree of similarity between the features of the input data and the features of the registered data (S204).
[0054] The state determination unit 130 compares the index value with a predetermined threshold value (S205). If the score is greater than the threshold value (Yes in S205), the state determination unit 130 determines that the person being determined is in a normal state (S206A). The state determination unit 130 outputs the determination result to the presentation unit 240. In this case, the presentation unit 240 displays information indicating that the person being determined is in a normal state on a display device (not shown) (S207A).
[0055] On the other hand, if the score is equal to or less than the threshold value (No in S205), the state determination unit 130 determines that the person being determined is in an abnormal state (S206B). The state determination unit 130 outputs the determination result to the presentation unit 240. In this case, the presentation unit 240 issues an alarm (S207B).
[0056] Additionally, in step S207B, the presenting unit 240 may display information indicating that the subject of the assessment is in an abnormal state on a display device (not shown). In one example, the presenting unit 240 acquires data of the index value (score) calculated in step S204 from the index value calculation unit 120, and displays the acquired score itself or information based on the score (in one example, a suggestion for re-examination) on the display device.
[0057] This completes the operation of the voice processing device 200 according to the third embodiment.
[0058] (Effects of this embodiment) According to the configuration of this embodiment, the feature extraction unit 110 extracts features of input data from input data based on the speech of the subject of assessment using a classifier that has undergone machine learning using speech data based on the speech of the subject of assessment when in a normal state as training data. The index value calculation unit 120 calculates an index value representing the degree of similarity between the features of the input data and the features of speech data based on the speech of the subject of assessment when in a normal state. The state determination unit 130 determines whether the subject of assessment is in a normal state or an abnormal state based on the index value. As a result, the voice processing device 200 can use the classifier to obtain an index value indicating the likelihood that the subject of assessment is in a normal state. The determination result based on this index value indicates how similar the speech of the subject of assessment is to the speech of that person when in a normal state. Therefore, the voice processing device 200 can easily determine the state of the subject of assessment (whether normal or abnormal) without the need for a user's interview with the subject of assessment or biometric data. Furthermore, when the determination result by the voice processing device 200 is output, the user can immediately check the state of the subject of assessment.
[0059] Furthermore, according to the configuration of this embodiment, the presentation unit 240 presents information indicating whether the person being assessed is in a normal state or an abnormal state based on the result of the assessment. Therefore, the user who sees the presented information can easily understand the state of the person being assessed. Then, the user can appropriately take measures (for example, re-interviewing with the crew or restricting work) according to the understood state of the person being assessed.
[0060] [Hardware configuration] Each of the components of the audio processing devices 100 and 200 described in the second and third embodiments is represented by a functional block. Some or all of these components are realized by an information processing device 900 as shown in Fig. 6. Fig. 6 is a block diagram showing an example of the hardware configuration of the information processing device 900.
[0061] As shown in FIG. 6, the information processing device 900 includes, for example, the following configuration.
[0062] ·CPU(Central Processing Unit)901 ROM (Read Only Memory) 902 ·RAM(Random Access Memory)903 Program 904 loaded into RAM 903 A storage device 905 for storing a program 904 A drive device 907 for reading and writing data from and to the recording medium 906 A communication interface 908 for connecting to a communication network 909 Input / output interface 910 for inputting and outputting data Bus 911 connecting each component Each of the components of the audio processing devices 100 and 200 described in the second and third embodiments is realized by the CPU 901 reading and executing a program 904 that realizes the functions of the components. The program 904 that realizes the functions of the components is stored in the storage device 905 or the ROM 902 in advance, for example, and is loaded into the RAM 903 and executed by the CPU 901 as needed. The program 904 may be supplied to the CPU 901 via the communication network 909, or may be stored in the recording medium 906 in advance, and the drive device 907 may read out the program and supply it to the CPU 901.
[0063] According to the above configuration, the audio processing devices 100 and 200 described in the second and third embodiments are realized as hardware, and therefore the same effects as those described in the second and third embodiments can be achieved.
[0064] [Common to Embodiments 2 and 3] An example of the configuration of a voice authentication system to which the voice processing device according to the second or third embodiment described above is commonly applied will be described.
[0065] (Voice Authentication System 1) An example of the configuration of the voice authentication system 1 will be described with reference to Fig. 7. Fig. 7 is a block diagram showing an example of the configuration of the voice authentication system 1.
[0066] 7, the voice authentication system 1 includes a voice processing device 100 (200) and a learning device 10. The voice authentication system 1 may also include one or more input devices. The voice processing device 100 (200) is the voice processing device 100 according to the second embodiment or the voice processing device 200 according to the third embodiment.
[0067] As shown in FIG. 7, the learning device 10 acquires training data from a database (DB) on a network or from a DB connected to the learning device 10. The learning device 10 uses the acquired training data to train a classifier. More specifically, the learning device 10 inputs speech data included in the training data to the classifier, provides correct answer information included in the training data to the output of the classifier, and calculates the value of a well-known loss function. The learning device 10 then updates the parameters of the classifier a predetermined number of times so as to reduce the calculated value of the loss function. Alternatively, the learning device 10 repeatedly updates the parameters of the classifier until the value of the loss function becomes equal to or less than a predetermined value.
[0068] As described in the second embodiment, the speech processing device 100 determines the state of the person to be determined by using a trained classifier. Similarly, the speech processing device 200 according to the third embodiment also determines the state of the person to be determined by using a trained classifier. [Industrial Applicability]
[0069] In one example, the present invention can be used in a voice authentication system that performs identity verification by analyzing voice data input using an input device. [Explanation of symbols]
[0070] 1. Voice authentication system 10 Learning Device 100 Audio processing device 110 Feature Extraction Unit 120 Index value calculation unit 130 Status determination unit 200 Audio processing device 240 Presentation section
Claims
1. a feature extraction means for extracting features for identifying the subject of judgment from input data based on the speech of the subject of judgment, using a classifier that has undergone machine learning using voice data based on the speech of the subject of judgment in a normal state as training data; and an index value calculation means for calculating an index value representing a degree of similarity between the features for identifying the subject of determination and the features of the voice data of the subject of determination extracted from the classifier to which voice data based on the speech of the subject of determination in a normal state is input; a state determination means for determining whether the person to be determined is in a normal state or an abnormal state based on the index value; Equipped with a learning device included in the voice authentication system uses voice data based on the speech of the subject of the determination when the subject is in a normal state as the training data to train the classifier; the classifier is a deep neural network (DNN), The feature extraction means extracts features of the input data from the intermediate layer of the DNN. Audio processing device.
2. The device further includes a display unit that displays information indicating whether the person being determined is in a normal state or an abnormal state based on the result of the determination.
2. The audio processing device according to claim 1, wherein:
3. If the subject of the determination is determined to be in an abnormal state, The presenting means presents information indicating the likelihood of the result of the determination based on the index value.
3. The audio processing device according to claim 2.
4. If the subject of the determination is determined to be in an abnormal state, The state determination means limits the authority of the person to be determined to operate the object.
2. The audio processing device according to claim 1, wherein:
5. extracting features for identifying the subject of judgment from input data based on the speech of the subject of judgment using a classifier that has undergone machine learning using speech data based on the speech of the subject of judgment when the subject of judgment was in a normal state as training data; calculating an index value representing a degree of similarity between the features for identifying the subject of determination and the features of the voice data of the subject of determination extracted from the classifier to which voice data based on the speech of the subject of determination in a normal state has been input; A voice processing method for determining whether a person to be determined is in a normal state or an abnormal state based on the index value, a learning device included in the voice authentication system uses voice data based on the speech of the subject of the determination when the subject is in a normal state as the training data to train the classifier; the classifier is a deep neural network (DNN), Extracting features of the input data from the intermediate layer of the DNN Audio processing methods.
6. extracting features for identifying the subject of judgment from input data based on the speech of the subject of judgment using a classifier that has undergone machine learning using speech data based on the speech of the subject of judgment when the subject of judgment was in a normal state as training data; and calculating an index value representing a degree of similarity between the features for identifying the subject of determination and the features of the voice data of the subject of determination extracted from the classifier to which voice data based on the speech of the subject of determination in a normal state is input; and determining whether the subject of the determination is in a normal state or an abnormal state based on the index value, the classifier is trained by a learning device included in the voice authentication system using, as the training data, voice data based on the speech of the person to be determined when the person is in a normal state; the classifier is a deep neural network (DNN), Extracting features of the input data from the intermediate layer of the DNN A program for causing the computer to execute the above.
7. A voice processing device according to any one of claims 1 to 4; a learning device that uses, as the training data, speech data based on the speech of the subject of the judgment when the subject is in a normal state, to train the classifier; A voice authentication system equipped with
Citation Information
Patent Citations
Drinking detection device for vehicles, and drinking detecting method for vehicles
JP2010015027A
Driver roll call system
JP2015069255A
Operation management system and taxi meter
JP2016201014A
Emotion recognition device and emotion recognition program
JP2020099367A
Method for judgment of drinking using differential frequency energy, recording medium and device for performing the method
US9907509B2