Impersonation detection device, impersonation detection system, impersonation detection method, and program

The impersonation detection system uses air and bone conduction sound data comparison to accurately identify synthesized speech, enhancing security in voice communications by preventing impersonation.

JP2026066750APending Publication Date: 2026-04-17NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing voice impersonation detection methods using synthesized speech lack accuracy in determining the presence or absence of impersonation.

Method used

An impersonation detection system that utilizes both air conduction sound data and bone conduction sound data to determine whether speech is human or synthesized, employing a comparison method to identify correspondence between these sound types and stored data.

Benefits of technology

Enables high-accuracy determination of synthesized speech impersonation by analyzing air and bone conduction sound data, preventing fraud and improving security in voice-based communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066750000001_ABST
    Figure 2026066750000001_ABST
Patent Text Reader

Abstract

The goal is to enable relatively high-accuracy detection of impersonation using synthesized speech. [Solution] The impersonation detection device includes a voice data analysis means that determines whether or not impersonation using synthesized speech is occurring based on a comparison between the air conduction sound data and bone conduction sound data to be detected and the stored air conduction sound data and bone conduction sound data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an impersonation determination device, an impersonation determination system, an impersonation determination method, and a program.

Background Art

[0002] Regarding person authentication using voice, Patent Document 1 describes performing authentication using the difference in the frequency spectra between air-conducted sound and bone-conducted sound (bone conduction sound). In the method described in Patent Document 1, the user is made to input a specified voice, and the phase difference at which the integrated amplitude of the difference waveform between the air-conducted sound and the bone-conducted sound simultaneously sampled from the input voice is minimized is calculated. Then, in the method described in Patent Document 1, the calculated phase difference is compared with the standard phase difference stored as master data, and if the deviation between the calculated phase difference and the standard phase difference is within the allowable range, the authentication flag is set to permitted, and if it is outside the range, it is set to non-permitted.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When determining the presence or absence of impersonation using synthesized voice, it is preferable to perform the determination with as high accuracy as possible.

[0005] An example of the object of the present disclosure is to provide an impersonation determination device, an impersonation determination system, an impersonation determination method, and a program capable of solving the above-described problems.

Means for Solving the Problems

[0006] According to a first aspect of this disclosure, the impersonation detection device includes voice data analysis means for determining whether or not impersonation using synthesized speech is occurring, based on a comparison between the air conduction sound data and bone conduction sound data to be determined and the stored air conduction sound data and bone conduction sound data.

[0007] According to a second aspect of the present disclosure, the impersonation detection system comprises a first terminal device, an impersonation detection device, and a second terminal device, wherein the first terminal device comprises air conduction sound data acquisition means for acquiring air conduction sound data to be determined, and bone conduction sound data acquisition means for acquiring bone conduction sound data to be determined, the impersonation detection device comprises voice data analysis means for determining whether or not impersonation using synthesized voice is occurring based on a comparison of the air conduction sound data and bone conduction sound data to be determined with stored air conduction sound data and bone conduction sound data, and the second terminal device comprises display means for displaying the determination result of whether or not impersonation using synthesized voice is occurring.

[0008] According to a third aspect of this disclosure, the impersonation detection method includes a computer determining whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data to be determined and the stored air conduction sound data and bone conduction sound data.

[0009] According to a fourth aspect of this disclosure, the program is a program that causes a computer to determine whether or not there is impersonation using synthesized speech based on a comparison between the air conduction sound data and bone conduction sound data to be determined and the stored air conduction sound data and bone conduction sound data. [Effects of the Invention]

[0010] According to one aspect of this disclosure, it is possible to determine with relatively high accuracy whether or not impersonation using synthesized speech is occurring. [Brief explanation of the drawing]

[0011] [Figure 1] This figure shows an example of the configuration of a spoofing detection system according to at least one embodiment. [Figure 2] This figure shows an example of the configuration of a first terminal device according to at least one embodiment. [Figure 3] This figure shows an example of the configuration of a second terminal device according to at least one embodiment. [Figure 4] This figure shows an example of the configuration of an analysis server device according to at least one embodiment. [Figure 5] This figure shows an example of the procedure for processing a call in a spoofing detection system according to at least one embodiment. [Figure 6] This figure shows an example of the procedure for determining whether or not an impersonation attempt is being made using synthesized speech, using an analysis server device according to at least one embodiment. [Figure 7] This figure shows an example of the configuration of a spoofing detection system according to at least one embodiment. [Figure 8] This figure shows an example of the configuration of a first terminal device according to at least one embodiment. [Figure 9] This figure shows an example of the configuration of an analysis server device according to at least one embodiment. [Figure 10] This figure shows an example of the procedure for processing a call in a spoofing detection system according to at least one embodiment. [Figure 11] This figure shows an example of the procedure for determining whether or not an impersonation attempt is being made using synthesized speech, using an analysis server device according to at least one embodiment. [Figure 12] This figure shows an example of the procedure for determining whether or not an impersonation attempt is being made using synthesized speech, using an analysis server device according to at least one embodiment. [Figure 13] This figure shows an example of the configuration of a spoofing detection device according to at least one embodiment. [Figure 14] This figure shows an example of the configuration of a spoofing detection system according to at least one embodiment. [Figure 15] This figure shows an example of the processing steps in a spoofing detection method according to at least one embodiment. [Figure 16]It is a diagram showing an example of the configuration of a computer according to at least one embodiment.

Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described. However, the following embodiments do not limit the invention according to the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention.

[0013] <First Embodiment> FIG. 1 is a diagram showing an example of the configuration of an impersonation determination system according to at least one embodiment. In the configuration shown in FIG. 1, the impersonation determination system 1 includes a first terminal device 100, a second terminal device 200, and an analysis server device 300.

[0014] The impersonation determination system 1 determines the presence or absence of impersonation using synthesized voice for a call using the first terminal device 100. Specifically, the impersonation determination system 1 determines whether the voice transmitted by the first terminal device 100 is due to human speech or synthesized voice. The first terminal device 100 acquires voice data of the caller and transmits it to the analysis server device 300. In particular, the first terminal device 100 acquires air-conducted sound data and bone-conducted sound data of the caller and transmits these data to the analysis server device. The first terminal device 100 may be configured using a computer such as a smartphone or a personal computer (PC).

[0015] The analysis server device 300 determines the presence or absence of impersonation using synthesized voice based on the air-conducted sound data and bone-conducted sound data from the first terminal device 100. The analysis server device 300 corresponds to an example of an impersonation determination device. Specifically, the analysis server device 300 determines whether the air conduction sound data and the bone conduction sound data correspond to each other. Correspondence between air conduction sound data and bone conduction sound data may mean that the air conduction sound data and bone conduction sound data were obtained from the same utterance. Alternatively, correspondence between air conduction sound data and bone conduction sound data may mean that the air conduction sound data and bone conduction sound data were obtained from the utterance of the same person (speaker). The air conduction sound data from the first terminal device 100 corresponds to the air conduction sound data subject to judgment (air conduction sound data subject to judgment on whether or not there is impersonation by synthesized voice). The bone conduction sound data from the first terminal device 100 corresponds to the bone conduction sound data subject to judgment (bone conduction sound data subject to judgment on whether or not there is impersonation by synthesized voice).

[0016] Furthermore, the analysis server device 300 compares the air conduction sound data and bone conduction sound data from the first terminal device 100 with the air conduction sound data and bone conduction sound data stored in the speech data storage unit 330. This allows the analysis server device 300 to determine whether the air conduction sound data and bone conduction sound data from the first terminal device 100 were obtained from the speech of the same person.

[0017] Specifically, the analysis server device 300 searches the air conduction sound data stored in the speech data storage unit 330 for air conduction sound data obtained from the speech of the same person as the air conduction sound data to be judged. If air conduction sound data is detected, the analysis server device 300 determines whether the bone conduction sound data paired with the detected air conduction sound data and the bone conduction sound data to be judged are data obtained from the speech of the same person.

[0018] In this case, the determination can also be understood as a determination of whether or not the air conduction sound data and bone conduction sound data from the first terminal device 100 correspond to each other. Similarly, with respect to air conduction sound data, if the air conduction sound data obtained from the speech of the same person corresponds to each other, it can also be said that the bone conduction sound data obtained from the speech of the same person corresponds to each other.

[0019] If the analysis server device 300 determines that the air conduction sound data and bone conduction sound data from the first terminal device 100 correspond, it determines that no spoofing using synthesized speech has occurred (no spoofing). On the other hand, if the analysis server device 300 determines that the air conduction sound data and bone conduction sound data from the first terminal device 100 do not correspond, it determines that spoofing using synthesized speech has occurred (spoofing present).

[0020] The analysis server device 300 then performs processing according to the determination result. For example, the analysis server device 300 may send the determination result to the second terminal device 200. Also, when a user of the first terminal device 100 and a user of the second terminal device 200 are making a call, the analysis server device 300 may terminate the call (disconnect the call) if it determines that there is impersonation using synthesized speech.

[0021] The analysis server device 300 may be configured using a computer such as a workstation or personal computer. The analysis server device 300 may be configured integrally with the first terminal device 100 or the second terminal device 200.

[0022] The following explanation will use the case where the impersonation detection system 1 is used for a call from a user of the first terminal device 100 to a user of the second terminal device 200 as an example. The user of the first terminal device 100 will also be referred to as the first user. The user of the second terminal device 200 will also be referred to as the second user. Furthermore, calls between users of terminal devices are also referred to as calls between terminal devices. For example, a call between user 1 and user 2 is also referred to as a call between terminal device 1 and terminal device 2.

[0023] Furthermore, the following explanation will use the case where the analysis server device 300 sends the result of the determination of whether or not impersonation using synthesized speech is occurring to the second terminal device 200 as an example. The second user may also terminate the call if they determine that it is an unauthorized call by referring to the determination result.

[0024] However, the use of the impersonation detection system 1 is not limited to a specific use. For example, the impersonation detection system 1 may be used for voiceprint authentication. In this case, the analysis server device 300 may acquire air conduction sound data and bone conduction sound data from the first terminal device 100 and, in addition to voiceprint authentication, perform a determination of whether or not impersonation is occurring using synthesized speech. The analysis server device 300 may then transmit the results of the voiceprint authentication and the determination of whether or not impersonation is occurring using synthesized speech to the second terminal device 200.

[0025] Furthermore, when the impersonation detection system 1 is used for calls between users, it may be used for one-way calls from one specific user to another, or for two-way calls between users. For example, when the impersonation detection system 1 is used for two-way calls between a first user and a second user, the second terminal device 200 may have a function to acquire the air conduction sound data and bone conduction sound data of the second user, and the analysis server device 300 may also determine whether or not the second user's speech is being impersonated using synthesized speech.

[0026] Furthermore, the impersonation detection system 1 may be used for two-way calls such as telephone calls, or for multi-way calls such as online meetings. When the impersonation detection system 1 is used for calls between multiple parties, the impersonation detection system 1 may be equipped with three or more terminal devices. For example, the impersonation detection system 1 may be equipped with multiple first terminal devices 100 and one or more second terminal devices 200. Alternatively, the impersonation detection system 1 may be equipped with one or more first terminal devices 100 and multiple second terminal devices 200. Alternatively, the impersonation detection system 1 may be equipped with three or more terminal devices that combine the functions of the first terminal device 100 and the functions of the second terminal devices 200.

[0027] The second terminal device 200 is used for communication with the first terminal device 100. The second terminal device 200 also displays the judgment results from the analysis server device 300. As described above, the second terminal device 200 may have the function of acquiring air conduction sound data and bone conduction sound data of the second user, and the analysis server device 300 may also determine whether or not the second user's speech is being impersonated using synthesized speech. The second terminal device 200 may be configured using, for example, a smartphone or a personal computer.

[0028] Figure 2 shows an example of the configuration of the first terminal device 100. In the configuration shown in Figure 2, the first terminal device 100 comprises a first terminal-side communication unit 110, an air conduction sound data acquisition unit 120, and a bone conduction sound data acquisition unit 130. The first terminal-side communication unit 110 communicates with other devices. In particular, the first terminal-side communication unit 110 transmits air conduction sound data and bone conduction sound data to the analysis server device 300.

[0029] The air conduction sound data acquisition unit 120 is configured using a microphone and acquires air conduction sound data of the voice emitted by the first user. The air conduction sound data acquisition unit 120 is an example of an air conduction sound data acquisition means. The bone conduction sound data acquisition unit 130 is configured using a microphone that is in direct contact with the human body, such as the neck, and acquires bone conduction sound data from the vibration of the vocal cords. The bone conduction sound data acquisition unit 130 is an example of a bone conduction sound data acquisition means.

[0030] The configuration of the microphones for the air conduction sound data acquisition unit 120 and the bone conduction sound data acquisition unit 130 is not limited to a specific configuration. For example, the microphones for the air conduction sound data acquisition unit 120 and the bone conduction sound data acquisition unit 130 may be integrated with the earphone, but are not limited to this.

[0031] Figure 3 shows an example of the configuration of the second terminal device 200. In the configuration shown in Figure 3, the second terminal device 200 comprises a second terminal side communication unit 210, an audio output unit 220, and a display unit 230. The second terminal-side communication unit 210 communicates with other devices. In particular, the second terminal-side communication unit 210 receives the results of the determination of whether or not impersonation using synthesized speech is occurring from the analysis server device 300. The second terminal-side communication unit 210 also receives air conduction sound data transmitted by the first terminal device 100. The second terminal-side communication unit 210 may also receive air conduction sound data via the analysis server device 300.

[0032] The audio output unit 220 is configured using a speaker and outputs sound. In particular, the audio output unit 220 outputs the voice of the first user by converting air conduction sound data transmitted from the first terminal device 100 into sound. The display unit 230 is equipped with a display screen such as an LCD panel, an LED (Light Emitting Diode) panel, or an organic EL (Organic Electro-Luminescence) panel, and displays various images. In particular, the display unit 230 displays the result of the determination by the analysis server device 300 regarding the presence or absence of impersonation using synthesized speech. The display unit 230 is an example of a display means.

[0033] Figure 4 shows an example of the configuration of the analysis server device 300. In the configuration shown in Figure 4, the analysis server device 300 comprises a server-side communication unit 310, an audio data analysis unit 320, an audio data storage unit 330, and an analysis processing unit 340. The server-side communication unit 310 communicates with other devices. In particular, the server-side communication unit 310 receives air conduction sound data and bone conduction sound data transmitted by the first terminal device 100. The server-side communication unit 310 also transmits the result of the determination of whether or not impersonation using synthesized speech is occurring to the second terminal device 200. Furthermore, the server-side communication unit 310 may also transmit air conduction sound data from the first terminal device 100 to the second terminal device 200. The air conduction sound data from the first terminal device 100 corresponds to the air conduction sound data to be judged.

[0034] The voice data storage unit 330 stores voice data obtained by combining air conduction sound data and bone conduction sound data. The voice data storage unit 330 may dynamically store combinations of air conduction sound data and bone conduction sound data obtained during a call using the impersonation detection system 1. Alternatively, the voice data storage unit 330 may store data that is predetermined as voice data obtained by combining air conduction sound data and bone conduction sound data. The audio data stored in the audio data storage unit 330, which consists of a combination of air conduction sound data and bone conduction sound data, is also referred to as registered data. The air conduction sound data within the registered data is also referred to as registered air conduction sound data. The bone conduction sound data within the registered data is also referred to as registered bone conduction sound data. The air conduction sound indicated by the registered air conduction sound data is also referred to as registered air conduction sound. The bone conduction sound indicated by the registered bone conduction sound data is also referred to as registered bone conduction sound.

[0035] The voice data analysis unit 320 compares the air conduction sound data and bone conduction sound data from the first terminal device 100 to determine whether or not these data correspond. The voice data analysis unit 320 is an example of a voice data analysis means. For example, the audio data analysis unit 320 may perform correction processing to compare air conduction sound data and bone conduction sound data on either the air conduction sound data or the bone conduction sound data, or on both of them.

[0036] As a correction process for comparing air conduction sound data and bone conduction sound data, the speech data analysis unit 320 may perform a process on the air conduction sound data to attenuate the high-frequency components of the air conduction sound. Bone conduction sound tends to have lower high-frequency components than air conduction sound, and it is expected that by attenuating the high-frequency components of the air conduction sound, the air conduction sound and bone conduction sound will become closer (more similar). As the air conduction sound and bone conduction sound become closer, it is expected that the speech data analysis unit 320 will be able to determine with relatively high accuracy whether or not the air conduction sound and bone conduction sound correspond to each other.

[0037] The voice data analysis unit 320 then determines whether the corrected air conduction sound data and bone conduction data correspond to each other. For example, the voice data analysis unit 320 may apply the corrected air conduction sound data and bone conduction sound data to an authentication algorithm for voiceprint authentication using air conduction sound to determine whether the corrected air conduction sound and bone conduction sound are the voices of the same person.

[0038] However, the method by which the audio data analysis unit 320 determines whether the air conduction sound data and the bone conduction sound data correspond is not limited to a specific method. For example, the audio data analysis unit 320 may perform pattern matching between the corrected air conduction sound waveform and the bone conduction waveform and determine whether the difference is within a predetermined range. In this case, in addition to performing a process on the air conduction sound data to attenuate the high-frequency components of the air conduction sound, or alternatively, the audio data analysis unit 320 may perform a process on the bone conduction sound data to amplify the high-frequency components of the bone conduction sound. Furthermore, the voice data analysis unit 320 may use a known authentication algorithm that utilizes air-conducted sound and bone-conducted sound to determine whether the air-conducted sound and bone-conducted sound are from the speech of the same person.

[0039] Furthermore, the voice data analysis unit 320 compares the air conduction sound data and bone conduction sound data to be determined (air conduction sound data and bone conduction sound data from the first terminal device 100) with the registered data (air conduction sound data and bone conduction sound data stored in the voice data storage unit 330). Based on this, the voice data analysis unit 320 determines whether the air conduction sound data and bone conduction sound data to be determined were obtained from the speech of the same person.

[0040] Specifically, the voice data analysis unit 320 compares the air conduction sound data to be judged with the registered air conduction sound data and searches for registered air conduction sound data obtained from the speech of the same person as the air conduction sound data to be judged.

[0041] If registered air conduction sound data obtained from the speech of the same person as the air conduction sound data to be judged is detected, the voice data analysis unit 320 acquires the registered bone conduction sound data (registered bone conduction sound data paired with the registered air conduction sound data) that is associated with the detected registered air conduction sound data. The voice data analysis unit 320 then determines whether the bone conduction sound data to be judged and the acquired registered bone conduction sound data are data obtained from the speech of the same person.

[0042] If the speech data analysis unit 320 determines that the bone conduction sound data to be analyzed and the acquired registered bone conduction sound data are data obtained from the speech of the same person, it determines that impersonation by speech synthesis has not occurred. If the bone conduction sound data to be analyzed and the acquired registered bone conduction sound data are obtained from the speech of the same person, it is considered that the air conduction sound data to be analyzed, the acquired registered air conduction sound data, the acquired registered bone conduction sound data, and the bone conduction sound data to be analyzed are all data obtained from the speech of the same person.

[0043] On the other hand, if the bone conduction sound data being analyzed and the acquired registered bone conduction sound data are determined not to be data obtained from the speech of the same person, the voice data analysis unit 320 determines that impersonation using speech synthesis is taking place.

[0044] The method by which the voice data analysis unit 320 determines whether the air conduction sound data to be judged and the registered air conduction sound data are obtained from the speech of the same person is not limited to a specific method. For example, the voice data analysis unit 320 may use a known voiceprint authentication method to determine whether the air conduction sound data to be judged and the registered air conduction sound data are obtained from the speech of the same person, but is not limited to this.

[0045] The method by which the voice data analysis unit 320 determines whether the bone conduction sound data to be judged and the registered bone conduction sound data are obtained from the speech of the same person is not limited to a specific method. For example, the voice data analysis unit 320 may amplify the high-frequency components of the bone conduction sound to be judged and the registered bone conduction sound, and apply a known voiceprint authentication method to determine whether the bone conduction sound data to be judged and the registered bone conduction sound data are obtained from the speech of the same person, but is not limited to this.

[0046] The analysis processing unit 340 outputs a determination result regarding the presence or absence of impersonation using synthesized speech, based on the determination result from the voice data analysis unit 320. For example, the analysis processing unit 340 may generate data notifying the determination result regarding the presence or absence of impersonation using speech synthesis and transmit it to the second terminal device 200 via the server-side communication unit 310. Furthermore, when the analysis server device 300 determines that there is impersonation using speech synthesis and terminates the call between the first terminal device 100 and the second terminal device 200, the voice data analysis unit 320 may perform the process of terminating the call, or the analysis processing unit 340 may perform the process of terminating the call.

[0047] Figure 5 shows an example of the procedure for processing a call in the impersonation detection system 1. In the process shown in Figure 5, the first terminal device 100, the second terminal device 200, and the analysis server device 300 initiate a call (step S101). For example, the first terminal device 100 and the analysis server device 300 establish a communication connection, and the analysis server device 300 establishes a communication connection with the second terminal device 200. During a call, the bone conduction sound data acquisition unit is assumed to be in direct contact with the body of the first user.

[0048] Next, the first terminal device 100, the second terminal device 200, and the analysis server device 300 repeat the call processing (step S102) for each step in real time until the call ends.

[0049] In the call processing per step (step S102), the air conduction sound data acquisition unit 120 of the first terminal device 100 acquires air conduction sound data of the voice emitted by the first user, and the bone conduction sound data acquisition unit 130 acquires bone conduction sound data due to the vibration of the first user's vocal cords (step S103). The first terminal-side communication unit 110 transmits the air conduction sound data and bone conduction sound data obtained in step S103 to the analysis server device 300 (step S104).

[0050] In the analysis server device 300, the server-side communication unit 310 transmits the air conduction sound data from the air conduction sound data and bone conduction sound data received from the first terminal device 100 to the second terminal device 200 (step S105). In the second terminal device 200, the audio output unit 220 outputs audio by converting the air conduction sound data received by the second terminal side communication unit 210 into speech (step S106).

[0051] Steps S103 to S106 correspond to one step of call processing (step S102). The first terminal device 100, the second terminal device 200, and the analysis server device 300 repeat the call processing (step S102) in real time so that the second terminal device 200 can continuously output audio in response to the first user's utterance.

[0052] When the call termination condition is met, the first terminal device 100, the second terminal device 200, and the analysis server device 300 terminate the call (step S107). The call termination condition may also be that the first user or the second user performs a call termination operation. In addition, if the analysis server device 300 determines that there is impersonation using synthesized speech, the call termination condition may be met and the call may be terminated. At the end of the call (step S107), for example, the first terminal device 100 and the analysis server device 300 disconnect their communication connection, and the analysis server device 300 and the second terminal device 200 disconnect their communication connection.

[0053] Figure 6 shows an example of the procedure for the analysis server device 300 to determine whether or not impersonation using synthesized speech has occurred. The analysis server device 300 may perform the process shown in Figure 6 a predetermined number of times during a single call, for example, at the start of a call. Alternatively, the analysis server device 300 may perform the process shown in Figure 6 repeatedly during a single call, for example, after a predetermined number of executions of the call processing per step (step S102) in Figure 5.

[0054] In the process shown in Figure 6, the server-side communication unit 310 receives the air conduction sound data and bone conduction sound data transmitted by the first terminal-side communication unit 110 in step S104 of Figure 5 (step S111). Next, the voice data analysis unit 320 analyzes the air conduction sound data and bone conduction sound data obtained in step S111 (step S112) and determines whether these data correspond to each other (step S113).

[0055] If, in step S113, it is determined that the air conduction sound data and the bone conduction sound data correspond (step S113: YES), the voice data analysis unit 320 determines whether there is registered air conduction sound data obtained from the speech of the same person as the air conduction sound data obtained in step S111 (step S121). For example, as described above, the voice data analysis unit 320 searches for registered air conduction sound data obtained from the speech of the same person as the air conduction sound data being determined.

[0056] In step S121, if it is determined that there is registered air conduction sound data obtained from the speech of the same person as the air conduction sound data obtained in step S111 (step S121: YES), the speech data analysis unit 320 acquires the registered bone conduction sound data that is paired with that registered air conduction sound (step S131). The voice data analysis unit 320 then analyzes the bone conduction sound data obtained in step S111 and the registered bone conduction sound data obtained in step S131 (step S132), and determines whether or not these bone conduction sound data were obtained from the speech of the same person (step S133).

[0057] If the bone conduction sound data obtained in step S111 and the registered bone conduction sound data obtained in step S131 are determined to be data obtained from the speech of the same person (step S133: YES), the analysis processing unit 340 performs a predetermined process for cases where there is no impersonation using speech synthesis (step S151). For example, the analysis processing unit 340 may transmit the determination result that there is no impersonation using synthesized speech to the second terminal device 200 via the server-side communication unit 310. The second terminal-side communication unit 210 of the second terminal device 200 then receives the determination result, and the display unit 230 displays the determination result received by the second terminal-side communication unit 210. In addition to, or instead of, the analysis processing unit 340, the voice data analysis unit 320 may perform processing in cases where there is no impersonation using speech synthesis. After step S151, the analysis server device 300 completes the process shown in Figure 6.

[0058] On the other hand, if in step S133 it is determined that the bone conduction sound data obtained in step S111 and the registered bone conduction sound data obtained in step S131 are not data obtained from the speech of the same person (step S133: NO), the analysis processing unit 340 performs a predetermined process for when there is impersonation using speech synthesis (step S161). For example, the analysis processing unit 340 may send a message to the second terminal device 200 via the server-side communication unit 310 indicating that impersonation using synthesized speech has occurred. Furthermore, the analysis processing unit 340 may terminate the call between the first terminal device 100 and the second terminal device 200. In addition to, or instead of, the analysis processing unit 340 may be configured to perform processing in cases of impersonation using speech synthesis. After step S161, the analysis server device 300 completes the process shown in Figure 6.

[0059] On the other hand, if in step S121 the speech data analysis unit 320 determines that there is no registered air conduction sound data obtained from the speech of the same person as the air conduction sound obtained in step S111 (step S121: NO), the analysis processing unit 340 adds the combination of air conduction sound data and bone conduction sound data obtained in step S111 to the registered data (step S141). That is, the analysis processing unit 340 stores the combination of air conduction sound data and bone conduction sound data obtained in step S111 as new registered data in the speech data storage unit 330. After step S141, the process transitions to step S151. On the other hand, if the voice data analysis unit 320 determines in step S113 that the air conduction sound data and bone conduction sound data do not correspond (step S113: NO), the process proceeds to step S161.

[0060] In steps S151 and S161 of Figure 6, the analysis server device 300 may notify the second terminal device 200 of the presence or absence of registered data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100. If the transition from step S113:NO to step S161 is made, the analysis server device 300 may notify the second terminal device 200 that it has not searched for registered data.

[0061] If there is registration data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100, it can be assumed that there is a call history by the user of the first terminal device 100, and it is considered more likely that impersonation has not occurred. Step S133: If YES, it corresponds to the case where there is registration data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100.

[0062] On the other hand, if there is no registration data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100, it cannot be said that there is a higher probability that impersonation has not occurred. Step S121: NO corresponds to the case where there is no registration data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100. Steps S113: NO and S133: NO can also be considered as cases where there is no registration data corresponding to the air conduction sound data and bone conduction sound data from the first terminal device 100.

[0063] Thus, the impersonation detection system 1 simultaneously acquires air conduction sound and bone conduction sound as the speaker's voice, and determines whether the air conduction sound and bone conduction sound correspond to each other, thereby determining whether the voice in the call is synthesized speech or spoken by a human. According to the impersonation detection system 1, by using multiple sounds such as air conduction sound, bone conduction sound, and past voice data of each, it is expected to be able to determine with high accuracy whether or not impersonation using speech synthesis is occurring. According to the impersonation detection system 1, for example, it is possible to detect impersonation using speech synthesis in remote calls. Furthermore, it is expected that the impersonation detection system 1 can prevent fraud caused by impersonation using speech synthesis.

[0064] Furthermore, if the analysis server device 300 determines that impersonation using speech synthesis is occurring, it is possible that other forms of fraud are actually taking place, such as human impersonation or the unauthorized attachment of a microphone for acquiring bone conduction data. For this reason, the analysis server device 300 may be configured to determine if there is a possibility of impersonation using speech synthesis. Alternatively, the analysis server device 300 may be configured to determine if some form of fraud is occurring.

[0065] As described above, the voice data analysis unit 320 determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data (registered data). According to the analysis server device 300, it is possible to determine with relatively high accuracy whether or not impersonation using synthesized speech is occurring. In particular, according to the analysis server device 300, it is possible to determine whether or not there is stored air conduction sound data and bone conduction sound data corresponding to the air conduction sound data and bone conduction sound data to be determined, and thereby determine whether or not there is a call history of the user of the first terminal device 100. If there is a call history of the user of the first terminal device 100, it can be determined that the possibility of impersonation using speech synthesis is lower than if there is no call history of the user of the first terminal device 100.

[0066] Furthermore, the voice data analysis unit 320 determines whether the air conduction sound data and bone conduction sound data to be evaluated are data obtained from the speech of the same person. If it determines that the air conduction sound data and bone conduction sound data to be evaluated are data obtained from the speech of the same person, it determines whether or not impersonation using synthesized speech has occurred based on a comparison of the air conduction sound data and bone conduction sound data to be evaluated with the stored air conduction sound data and bone conduction sound data. According to the analysis server device 300, the presence or absence of impersonation using speech synthesis is determined not only by comparing the air conduction sound data and bone conduction sound data to be judged with the stored air conduction sound data and bone conduction sound data, but also by determining whether the air conduction sound data and bone conduction sound data to be judged were obtained from the speech of the same person. This is expected to enable more accurate determination of whether or not impersonation using speech synthesis is occurring.

[0067] <Second Embodiment> The impersonation detection system may further use biometric information to determine whether or not impersonation is occurring using synthesized speech. This point will be explained in the second embodiment. Figure 7 shows an example of the configuration of a spoofing detection system according to at least one embodiment. In the configuration shown in Figure 7, the spoofing detection system 2 comprises a first terminal device 102, a second terminal device 200, and an analysis server device 302.

[0068] In Figure 7, parts that have the same function as those in Figure 1 are denoted by the same reference numeral (200), and detailed explanations are omitted here. In the impersonation detection system 2, the first terminal device 102 acquires biometric information of the first user in addition to the processing performed by the first terminal device 100. Furthermore, the analysis server device 302, in addition to the processing performed by the analysis server device 300, uses the biometric information to determine whether the speaker is human or not (whether the air conduction sound data and bone conduction sound data to be determined are data obtained from human speech or not). The analysis server device 302 also uses the air conduction sound data or bone conduction sound data, or both, along with the biometric information to determine whether or not there is impersonation by speech synthesis. In all other respects, the impersonation detection system 2 is the same as the impersonation detection system 1.

[0069] Figure 8 shows an example of the configuration of the first terminal device 102. In the configuration shown in Figure 8, the first terminal device 102 comprises a first terminal side communication unit 110, an air conduction sound data acquisition unit 120, a bone conduction sound data acquisition unit 130, and a biological information acquisition unit 140.

[0070] In Figure 8, parts that have the same function as those in Figure 2 are denoted by the same reference numerals (110, 120, 130), and detailed explanations are omitted here. The first terminal device 102 includes a biometric information acquisition unit 140 in addition to the configuration of the first terminal device 100. Accordingly, in the first terminal device 102, the first terminal-side communication unit 110 transmits biometric information in addition to air conduction sound data and bone conduction sound data to the analysis server device 302. In all other respects, the first terminal device 102 is the same as the first terminal device 100.

[0071] Similar to the case of the first terminal device 100, the user of the first terminal device 102 is also referred to as the first user. Similar to the case of the first terminal device 100, the air conduction sound data from the first terminal device 102 corresponds to the air conduction sound data subject to determination (air conduction sound data subject to determination of whether or not there is impersonation by synthesized voice). The bone conduction sound data from the first terminal device 102 corresponds to the bone conduction sound data subject to determination (bone conduction sound data subject to determination of whether or not there is impersonation by synthesized voice).

[0072] The biometric information acquisition unit 140 acquires the biometric information of the first user. The biometric information acquired by the biometric information acquisition unit 140 is not limited to a specific type. For example, the biometric information acquisition unit 140 may acquire the first user's body temperature, pulse, and blood oxygen concentration, or some of these, but is not limited to this. The more types of biometric information acquired by the biometric information acquisition unit 140, the more the accuracy of the judgment performed by the analysis server device 302 is expected to improve. The biological information acquisition unit 140 may be integrated with the bone conduction sound data acquisition unit 130, or they may be configured separately.

[0073] Figure 9 shows an example of the configuration of the analysis server device 302. In the configuration shown in Figure 9, the analysis server device 302 comprises a server-side communication unit 310, a voice data analysis unit 320, a voice data storage unit 330, a biometric information analysis unit 350, and an analysis processing unit 360.

[0074] In Figure 9, parts that have the same function as those in Figure 4 are denoted by the same reference numerals (310, 320, 330), and detailed explanations are omitted here. The analysis server device 302 includes a biometric information analysis unit 350 in addition to the configuration of the analysis server device 300. The analysis processing unit 360 determines whether or not impersonation using synthesized speech is occurring based on air conduction sound data, bone conduction sound data, and biometric information. In the analysis server device 302, the server-side communication unit 310 receives air conduction sound data, bone conduction sound data, and biometric information transmitted by the first terminal device 102. In all other respects, the analysis server device 302 is the same as the analysis server device 300.

[0075] The biometric information analysis unit 350 determines whether there are any abnormal values ​​in the biometric information values ​​received by the server-side communication unit 310. Based on this, the biometric information analysis unit 350 determines whether the first user is human or not. The biometric information analysis unit 350 is an example of a biometric information analysis means. For example, the biometric information analysis unit 350 acquires the body temperature data of the first user from the first terminal device 102 via the server-side communication unit 310. The biometric information analysis unit 350 then compares the first user's body temperature with predetermined upper and lower thresholds to determine whether the first user's body temperature is within the normal range (within the range predetermined as a normal human body temperature).

[0076] However, the biological information used by the biological information analysis unit 350 to determine whether or not there are abnormal values ​​is not limited to a specific type of biological information. For example, the biological information analysis unit 350 may determine, in addition to or instead of the first user's body temperature, whether or not there are abnormalities in the first user's heart rate (whether or not the heart rate is an abnormal value).

[0077] When a user of the first terminal device 102 and a user of the second terminal device 200 make a call, the biometric information analysis unit 350 may terminate the call if it determines that there is an abnormal value in the biometric information. In this case, the process of terminating the call may be performed by the voice data analysis unit 320, the biometric information analysis unit 350, or the analysis processing unit 360. Alternatively, the biological information analysis unit 350 may transmit the result of determining whether there are any abnormal values ​​in the biological information to the second terminal device 200 via the server-side communication unit 310.

[0078] The analysis processing unit 360 determines whether or not impersonation using speech synthesis is occurring, based on biological information and air conduction sound data or bone conduction sound data, or both. The analysis processing unit 360 is an example of an analysis processing means. Specifically, the analysis processing unit 360 estimates the attributes of the first user using biometric information. The analysis processing unit 360 also estimates the attributes of the first user using air conduction sound data, bone conduction sound data, or both. The analysis processing unit 360 then determines whether the attributes estimated using air conduction sound data, bone conduction sound data, or both match the attributes estimated using biometric information. If the analysis processing unit 360 determines that the attributes estimated using biometric information match the attributes estimated using air conduction sound data, bone conduction sound data, or both, the analysis processing unit 360 determines that there is no impersonation using speech synthesis. On the other hand, if the analysis processing unit 360 determines that the attributes estimated using biometric information do not match the attributes estimated using air conduction sound data, bone conduction sound data, or both, the analysis processing unit 360 determines that there is impersonation using speech synthesis.

[0079] For example, the analysis processing unit 360 acquires information on the age and gender of the first user. The analysis processing unit 360 may also estimate the age and gender of the first user based on the biometric information of the first user acquired from the first terminal device 102 via the server-side communication unit 310. Furthermore, for example, the analysis processing unit 360 may acquire facial image data of the first user and estimate the gender and age of the first user based on the facial image.

[0080] However, the attributes of the first user estimated by the analysis processing unit 360 are not limited to a specific type. The biometric information that the analysis processing unit 360 uses to estimate attributes is also not limited to a specific type of biometric information, and can be various types of biometric information depending on the attribute to be estimated. Furthermore, the method used by the analysis processing unit 360 to estimate the attributes of the first user is not limited to a specific method. For example, the analysis processing unit 360 may use a known method to estimate the attributes of the first user.

[0081] Furthermore, the analysis processing unit 360 estimates the age and gender of the first user based on the air conduction sound data to be determined. The analysis processing unit 360 may also estimate the age and gender of the first user based on bone conduction sound data to be determined, in addition to or instead of the air conduction sound data to be determined. The method by which the analysis processing unit 360 estimates the age and gender of the first user based on the air conduction sound data to be determined, the bone conduction sound data to be determined, or a combination thereof, is not limited to a specific method. For example, the analysis processing unit 360 may estimate the attributes of the first user using a known method. The air conduction sound data, bone conduction sound data, or a combination thereof that is subject to evaluation is also referred to as the audio data subject to evaluation.

[0082] The analysis processing unit 360 then determines whether the attributes estimated from the biometric information match the attributes estimated from the voice data to be judged. If it determines that the attributes estimated from the biometric information match the attributes estimated from the voice data to be judged, the analysis processing unit 360 determines that there is no impersonation using speech synthesis. On the other hand, if it determines that the attributes estimated from the biometric information do not match the attributes estimated from the voice data to be judged, the analysis processing unit 360 determines that there is impersonation using speech synthesis.

[0083] If the analysis processing unit 360 determines whether each of several attributes, such as age and gender, matches, it may determine that there is no impersonation by speech synthesis if all of the attributes match the attributes estimated from the biometric information and the attributes estimated from the voice data being judged. On the other hand, if the analysis processing unit 360 does not match the attributes estimated from the biometric information and the attributes estimated from the voice data being judged for any one or more of the multiple attributes, it may determine that there is impersonation by speech synthesis.

[0084] Alternatively, the analysis processing unit 360 may determine that there is no voice synthesis impersonation if, for one or more of the multiple attributes, the attribute estimated from the biometric information matches the attribute estimated from the voice data to be judged. On the other hand, the analysis processing unit 360 may determine that there is voice synthesis impersonation if, for all of the multiple attributes, the attribute estimated from the biometric information does not match the attribute estimated from the voice data to be judged.

[0085] Alternatively, the analysis processing unit 360 may estimate the likelihood of speech synthesis impersonation. For example, the analysis processing unit 360 may calculate the likelihood of speech synthesis impersonation as the ratio obtained by dividing the number of attributes that are determined not to match the attributes estimated from the biometric information by the number of attributes that are subject to matching.

[0086] In this case, the analysis processing unit 360 may notify the second terminal device 200 via the server-side communication unit 310 of the high probability that voice synthesis impersonation is occurring. Furthermore, the analysis processing unit 360 may compare the likelihood of speech synthesis impersonation with a predetermined threshold. If the analysis processing unit 360 determines that the likelihood of speech synthesis impersonation is higher than or equal to the threshold, it may determine that speech synthesis impersonation has occurred. On the other hand, if the analysis processing unit 360 determines that the likelihood of speech synthesis impersonation is lower than the threshold, it may determine that speech synthesis impersonation has not occurred.

[0087] The analysis server device 302 (particularly the combination of the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360) comprehensively determines whether or not impersonation using synthesized speech is occurring by using voice data such as air conduction sound data and bone conduction sound data, and biometric information. As described above, the voice data analysis unit 320 determines whether or not impersonation using speech synthesis is occurring based on whether or not the air conduction sound data to be judged and the bone conduction sound data to be judged correspond to each other. The biometric information analysis unit 350 determines whether or not impersonation using speech synthesis is occurring based on whether or not there are abnormal values ​​in the biometric information values ​​for a given user. The analysis processing unit 360 determines whether or not impersonation using speech synthesis is occurring based on whether or not there is registered data corresponding to the air conduction sound data and bone conduction sound data to be judged.

[0088] For example, if the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 all determine that there is no impersonation using speech synthesis, the analysis server device 302 may determine that there is no impersonation using speech synthesis. On the other hand, if one or more of the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 determine that there is impersonation using speech synthesis, the analysis server device 302 may determine that there is impersonation using speech synthesis.

[0089] Alternatively, if one or more of the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 determine that there is no impersonation using speech synthesis, the analysis server device 302 may determine that there is no impersonation using speech synthesis. On the other hand, if any of the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 determine that there is impersonation using speech synthesis, the analysis server device 302 may determine that there is impersonation using speech synthesis.

[0090] Alternatively, the analysis server device 302 may be configured to estimate the likelihood of speech synthesis impersonation. For example, the analysis server device 302 may calculate the likelihood of speech synthesis impersonation as the ratio obtained by dividing the number of parts of the speech data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 that are determined to have speech synthesis impersonation by 3 (the number of these parts).

[0091] In this case, the analysis server device 302 may be configured to notify the second terminal device 200 of the high probability that speech synthesis-based impersonation is occurring. Furthermore, the analysis server device 302 may compare the likelihood of speech synthesis impersonation with a predetermined threshold. If the analysis server device 302 determines that the likelihood of speech synthesis impersonation is higher than the threshold, it may determine that speech synthesis impersonation is occurring. On the other hand, if the analysis server device 302 determines that the likelihood of speech synthesis impersonation is below the threshold, it may determine that speech synthesis impersonation is not occurring.

[0092] The voice data analysis unit 320 may perform a comprehensive determination of whether or not there is impersonation using speech synthesis. Alternatively, the biometric information analysis unit 350 may perform a comprehensive determination of whether or not there is impersonation using speech synthesis. Alternatively, the analysis processing unit 360 may perform a comprehensive determination of whether or not there is impersonation using speech synthesis.

[0093] Alternatively, the combination of the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 may be configured to make a comprehensive determination of whether or not there is impersonation by voice synthesis. For example, the voice data analysis unit 320, the biometric information analysis unit 350, and the analysis processing unit 360 may perform the comprehensive determination of whether or not there is impersonation by voice synthesis using distributed processing.

[0094] Figure 10 shows an example of the procedure for processing a call in the impersonation detection system 2. In the process shown in Figure 10, the first terminal device 102, the second terminal device 200, and the analysis server device 302 initiate a call (step S201). For example, the first terminal device 102 and the analysis server device 302 establish a communication connection, and the analysis server device 302 establishes a communication connection with the second terminal device 200. During a call, the bone conduction sound data acquisition unit 130 is assumed to be in direct contact with the body of the first user.

[0095] Next, the first terminal device 102, the second terminal device 200, and the analysis server device 302 repeat the call processing per step (step S202) in real time until the call ends.

[0096] In the call processing per step (step S202), the air conduction sound data acquisition unit 120 of the first terminal device 102 acquires air conduction sound data of the voice emitted by the first user, the bone conduction sound data acquisition unit 130 acquires bone conduction sound data due to the vibration of the first user's vocal cords, and the biometric information acquisition unit 140 acquires the biometric information of the first user (step S203). The first terminal-side communication unit 110 transmits the air conduction sound data, bone conduction sound data, and biological information obtained in step S203 to the analysis server device 300 (step S204).

[0097] In the analysis server device 302, the server-side communication unit 310 transmits the air conduction sound data, bone conduction sound data, and biometric information received from the first terminal device 102 to the second terminal device 200 (step S205). In the second terminal device 200, the audio output unit 220 outputs audio by converting the air conduction sound data received by the second terminal side communication unit 210 into speech (step S206).

[0098] Steps S203 to S206 correspond to one step of call processing (step S202). The first terminal device 102, the second terminal device 200, and the analysis server device 302 repeat the call processing (step S202) in real time so that the second terminal device 200 can continuously output audio in response to the first user's speech.

[0099] When the call termination condition is met, the first terminal device 102, the second terminal device 200, and the analysis server device 302 terminate the call (step S207). The call termination condition may also be that the first user or the second user performs a call termination operation. Furthermore, if the analysis server device 302 determines that there is an impersonation attempt using synthesized speech, the call termination condition may be met and the call may be terminated. If the analysis server device 302 determines that there is an abnormal value in the biometric information, the call termination condition may be met and the call may be terminated. At the end of the call (step S207), for example, the first terminal device 102 and the analysis server device 302 disconnect their communication connection, and the analysis server device 302 and the second terminal device 200 disconnect their communication connection.

[0100] Figures 11 and 12 show examples of the procedure for determining whether or not impersonation using synthesized speech is occurring in the analysis server device 302. The analysis server device 302 may perform the processes shown in Figures 11 and 12 a predetermined number of times during a single call, for example, at the start of a call. Alternatively, the analysis server device 300 may perform the processes shown in Figures 11 and 12 repeatedly during a single call, for example, after a predetermined number of executions of the call processing per step (step S202) in Figure 10.

[0101] In the processes shown in Figures 11 and 12, the server-side communication unit 310 receives the air conduction sound data, bone conduction sound data, and biological information transmitted by the first terminal-side communication unit 110 in step S204 of Figure 10 (step S211). Next, the voice data analysis unit 320 analyzes the air conduction sound data and bone conduction sound data obtained in step S211 (step S212) and determines whether these data correspond to each other (step S213).

[0102] If, in step S213, it is determined that the air conduction sound data and bone conduction sound data correspond (step S213: YES), the speech data analysis unit 320 determines whether there is registered air conduction sound data obtained from the speech of the same person as the air conduction sound data obtained in step S211 (step S221). For example, the speech data analysis unit 320 searches for registered air conduction sound data obtained from the speech of the same person as the air conduction sound data being determined.

[0103] In step S221, if it is determined that there is registered air conduction sound data obtained from the speech of the same person as the air conduction sound data obtained in step S211 (step S221: YES), the voice data analysis unit 320 acquires the registered bone conduction sound data that is paired with that registered air conduction sound (step S231). Then, the voice data analysis unit 320 analyzes the bone conduction sound data obtained in step S211 and the registered bone conduction sound data obtained in step S231 (step S232), and determines whether or not these bone conduction sound data were obtained from the speech of the same person (step S233).

[0104] If the bone conduction sound data obtained in step S211 and the registered bone conduction sound data obtained in step S231 are determined to be data obtained from the speech of the same person (step S233: YES), the bio-information analysis unit 350 determines whether there are any abnormal values ​​in the bio-information values ​​obtained in step S211 as human bio-information (step S251). For example, the bio-information analysis unit 350 may compare the bio-information with upper and lower threshold values ​​predetermined for each type of bio-information. The bio-information analysis unit 350 may then determine that there are abnormal values ​​if it determines that there is bio-information that is greater than the upper threshold or bio-information that is smaller than the lower threshold.

[0105] If the biological information analysis unit 350 determines in step S251 that there are no abnormal values ​​in the biological information (step S251: NO), the analysis processing unit 360 estimates the attributes of the first user based on the data obtained in step S211 (step S261). Specifically, the analysis processing unit 360 estimates the attributes of the first user based on the air conduction sound data or bone conduction sound data, or both, obtained in step S211. The analysis processing unit 360 also estimates the attributes of the first user based on the biological information obtained in step S211.

[0106] Then, the analysis processing unit 360 determines whether the attributes estimated based on the air conduction sound data or bone conduction sound data obtained in step S211, or both thereof, match the attributes estimated based on the biological information obtained in step S211 (step S262).

[0107] If it is determined that the attribute values ​​match (step S262: YES), the analysis processing unit 360 performs a predetermined process for cases where there is no impersonation using speech synthesis (step S271). For example, the analysis processing unit 360 may transmit the determination result that there is no impersonation using synthesized speech to the second terminal device 200 via the server-side communication unit 310. The second terminal-side communication unit 210 of the second terminal device 200 then receives the determination result, and the display unit 230 displays the determination result received by the second terminal-side communication unit 210.

[0108] In addition to, or instead of, the analysis processing unit 360, the voice data analysis unit 320 or the biometric information analysis unit 350 may perform processing when there is no impersonation using speech synthesis. The analysis processing unit 360, the voice data analysis unit 320, and the biometric information analysis unit 350 may all perform processing when there is no impersonation using speech synthesis. After step S271, the analysis server device 302 completes the processing shown in Figures 11 and 12.

[0109] On the other hand, if step S262 determines that the attribute values ​​do not match (step S262: NO), the analysis processing unit 360 performs a predetermined process for when there is impersonation using speech synthesis (step S281). For example, the analysis processing unit 360 may send a message to the second terminal device 200 via the server-side communication unit 310 indicating that impersonation using synthesized speech has occurred. Furthermore, the analysis processing unit 340 may terminate the call between the first terminal device 102 and the second terminal device 200.

[0110] In addition to, or instead of, the analysis processing unit 360, the voice data analysis unit 320 or the biometric information analysis unit 350 may perform processing in cases of impersonation using speech synthesis. The analysis processing unit 360, the voice data analysis unit 320, and the biometric information analysis unit 350 may all perform processing in cases of impersonation using speech synthesis.

[0111] Furthermore, if the analysis processing unit 360 has obtained only a portion of the air conduction sound data, bone conduction sound data, and biological information in step S211, it sends a message to that effect to the second terminal device 200 via the server-side communication unit 310 (step S282). The display unit 230 of the second terminal device 200 may display this message, allowing the second user to decide whether or not to terminate the call.

[0112] The reason why only a portion of the air conduction sound data, bone conduction sound data, and biological information is obtained in step S211 is not limited to any specific reason. For example, the reason why only a portion of the air conduction sound data, bone conduction sound data, and biological information is obtained in step S211 may be because the first terminal device 102 is equipped with only a portion of the air conduction sound data acquisition unit 120, the bone conduction sound data acquisition unit 130, and the biological information acquisition unit 140, but is not limited to this. After step S282, the analysis server device 302 terminates the process shown in Figure 11.

[0113] On the other hand, if the biological information analysis unit 350 determines in step S251 that there is an abnormal value in the biological information (step S251: YES), the process proceeds to step S281. If, in step S233, the bio-information analysis unit 350 determines that the bone conduction sound data obtained in step S211 and the registered bone conduction sound data obtained in step S231 are not data obtained from the speech of the same person (step S233: NO), the process also proceeds to step S281. If the voice data analysis unit 320 determines in step S213 that the air conduction sound data and bone conduction sound data do not correspond (step S213: NO), the process also proceeds to step S281.

[0114] On the other hand, if the voice data analysis unit 320 determines in step S221 that there is no registered data corresponding to the combination of air conduction sound data and bone conduction sound data obtained in step S211 (step S221: NO), the analysis processing unit 360 adds the combination of air conduction sound data and bone conduction sound data obtained in step S211 to the registered data (step S241). That is, the analysis processing unit 360 stores the combination of air conduction sound data and bone conduction sound data obtained in step S211 as new registered data in the voice data storage unit 330. After step S241, the process transitions to step S251.

[0115] Thus, the impersonation detection system 2 uses biometric information in addition to air conduction sound data and bone conduction sound data to determine whether or not impersonation using synthesized speech is occurring. In this respect, the impersonation detection system 2 is expected to be able to determine whether or not impersonation using synthesized speech is occurring with higher accuracy.

[0116] As described above, the biometric information analysis unit 350 determines whether or not impersonation using synthesized speech has occurred based on whether or not the value of the biometric information is within a predetermined range for that type of biometric information. According to the analysis server device 302, the presence or absence of impersonation using speech synthesis can be determined with greater accuracy by comparing the air conduction sound data and bone conduction sound data to be judged with the stored air conduction sound data and bone conduction sound data, as well as by determining whether the value of the biological information is within a predetermined range for the value of that type of biological information.

[0117] Furthermore, the analysis processing unit 360 determines whether or not impersonation using speech synthesis is occurring, based on the attributes of the person being judged estimated based on the biometric information. According to the analysis server device 302, it is expected that the presence or absence of impersonation using speech synthesis can be determined with higher accuracy by comparing the air conduction sound data and bone conduction sound data to be judged with the stored air conduction sound data and bone conduction sound data, as well as by determining the presence or absence of impersonation using speech synthesis based on the attributes of the person being judged estimated based on biological information.

[0118] <Third Embodiment> Figure 13 shows an example of the configuration of a spoofing detection device according to at least one embodiment. In the configuration shown in Figure 13, the spoofing detection device 610 includes a voice data analysis unit 611.

[0119] In this configuration, the voice data analysis unit 611 determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data. The voice data analysis unit 611 is an example of a voice data analysis means.

[0120] The impersonation detection device 610 can determine with relatively high accuracy whether or not impersonation is occurring using synthesized speech. In particular, the impersonation detection device 610 can determine whether or not there is stored air conduction sound data and bone conduction sound data corresponding to the air conduction sound data and bone conduction sound data to be detected, thereby determining whether or not there is a call history of the speaker. If there is a call history of the speaker, it can be determined that the possibility of impersonation using speech synthesis is lower than if there is no call history of the speaker.

[0121] <Fourth Embodiment> Figure 14 shows an example of the configuration of a spoofing detection system according to at least one embodiment. In the configuration shown in Figure 14, the spoofing detection system 620 comprises a first terminal device 621, a spoofing detection device 624, and a second terminal device 626. The first terminal device 621 comprises an air conduction sound data acquisition unit 622 and a bone conduction sound data acquisition unit 623. The spoofing detection device 624 comprises a voice data analysis unit 625. The second terminal device 626 comprises a display unit 627.

[0122] In this configuration, the air conduction sound data acquisition unit 622 acquires the air conduction sound data to be judged. The bone conduction sound data acquisition unit 623 acquires the bone conduction sound data to be judged. The voice data analysis unit 625 determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction data. The display unit 627 displays the result of the determination of whether or not impersonation using synthesized speech has occurred. The air conduction sound data acquisition unit 622 is an example of an air conduction sound data acquisition means. The bone conduction sound data acquisition unit 623 is an example of a bone conduction sound data acquisition means. The voice data analysis unit 625 is an example of a voice data analysis means. The display unit 627 is an example of a display means.

[0123] The impersonation detection system 620 can determine with relatively high accuracy whether or not impersonation is occurring using synthesized speech. In particular, the impersonation detection system 620 can determine whether or not there is stored air conduction sound data and bone conduction sound data corresponding to the air conduction sound data and bone conduction sound data being detected, thereby determining whether or not there is a call history of the speaker. If there is a call history of the speaker, it can be determined that the possibility of impersonation using speech synthesis is lower than if there is no call history of the speaker.

[0124] <Fifth Embodiment> Figure 15 shows an example of the processing steps in a spoofing detection method according to at least one embodiment. The spoofing detection method shown in Figure 15 includes analyzing voice data (step S611). In analyzing the audio data (step S611), the computer determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data.

[0125] The impersonation detection method shown in Figure 15 can determine with relatively high accuracy whether or not impersonation using synthesized speech is occurring. In particular, the impersonation detection method shown in Figure 15 can determine whether or not there is stored air conduction sound data and bone conduction sound data corresponding to the air conduction sound data and bone conduction sound data being detected, thereby determining whether or not there is a call history of the speaker. If there is a call history of the speaker, the possibility of impersonation using speech synthesis can be determined to be lower than if there is no call history of the speaker.

[0126] Figure 16 shows an example of a computer configuration according to at least one embodiment. In the configuration shown in Figure 16, the computer 700 comprises a CPU 710, a main memory 720, an auxiliary memory 730, an interface 740, and a non-volatile recording medium 750.

[0127] One or more of the above-mentioned first terminal device 100, first terminal device 102, second terminal device 200, analysis server device 300, analysis server device 302, impersonation detection device 610, first terminal device 621, impersonation detection device 624, and second terminal device 626, or a part thereof, may be implemented in the computer 700. In that case, the operation of each processing unit described above is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, expands it in the main memory device 720, and executes the above processing according to the program. The CPU 710 also reserves memory areas in the main memory device 720 corresponding to each of the above-mentioned storage units according to the program. Communication between each device and other devices is performed by the interface 740 having a communication function and performing communication according to the control of the CPU 710. The interface 740 also has a port for the non-volatile recording medium 750 and reads information from the non-volatile recording medium 750 and writes information to the non-volatile recording medium 750.

[0128] When the first terminal device 100 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0129] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the first terminal device 100 according to the program. Communication between the first terminal device 100 and other devices by the first terminal-side communication unit 110 is performed by the interface 740 having a communication function and operating according to the control of the CPU 710. Interaction between the first terminal device 100 and the user by the air conduction sound data acquisition unit 120 and the bone conduction sound data acquisition unit 130, etc. is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device according to the control of the CPU 710.

[0130] When the first terminal device 102 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0131] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the first terminal device 102 according to the program. Communication between the first terminal device 102 and other devices by the first terminal-side communication unit 110 is performed by the interface 740 having a communication function and operating according to the control of the CPU 710. Interaction between the first terminal device 102 and the user by the air conduction sound data acquisition unit 120 and the bone conduction sound data acquisition unit 130, etc. is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device according to the control of the CPU 710.

[0132] When the second terminal device 200 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0133] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the second terminal device 200 according to the program. Communication between the second terminal device 200 and other devices by the second terminal-side communication unit 210 is performed by the interface 740 having a communication function and operating according to the control of the CPU 710. Interaction between the second terminal device 200 and the user by the audio output unit 220 and the display unit 230, etc. is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device according to the control of the CPU 710.

[0134] When the analysis server device 300 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0135] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the analysis server device 300, such as the voice data storage unit 330, according to the program. Communication between the analysis server device 300 and other devices by the server-side communication unit 310 is performed by the interface 740 having a communication function and operating according to the control of the CPU 710. Interaction between the server-side communication unit 310 and the user is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device according to the control of the CPU 710.

[0136] When the impersonation detection device 610 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0137] Furthermore, the CPU 710 reserves memory in the main memory 720 for the impersonation detection device 610 to process according to the program. Communication between the impersonation detection device 610 and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Interaction between the impersonation detection device 610 and the user is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device under the control of the CPU 710.

[0138] When the first terminal device 621 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0139] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the first terminal device 102 according to the program. Communication between the first terminal device 621 and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Interaction between the first terminal device 621 and the user by the air conduction sound data acquisition unit 622 and the bone conduction sound data acquisition unit 623, etc. is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device under the control of the CPU 710.

[0140] When the impersonation detection device 624 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0141] Furthermore, the CPU 710 reserves memory in the main memory 720 for the impersonation detection device 624 to process according to the program. Communication between the impersonation detection device 624 and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Interaction between the impersonation detection device 624 and the user is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device under the control of the CPU 710.

[0142] When the second terminal device 626 is implemented in the computer 700, its processing is stored in auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from auxiliary storage device 730, loads it into main memory 720, and executes the above processing according to the program.

[0143] Furthermore, the CPU 710 reserves memory in the main memory 720 for processing by the second terminal device 626 according to the program. Communication between the second terminal device 626 and other devices is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Interaction between the second terminal device 626 and the user via the display unit 627, etc., is performed by the interface 740 having input and output devices, presenting information to the user via the output device and accepting user operations via the input device under the control of the CPU 710.

[0144] One or more of the above-mentioned programs may be recorded on the non-volatile recording medium 750. In this case, the interface 740 may read the program from the non-volatile recording medium 750. The CPU 710 may then either directly execute the program read by the interface 740, or temporarily save it in the main memory 720 or auxiliary memory 730 before executing it.

[0145] Alternatively, the processing of each component may be performed by recording a program for executing all or part of the processing performed by the first terminal device 100, the first terminal device 102, the second terminal device 200, the analysis server device 300, the analysis server device 302, the impersonation detection device 610, the first terminal device 621, the impersonation detection device 624, and the second terminal device 626 on a computer-readable recording medium, and then loading and executing the program recorded on this recording medium into a computer system. The term "computer system" here includes hardware such as an OS (Operating System) and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, magneto-optical disks, ROMs (Read Only Memory), CD-ROMs (Compact Disc Read Only Memory), and storage devices such as hard disks built into computer systems. The above-mentioned program may be intended to implement only a part of the functions described above, and may also be able to implement the above-mentioned functions in combination with programs already recorded in the computer system.

[0146] Although embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments and includes designs and the like that do not depart from the spirit of this invention. Furthermore, the above-described embodiments may be combined with other embodiments as appropriate.

[0147] Some or all of the above embodiments may also be described as follows, but are not limited to these.

[0148] (Note 1) A voice data analysis means that determines whether or not impersonation using synthesized speech is occurring, based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data. A device for detecting impersonation.

[0149] (Note 2) A biometric information analysis means that determines whether or not impersonation using synthesized speech has occurred, based on whether or not the value of the biometric information falls within a predetermined range for that type of biometric information. The impersonation detection device described in Appendix 1 is further equipped with the following:

[0150] (Note 3) An analysis processing means that determines whether or not impersonation using speech synthesis is occurring, based on the attributes of the person being judged, which are estimated based on biometric information. An impersonation detection device as described in Appendix 1 or Appendix 2, further comprising the above.

[0151] (Note 4) The aforementioned voice data analysis means determines whether or not the air conduction sound data and bone conduction sound data to be determined are data obtained from the speech of the same person, and determines whether or not there is impersonation using synthesized speech based on a comparison of the air conduction sound data and bone conduction sound data to be determined with stored air conduction sound data and bone conduction data. A device for detecting impersonation, as described in any one of the appendices 1 to 3.

[0152] (Note 5) The system comprises a first terminal device, a spoofing detection device, and a second terminal device. The first terminal device, A means for acquiring air conduction sound data to acquire air conduction sound data to be judged, A means for acquiring bone conduction sound data to acquire bone conduction sound data to be judged, Equipped with, The aforementioned impersonation detection device is A voice data analysis means that determines whether or not impersonation using synthesized speech is occurring, based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data. Equipped with, The aforementioned second terminal device is Display means for displaying the result of determining whether or not impersonation using synthesized speech has occurred. Equipped with, A system for detecting impersonation.

[0153] (Note 6) The aforementioned impersonation detection device is A biometric information analysis means that determines whether or not impersonation using synthesized speech has occurred, based on whether or not the value of the biometric information falls within a predetermined range for that type of biometric information. The impersonation detection system described in Appendix 5 is further equipped with the following features.

[0154] (Note 7) The aforementioned impersonation detection device is An analysis processing means that determines whether or not impersonation using speech synthesis is occurring, based on the attributes of the person being judged, which are estimated based on biometric information. An impersonation detection system further comprising the features described in Appendix 5 or Appendix 6.

[0155] (Note 8) The aforementioned voice data analysis means determines whether or not the air conduction sound data and bone conduction sound data to be determined are data obtained from the speech of the same person, and determines whether or not there is impersonation using synthesized speech based on a comparison of the air conduction sound data and bone conduction sound data to be determined with stored air conduction sound data and bone conduction data. The impersonation detection system described in any one of the appendices 5 to 7.

[0156] (Note 9) Computers The system determines whether or not impersonation using synthesized speech is occurring based on a comparison between the air conduction sound data and bone conduction sound data being evaluated and the stored air conduction sound data and bone conduction sound data. A method for detecting impersonation, including the following.

[0157] (Note 10) The aforementioned computer, The determination of whether or not impersonation using synthesized speech has occurred is based on whether or not the value of the biometric information falls within a predetermined range for that type of biometric information. The impersonation detection method described in Appendix 9, which further includes the method described in Appendix 9.

[0158] (Note 11) Determining whether or not someone is impersonating another person using speech synthesis, based on the attributes of the person being judged estimated from their biometric information. The impersonation detection method described in Appendix 9 or Appendix 10, further including the above.

[0159] (Note 12) Determining whether or not impersonation using synthesized speech has occurred based on a comparison of the air conduction sound data and bone conduction sound data to be determined with the stored air conduction sound data and bone conduction sound data includes the computer determining whether or not the air conduction sound data and bone conduction sound data to be determined are data obtained from the speech of the same person, and determining whether or not impersonation using synthesized speech has occurred based on a comparison of the air conduction sound data and bone conduction data to be determined with the stored air conduction sound data and bone conduction data. The impersonation detection method described in any one of the appendices 9 to 11.

[0160] (Note 13) On the computer, The system determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data being evaluated and the stored air conduction sound data and bone conduction sound data. A program that executes the command.

[0161] (Note 14) To the aforementioned computer, The determination of whether or not impersonation using synthesized speech has occurred is based on whether or not the value of the biometric information falls within a predetermined range for that type of biometric information. The program described in Appendix 13 further executes the following.

[0162] (Note 15) To the aforementioned computer, Determining whether or not someone is impersonating another person using speech synthesis, based on the attributes of the person being judged estimated from their biometric information. The program described in Appendix 13 or Appendix 14, which further executes the above.

[0163] (Note 16) In determining whether or not impersonation using synthesized speech has occurred based on a comparison of the air conduction sound data and bone conduction sound data to be determined with the stored air conduction sound data and bone conduction sound data, the computer is instructed to determine whether or not the air conduction sound data and bone conduction sound data to be determined are data obtained from the speech of the same person, and to determine whether or not impersonation using synthesized speech has occurred based on a comparison of the air conduction sound data and bone conduction data to be determined with the stored air conduction sound data and bone conduction data. The program described in any one of the appendices 13 to 15. [Explanation of symbols]

[0164] 1, 2, 620 Impersonation Detection System 100, 102, 621 First Terminal Device 110 First terminal side communication unit 120, 622 Air conduction sound data acquisition unit 130, 623 Bone conduction sound data acquisition unit 140 Biological Information Acquisition Unit 200, 626 Second Terminal Device 210 Second Terminal Communication Unit 220 Audio output section 230, 627 display section 300, 302 Analysis Server Equipment 310 Server-side communication unit 320, 611, 625 Voice Data Analysis Department 330 Audio data storage unit 340, 360 Analytical Processing Unit 350 Department of Biological Information Analysis 610, 624 Impersonation detection device

Claims

1. A voice data analysis means that determines whether or not impersonation using synthesized speech is occurring, based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data. A device for detecting impersonation.

2. A biometric information analysis means that determines whether or not impersonation using synthesized speech has occurred, based on whether or not the value of the biometric information falls within a predetermined range for that type of biometric information. The impersonation detection device according to claim 1, further comprising:

3. An analysis processing means that determines whether or not impersonation using speech synthesis is occurring, based on the attributes of the person being judged, which are estimated based on biometric information. The impersonation detection device according to claim 1, further comprising:

4. The aforementioned voice data analysis means determines whether or not the air conduction sound data and bone conduction sound data to be determined are data obtained from the speech of the same person, and determines whether or not there is impersonation using synthesized speech based on a comparison of the air conduction sound data and bone conduction sound data to be determined with stored air conduction sound data and bone conduction data. The impersonation detection device according to claim 1.

5. The system comprises a first terminal device, a spoofing detection device, and a second terminal device. The first terminal device is A means for acquiring air conduction sound data to acquire air conduction sound data to be judged, A means for acquiring bone conduction sound data to acquire bone conduction sound data to be judged, Equipped with, The aforementioned impersonation detection device is A voice data analysis means that determines whether or not impersonation using synthesized speech is occurring, based on a comparison between the air conduction sound data and bone conduction sound data to be judged and the stored air conduction sound data and bone conduction sound data. Equipped with, The second terminal device is Display means for displaying the result of determining whether or not impersonation using synthesized speech has occurred. Equipped with, A system for detecting impersonation.

6. Computers The system determines whether or not impersonation using synthesized speech is occurring based on a comparison between the air conduction sound data and bone conduction sound data being evaluated and the stored air conduction sound data and bone conduction sound data. A method for detecting impersonation, including the following.

7. On the computer, The system determines whether or not impersonation using synthesized speech has occurred based on a comparison between the air conduction sound data and bone conduction sound data being evaluated and the stored air conduction sound data and bone conduction sound data. A program that executes the command.

Citation Information

Patent Citations

  • Personal identification system

    JP2006010809A