Information processing device and program

The information processing device and program address the challenge of impersonation in voice conversations by generating indistinguishable fake voices and verifying identity, effectively preventing fraud.

JP2025145929APending Publication Date: 2025-10-03TOSHIBA TEC KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024046446
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing methods to verify identity during voice conversations are ineffective against impersonation due to the difficulty in distinguishing synthesized voices from real ones, especially in web conferences, and are limited in effectiveness for unaware users.

Method used

An information processing device and program that includes an acquisition unit to collect voice data, a voice processing unit to generate fake voice data, and a communication unit to transmit this data, altering pronunciation and intonation to create indistinguishable fake voices, and a determination unit to verify the authenticity of voice data.

Benefits of technology

Effectively prevents impersonation by generating fake voice data that is difficult to distinguish and verifies the identity of conversation partners, thereby safeguarding against voice fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025145929000001_ABST
    Figure 2025145929000001_ABST
Patent Text Reader

Abstract

To provide a technique for dealing with impersonation of another person in a conversation using a conversation application.SOLUTION: In one embodiment, an information processing device for a user to have a conversation with a conversation partner using a conversation application includes an acquisition unit, a voice processing unit, and a communication processing unit. The acquisition unit acquires voice data after passing through the conversation application based on the user's voice. The voice processing unit processes the voice data acquired by the acquisition unit. The communication processing unit transmits voice data generated based on the processing of the voice data by the voice processing unit to the conversation partner.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an information processing device and a program. [Background technology]

[0002] Recently, in addition to voice conversations using mobile phones, there has been an increase in web conferences that allow people far away to exchange video and audio using the Internet, making it easier to obtain the speaker's voice.

[0003] Meanwhile, advances in AI (Artificial Intelligence) technology have made it possible to inexpensively generate high-quality synthesized voices from anyone's voice by collecting their own voice data, allowing them to speak or sing as desired. This has made it easier for fraudsters to exploit these technologies. However, these synthesized voices contain less information than videos, making it difficult to spot differences or inconsistencies from the real thing, making it difficult to distinguish them as fakes.

[0004] In the past, methods to verify identity of fraudsters pretending to be the person in question during phone calls or web conferences have been to call back to verify the identity or to ask questions about information that only the person in question would know. However, both of these methods are only effective for a small number of people who are aware of the dangers involved, and it is expected that fraudsters pretending to be the person in question will become even more prevalent in the future. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-204368 Summary of the Invention [Problem to be solved by the invention]

[0006] The problem to be solved by the embodiments of the present invention is to provide a technique for dealing with impersonation of another person in a conversation using a conversation app. [Means for solving the problem]

[0007] In one embodiment, an information processing device for a user to have a conversation with a conversation partner using a conversation app includes an acquisition unit, a voice processing unit, and a communication processing unit. The acquisition unit acquires voice data after passing through the conversation app based on the user's voice. The voice processing unit processes the voice data acquired by the acquisition unit. The communication processing unit transmits voice data generated based on the processing of the voice data by the voice processing unit to the conversation partner. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating a conversation processing system according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating a first terminal according to the embodiment. [Figure 3] FIG. 3 is a block diagram illustrating a second terminal according to the embodiment. [Figure 4] FIG. 4 is a flowchart illustrating an example of conversation processing by the processing circuit of the first terminal according to the embodiment. [Figure 5] FIG. 5 is a flowchart illustrating an example of a process for generating fake voice data by the processing circuit of the first terminal according to the embodiment. [Figure 6] FIG. 6 is a flowchart illustrating an example of conversation processing by the processing circuit of the second terminal according to the embodiment. [Figure 7] FIG. 7 is a diagram for explaining an example in which the second terminal according to the embodiment determines that the conversation partner is not the specific user himself / herself. [Figure 8] FIG. 8 is a diagram for explaining an example in which the second terminal according to the embodiment determines that the conversation partner is the specific user himself / herself. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, several embodiments will be described with reference to the drawings. Note that the scale of each part in each drawing used in the following description of the embodiments may be changed as appropriate. Also, for the sake of explanation, each drawing used in the following description of the embodiments may omit components.

[0010] [Embodiment] (Configuration example) FIG. 1 is a block diagram illustrating a conversation processing system S. The conversation processing system S is a system for realizing conversation via a device. The conversation may be a conversation without accompanying video or a conversation with accompanying video. The conversation includes the meaning of a voice call.

[0011] Here, the following description will be given taking as an example a conversation realized via a network NW without using a telephone number, such as a social networking service (SNS) or a web conference, but is not limited to this. The conversation may also be realized via a telephone line using a telephone number.

[0012] The conversation processing system S includes a first terminal 1, a second terminal 2, a third terminal 3, and a server 4. The first terminal 1, the second terminal 2, the third terminal 3, and the server 4 are communicably connected to each other via a network NW. The network NW includes one or more of various networks such as the Internet, a mobile communication network, and a LAN (Local Area Network). The LAN may be a wireless LAN or a wired LAN.

[0013] The first terminal 1, the second terminal 2, and the third terminal 3 are devices that allow a user to converse with a conversation partner using a conversation application described below. The conversation partner may be a person or a computer that generates voice data, such as AI. If the conversation partner is a person, the conversation partner also refers to another user. If the conversation partner is a computer, the conversation partner also refers to another user who operates the computer. For example, the first terminal 1, the second terminal 2, and the third terminal 3 may be, but are not limited to, a smartphone, a tablet terminal, or a PC (Personal Computer). Each of the first terminal 1, the second terminal 2, and the third terminal 3 is an example of an information processing device.

[0014] The first terminal 1 is assumed to be a device used by user A. The second terminal 2 is assumed to be a device used by user B. The third terminal 3 is assumed to be a device used by user C. It is assumed that users A and B are close friends or have important conversations with each other. For example, user C is an attacker attempting to commit voice fraud. Here, we will explain an example in which user C impersonates user A and uses the third terminal 3 to converse with user B.

[0015] The server 4 is a device that relays voice data in order to realize a conversation via the device.

[0016] FIG. 2 is a block diagram illustrating the first terminal 1. As shown in FIG. The first terminal 1 includes a processing circuit 10, a main memory 11, an auxiliary storage device 12, a communication interface 13, an input / output I / F 14, a display device 15, an audio output device 16, an audio input device 17, and an imaging device 18. The processing circuit 10, the main memory 11, the auxiliary storage device 12, the communication interface 13, the input / output I / F 14, the display device 15, the audio output device 16, the audio input device 17, and the imaging device 18 are connected to each other so that signals can be input and output. In Fig. 2, the interface is written as "I / F".

[0017] The processing circuit 10 corresponds to the central part of the first terminal 1. The processing circuit 10 is an element that constitutes the computer of the first terminal 1. The processing circuit 10 includes one or more circuits that perform multiple processes based on multiple functions. For example, the circuit is, but is not limited to, a processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field-Programmable Gate Array). For example, the processor is, but is not limited to, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The processing circuit 10 loads a program stored in the main memory 11 or the auxiliary storage device 12 into the main memory 11. The processing circuit 10 executes the program loaded into the main memory 11, thereby enabling various processes to be performed.

[0018] The main memory 11 includes elements corresponding to the main memory portion of the first terminal 1. The main memory 11 is an element that constitutes the computer of the first terminal 1. The main memory 11 includes a non-volatile memory area and a volatile memory area. The main memory 11 stores an operating system or a program in the non-volatile memory area. The main memory 11 uses the volatile memory area as a work area where data is rewritten as appropriate by the processing circuit 10. For example, the main memory 11 includes ROM (Read Only Memory) as a non-volatile memory area. For example, the main memory 11 includes RAM (Random Access Memory) as a volatile memory area. The main memory 11 is an example of a storage unit of the first terminal 1.

[0019] The auxiliary storage device 12 corresponds to the auxiliary storage portion of the first terminal 1. The auxiliary storage device 12 includes one or more storage devices. The storage devices may be, but are not limited to, an EEPROM (registered trademark) (Electric Erasable Programmable Read-Only Memory), a HDD (Hard Disc Drive), an SSD (Solid State Drive), or a semiconductor memory. The auxiliary storage device 12 stores the above-mentioned programs, data used by the processing circuit 10 to perform various processes, and data generated by the processing in the processing circuit 10. The auxiliary storage device 12 is an example of a storage portion of the first terminal 1.

[0020] The auxiliary storage device 12 includes a conversation application storage area 121. The conversation application storage area 121 stores a conversation application program. Hereinafter, the conversation application may be abbreviated as a conversation app. The conversation app is an application that can execute processing for realizing a conversation via a device. Here, the conversation app is described as an application for realizing a conversation via a network NW without using a phone number, such as an SNS or web conference, but is not limited to this. The conversation app may also be an application for realizing a conversation via a telephone line using a phone number.

[0021] The auxiliary storage device 12 includes a voice fraud prevention app storage area 122. The voice fraud prevention app storage area 122 stores a voice fraud prevention application program. Hereinafter, the voice fraud prevention application may be abbreviated as a voice fraud prevention app. The voice fraud prevention app is an application that can execute processing to prevent voice fraud. The voice fraud prevention app may be part of a conversation app.

[0022] The communication interface 13 includes various interfaces that connect the first terminal 1 to other devices via the network NW in accordance with a predetermined communication protocol so that the first terminal 1 can communicate with other devices. The communication interface 13 is an example of a communication unit of the first terminal 1.

[0023] The input / output interface 14 includes various interfaces that connect external devices to the first terminal 1 so that signals can be input and output. For example, external devices include, but are not limited to, a display, headphones, a microphone, and a camera. The input / output interface 14 is an example of a connection unit of the first terminal 1.

[0024] The display device 15 is a device capable of displaying various images under the control of the processing circuit 10. For example, the display device 15 is a liquid crystal display or an EL (Electroluminescence) display. The display device 15 is an example of a display unit of the first terminal 1.

[0025] The audio output device 16 is a device that can output audio based on audio data. For example, the audio output device 16 is a speaker. The audio output device 16 is an example of an audio output unit of the first terminal 1.

[0026] The voice input device 17 is a device that can input voice to the first terminal 1. For example, the voice input device 17 is a microphone. The voice input device 17 is an example of a voice input unit of the first terminal 1.

[0027] The imaging device 18 is a device that captures images by taking pictures. For example, the imaging device 18 is a camera. The imaging device 18 is an example of an imaging unit of the first terminal 1.

[0028] The input device 19 is a device that can input instructions or information to the first terminal 1. The input device 19 may form a touch screen together with the display device 15. The input device 19 may include a keyboard. The input device 19 is an example of an input unit of the first terminal 1.

[0029] The hardware configuration of the first terminal 1 is not limited to the above configuration. The first terminal 1 may appropriately omit or change the above components and add new components.

[0030] Each unit realized by the processing circuit 10 will be described. The processing circuit 10 realizes an application processing unit 101, an acquisition unit 102, a voice processing unit 103, and a communication processing unit 104. The acquisition unit 102, the voice processing unit 103, and the communication processing unit 104 may be realized by the processing circuit 10 by executing a voice fraud prevention application. Each unit realized by the processing circuit 10 can also be referred to as a function. Each unit realized by the processing circuit 10 can also be referred to as being realized by a control unit including the processing circuit 10 and the main memory 11.

[0031] The application processing unit 101 processes an application. For example, the application processing unit 101 can start an application. Starting an application includes making it possible to execute processing related to the application. The application processing unit 101 can terminate a running application.

[0032] The acquisition unit 102 acquires voice data of user A. For example, the acquisition unit 102 can acquire voice data after passing through a conversation app based on the voice of the first user. The voice data after passing through the conversation app is voice data after being digitized by the conversation app based on the voice of user A input via the voice input device 17. Hereinafter, the voice data after passing through the conversation app acquired by the acquisition unit 102 is also referred to as input voice data.

[0033] The voice processing unit 103 processes the input voice data. The voice processing unit 103 generates voice data based on the processing of the input voice data. The word "processing" includes the meaning of "change." Hereinafter, the voice data generated based on the processing of the input voice data by the voice processing unit 103 is also referred to as fake voice data. The fake voice data is data in which voice is processed to a level that is difficult for a listener to understand, without processing character strings such as words indicated in the input voice data.

[0034] The fake voice data may include voice data obtained by processing the pronunciation of words indicated in the input voice data. The pronunciation is the way words are pronounced. For example, the pronunciation includes the strength, pitch, and length of the pronunciation. The pronunciation strength is the strength of the voice when a word is spoken. The pronunciation strength is the pitch of the voice when a word is spoken. The pronunciation length is the pause between words when a word is spoken.

[0035] In one example, the pronunciation manner is an accent. The pronunciation manner of the accent is predetermined for each word. Words that are pronounced "hashi" include "hashi," "hashi," and "hashi." The positions where stress is added differ among these. Words that are pronounced "fuku" include "fuku" and "fuku." The positions where stress is added differ among these.

[0036] In another example, the speech manner is intonation. The speech manner of intonation is attributed to the user. Intonation includes the meaning of prosody and rhythm. Some users do not pronounce any of the parts of "arigatou" strongly, some users pronounce "ri" strongly, and some users pronounce "to" strongly.

[0037] The fake audio data may include audio data obtained by synthesizing other audio data with input audio data. Hereinafter, the other audio data synthesized with input audio data is also referred to as additional audio data. The additional audio data is audio data that is different from the input audio data. The additional audio data is data of any sound. For example, the additional audio data is data of inaudible noise of 20 kHz or more, but is not limited to this. The additional audio data may also be data of audible sound. Synthesis includes the meaning of mixing.

[0038] The voice processing unit 103 can perform linguistic analysis of the input voice data to generate fake voice data. The voice processing unit 103 divides the input voice data into processing units by linguistic analysis of the input voice data. A processing unit is a collection of input voice data for generating fake voice data. A processing unit may be a sentence unit or a paragraph unit. The voice processing unit 103 generates fake voice data based on the input voice data for each processing unit. The voice processing unit 103 can process the input voice data as follows to generate fake voice data.

[0039] In one example, processing the input voice data includes processing the manner in which words represented by the input voice data are spoken.

[0040] A case where the speech mode is an accent will be described. The speech processing unit 103 refers to dictionary data and analyzes the pronunciation and accent of words indicated in the input speech data. The dictionary data is data in which the pronunciation and predetermined accent of each word are registered. The dictionary data may be stored in the auxiliary storage device 12.

[0041] The speech processing unit 103 processes the accent of words indicated by the input speech data. For example, the speech processing unit 103 changes the accent of a word indicated by the input speech data to an accent determined for another word that has the same pronunciation as this word. An example will be explained using "I bought chopsticks today." The speech processing unit 103 can change the accent of "chopsticks" from the accent determined for "chopsticks" to the accent determined for "bridge."

[0042] A case where the speech mode is intonation will be described. The speech processing unit 103 processes the intonation of words indicated by the input speech data. An example will be described using "Thank you." Here, the intonation of "Thank you" indicated by the input speech data is assumed to be soft. The speech processing unit 103 can change the intonation of "Thank you" from an intonation in which none of the sounds are strongly pronounced to an intonation in which the "ri" is strongly pronounced.

[0043] In another example, processing the input voice data includes synthesizing additional voice data with the input voice data. For example, the voice processing unit 103 can synthesize an inaudible sound with the word "thank you."

[0044] The communication processing unit 104 receives voice data of the conversation partner from the conversation partner via the network NW. Receiving from the conversation partner includes receiving from the conversation partner's terminal. The communication processing unit 104 transmits user A's voice data to the conversation partner via the network NW. Transmitting to the conversation partner includes transmitting to the conversation partner's terminal. For example, the communication processing unit 104 can transmit fake voice data of user A to the conversation partner.

[0045] FIG. 3 is a block diagram illustrating the second terminal 2. As shown in FIG. The second terminal 2 includes a processing circuit 20, a main memory 21, an auxiliary storage device 22, a communication interface 23, an input / output I / F 24, a display device 25, an audio output device 26, an audio input device 27, and an imaging device 28. The processing circuit 20, the main memory 21, the auxiliary storage device 22, the communication interface 23, the input / output I / F 24, the display device 25, the audio output device 26, the audio input device 27, and the imaging device 28 are connected to each other so that signals can be input and output. In Fig. 3, the interface is indicated as "I / F".

[0046] The processing circuit 20 corresponds to the central part of the second terminal 2. The processing circuit 20 is an element that constitutes the computer of the second terminal 2. The processing circuit 20 is configured similarly to the processing circuit 10. The processing circuit 20 loads a program stored in the main memory 21 or the auxiliary storage device 22 into the main memory 21. The processing circuit 20 executes the program loaded into the main memory 21, thereby enabling various processes to be executed.

[0047] The main memory 21 includes elements corresponding to the main storage portion of the second terminal 2. The main memory 21 is an element that constitutes the computer of the second terminal 2. The main memory 21 is configured in the same manner as the main memory 11. The main memory 21 is an example of a storage unit of the second terminal 2.

[0048] The auxiliary storage device 22 corresponds to the auxiliary storage portion of the second terminal 2. The auxiliary storage device 22 has the same configuration as the auxiliary storage device 12. The auxiliary storage device 22 is an example of a storage unit of the second terminal 2.

[0049] The auxiliary storage device 22 includes a conversation application storage area 221. The conversation application storage area 221 stores a conversation application program.

[0050] The auxiliary storage device 22 includes a voice fraud prevention application storage area 222. The voice fraud prevention application storage area 222 stores a program for a voice fraud prevention application.

[0051] The auxiliary storage device 22 includes a registered voice data storage area 223. The registered voice data storage area 223 stores the registered voice data of each user. Here, it is assumed that the registered voice data storage area 223 stores the registered voice data of user A. The registered voice data is raw voice data registered in the second terminal 2. The raw voice data is unprocessed voice data obtained based on the user's voice after passing through a conversation app. The raw voice data is similar to the input voice data. The registered voice data storage area 223 stores identification information that is linked to the registered voice data and can uniquely identify the user. For example, the identifiable information is the user's name, but is not limited to this. The identifiable information may be the user's account for the conversation app or the user's phone number.

[0052] For example, the second terminal 2 registers the raw voice data of each user as registration voice data as follows. Here, the description will be given using the raw voice data of user A as an example. When user A reads a predetermined sentence, the first terminal 1 generates raw voice data of user A. The first terminal 1 transmits the raw voice data of user A to the second terminal 2. The second terminal 2 receives the raw voice data of user A from the first terminal 1. The second terminal 2 saves the received raw voice data of user A in the registration voice data storage area 223 as registration voice data of user A. This allows the second terminal 2 to register the raw voice data of user A in the second terminal 2 as registration voice data of user A.

[0053] The communication interface 23 includes various interfaces that connect the second terminal 2 to other devices via the network NW in accordance with a predetermined communication protocol so that the second terminal 2 can communicate with other devices. The communication interface 23 is an example of a communication unit of the second terminal 2.

[0054] The input / output interface 24 includes various interfaces that connect external devices to the second terminal 2 so as to input and output signals. The input / output interface 24 is an example of a connection unit of the second terminal 2.

[0055] The display device 25 is a device that can display various images under the control of the processing circuit 20. The display device 25 has the same configuration as the display device 15. The display device 25 is an example of a display unit of the second terminal 2.

[0056] The audio output device 26 is a device capable of outputting audio based on audio data. The audio output device 26 has the same configuration as the audio output device 16. The audio output device 16 is an example of the audio output unit of the second terminal 2.

[0057] The voice input device 27 is a device that can input voice to the second terminal 2. The voice input device 27 has the same configuration as the voice input device 17. The voice input device 27 is an example of a voice input unit of the second terminal 2.

[0058] The imaging device 28 is a device that captures images by taking pictures. The imaging device 28 has the same configuration as the imaging device 18. The imaging device 28 is an example of the imaging unit of the second terminal 2.

[0059] The input device 29 is a device that can input instructions or information to the second terminal 2. The input device 29 has the same configuration as the input device 19. The input device 29 is an example of an input unit of the second terminal 2.

[0060] The hardware configuration of the second terminal 2 is not limited to the above configuration. The second terminal 2 may appropriately omit or change the above components and add new components.

[0061] Each unit realized by the processing circuit 20 will be described. The processing circuit 20 realizes an application processing unit 201, a selection unit 202, a communication processing unit 203, a determination unit 204, and a notification unit 205. The selection unit 202, the communication processing unit 203, the determination unit 204, and the notification unit 205 may be realized by the processing circuit 10 by executing a voice fraud prevention application. Each unit realized by the processing circuit 20 can also be referred to as a function. Each unit realized by the processing circuit 20 can also be referred to as being realized by a control unit including the processing circuit 20 and the main memory 21.

[0062] The application processing unit 201 processes an application. For example, the application processing unit 201 can start an application. The application processing unit 201 can terminate an application that is currently running.

[0063] The selection unit 202 selects the registered voice data of a specific user. The specific user is a person identified as a conversation partner of user B. The conversation partner may be the specific user, or may not be the specific user but a computer masquerading as the specific user.

[0064] The communication processing unit 203 receives voice data of the conversation partner from the conversation partner via the network NW. The communication processing unit 203 transmits voice data of user B to the conversation partner via the network NW.

[0065] The determination unit 204 determines whether the conversation partner is a specific user. For example, the determination unit 204 can determine whether the conversation partner is a specific user based on the voice data of the conversation partner and the raw voice data of the specific user.

[0066] The determination unit 204 compares the voice data of the conversation partner with the raw voice data of the specific user. The determination unit 204 may determine whether the conversation partner is a specific user based on a comparison of the pronunciation of the same words included in both the voice data of the conversation partner and the raw voice data. If the pronunciation of the words included in the voice data of the conversation partner is the same as the pronunciation of the same words included in the raw voice data, the determination unit 204 can determine that the conversation partner is a specific user. If the voice data of the conversation partner is generated based on fake voice data of the specific user, the pronunciation of the words included in the voice data of the conversation partner is different from the pronunciation of the same words included in the raw voice data. In this case, the determination unit 204 can determine that the conversation partner is not a specific user.

[0067] Note that the determination unit 204 is not limited to comparing the pronunciation patterns of the same words contained in both the conversation partner's voice data and the raw voice data. The determination unit 204 may compare the overall tendency of the pronunciation patterns in the conversation partner's voice data with the overall tendency of the pronunciation patterns in the enrollment voice data.

[0068] The determination unit 204 may determine whether the conversation partner is a specific user based on the presence or absence of additional voice data. If the voice data of the conversation partner does not include additional voice data, the determination unit 204 can determine that the conversation partner is a specific user. If the voice data of the conversation partner is generated based on fake voice data of the specific user, the voice data of the conversation partner includes additional voice data. If the voice data of the conversation partner includes additional voice data but the registered voice data of the specific user does not include this additional voice data, the determination unit 204 can determine that the conversation partner is not a specific user.

[0069] The notification unit 205 notifies the determination result by the determination unit 204. The determination result is information indicating whether the conversation partner is a specific user or not. The fact that the conversation partner is not a specific user implies that the conversation partner is impersonating a specific user. The notification unit 205 may notify the determination result by displaying the determination result on the display device 25. The notification unit 205 may notify the determination result by outputting the determination result as sound from the audio output device 26.

[0070] (Example of operation) The processing performed by the processing circuits of the devices included in the conversation processing system S will now be described. The processing procedures described below are merely examples, and each process may be modified as much as possible. Furthermore, steps may be omitted, replaced, or added as appropriate depending on the embodiment.

[0071] FIG. 4 is a flowchart showing an example of conversation processing by the processing circuit 10 of the first terminal 1. The conversation process in the first terminal 1 is a process of generating fake voice data in a conversation between user A and a conversation partner.

[0072] Here, it is assumed that User A is having a conversation with User C, who is not very familiar with User A or has met User C for the first time. User A uses a voice fraud prevention app to prevent User A's raw voice data from being passed to User C.

[0073] The processing circuit 10 detects a voice fraud prevention app launch operation by user A using the input device 19 (ACT1). ACT1 may be processing by the application processing unit 101. If the processing circuit 10 does not detect a voice fraud prevention app launch operation (ACT1, NO), the processing circuit 10 waits for a voice fraud prevention app launch operation. If the processing circuit 10 detects a voice fraud prevention app launch operation (ACT1, YES), the processing transitions from ACT1 to ACT2. The processing circuit 10 launches the voice fraud prevention app based on the voice fraud prevention app launch operation. By launching the voice fraud prevention app, the processing circuit 10 enables processing related to the voice fraud prevention app to be executed.

[0074] The processing circuit 10 detects a conversation app launch operation by user A using the input device 19 (ACT2). ACT2 may be processing by the application processing unit 101. If the processing circuit 10 does not detect a conversation app launch operation (ACT2, NO), the processing circuit 10 waits for a conversation app launch operation. If the processing circuit 10 detects a conversation app launch operation (ACT2, YES), the processing transitions from ACT2 to ACT3. The processing circuit 10 launches the conversation app based on the conversation app launch operation. By launching the conversation app, the processing circuit 10 realizes a conversation between user A using the first terminal 1 and a conversation partner using the third terminal 3.

[0075] The processing circuit 10 stores the input voice data of user A (ACT3). ACT3 may be processing by the acquisition unit 102. In ACT3, for example, the processing circuit 10 acquires the input voice data. The processing circuit 10 stores the acquired input voice data in the main memory 11 or the auxiliary storage device 12.

[0076] The processing circuit 10 generates fake voice data based on processing of the input voice data (ACT4). ACT4 may be processing by the voice processing unit 103.

[0077] The processing circuit 10 transmits the fake voice data to the conversation partner (ACT5). ACT5 may be processing by the communication processing unit 104. In ACT5, for example, the processing circuit 10 transmits the fake voice data to the third terminal 3 of user B so that the fake voice data is received by the third terminal 3 via the server 4.

[0078] The processing circuit 10 detects an operation by user A to terminate the conversation app using the input device 19 (ACT6). ACT6 may be processing by the application processing unit 101. If the processing circuit 10 does not detect an operation to terminate the conversation app (ACT6, NO), the processing transitions from ACT6 to ACT3. If the processing circuit 10 detects an operation to terminate the conversation app (ACT6, YES), the processing circuit 10 terminates the running conversation app. By terminating the running conversation app, the processing circuit 10 cuts off the conversation between user A using the first terminal 1 and the conversation partner using the third terminal 3. The processing circuit 10 terminates the voice fraud prevention application based on the termination of the conversation app.

[0079] As described above, the first terminal 1 can transmit fake voice data of user A to the conversation partner during the conversation. This allows the first terminal 1 to prevent the conversation partner from passing on the raw voice data of user A to the conversation partner. Therefore, a person other than user A cannot impersonate user A by using the raw voice data of user A. In this way, the first terminal 1 can deal with impersonation of another person in a conversation using a conversation app.

[0080] An example of the process for generating fake voice data in ACT4 will be described. FIG. 5 is a flowchart showing an example of a process for generating fake voice data by the processing circuit 10 of the first terminal 1.

[0081] The processing circuit 10 performs linguistic analysis of the input speech data (ACT41). In ACT41, for example, the processing circuit 10 divides the input speech data into processing units by linguistic analysis of the input speech data. The processing circuit 10 processes the input speech data for each processing unit as follows.

[0082] The processing circuit 10 determines whether to process the accent (ACT42). The processing circuit 10 may process the accent of all words indicated in the input speech data, or may process the accent of arbitrarily selected words. The target of accent processing by the processing circuit 10 can be set as appropriate.

[0083] If the processing circuit 10 determines that the accent should be processed (ACT42, YES), the process proceeds from ACT42 to ACT43. The processing circuit 10 processes the accent of the word indicated by the input voice data (ACT43). If the processing circuit 10 determines that the accent should not be processed (ACT42, NO), the process proceeds from ACT42 to ACT44.

[0084] The processing circuit 10 determines whether to process the intonation (ACT44). The processing circuit 10 may process the intonation of all words represented by the input speech data, or may process the intonation of arbitrarily selected words. The target of intonation processing by the processing circuit 10 can be set as appropriate.

[0085] If the processing circuit 10 determines that the intonation should be processed (ACT44, YES), the process proceeds from ACT44 to ACT45. The processing circuit 10 processes the intonation of the words represented by the input voice data (ACT45). If the processing circuit 10 determines that the intonation should not be processed (ACT44, NO), the process proceeds from ACT44 to ACT46.

[0086] The processing circuit 10 processes the sentences (ACT 46). In ACT 46, for example, the processing circuit 10 processes the processed sentences based on the input voice data so as to smoothly connect the sentences.

[0087] The processing circuit 10 determines whether to synthesize additional voice data with the input voice data (ACT47). The processing circuit 10 may synthesize additional voice data with all words indicated in the input voice data, or may synthesize additional voice data with arbitrarily selected words. The target of accent processing by the processing circuit 10 can be set as appropriate.

[0088] If the processing circuit 10 determines that the additional audio data is to be synthesized with the input audio data (ACT47, YES), the process proceeds from ACT47 to ACT48. The processing circuit 10 synthesizes the additional audio data with the input audio data (ACT48). If the processing circuit 10 determines that the additional audio data is not to be synthesized with the input audio data (ACT47, YES), the process ends.

[0089] The processing circuit 10 can generate fake voice data based on the processing of input voice data as described above.

[0090] FIG. 6 is a flowchart showing an example of conversation processing by the processing circuit 20 of the second terminal 2. The conversation process in the second terminal 2 is a process for determining whether or not the conversation partner is a specific user in a conversation between user B and the conversation partner.

[0091] Here, the specific user is assumed to be User A. User B is assumed to be conversing with someone claiming to be User A. User B's conversation partner may be User A, or it may not be User A but a computer pretending to be User A. User B uses a voice fraud prevention app to find out whether the conversation partner is User A.

[0092] The processing circuit 20 detects a voice fraud prevention app launch operation by user B using the input device 29 (ACT11). ACT11 may be processing by the application processing unit 201. If the processing circuit 20 does not detect a voice fraud prevention app launch operation (ACT11, NO), the processing circuit 20 waits for a voice fraud prevention app launch operation. If the processing circuit 20 detects a voice fraud prevention app launch operation (ACT11, YES), the processing transitions from ACT11 to ACT12. The processing circuit 20 launches the voice fraud prevention app based on the voice fraud prevention app launch operation. By launching the voice fraud prevention app, the processing circuit 20 enables processing related to the voice fraud prevention app to be executed.

[0093] The processing circuit 20 detects a conversation app launch operation by user A using the input device 29 (ACT12). ACT12 may be processed by the application processing unit 201. If the processing circuit 20 does not detect a conversation app launch operation (ACT12, NO), the processing circuit 20 waits for a conversation app launch operation. If the processing circuit 20 detects a conversation app launch operation (ACT12, YES), the processing transitions from ACT12 to ACT13. The processing circuit 20 launches the conversation app based on the conversation app launch operation. By launching the conversation app, the processing circuit 20 realizes a conversation between user B who uses the second terminal 2 and the conversation partner.

[0094] The processing circuit 20 selects registered voice data (ACT13). ACT13 may be processing by the selection unit 202. In ACT13, for example, the processing circuit 20 selects registered voice data of a specific user from the registered voice data of each user stored in the registered voice data storage area 223. The processing circuit 20 may select the registered voice data of a specific user based on the name of the specific user given by the conversation partner.

[0095] The processing circuit 20 stores the received voice data of the conversation partner (ACT14). ACT14 may be processing by the communication processing unit 203. In ACT14, for example, the processing circuit 20 receives the voice data of the conversation partner via the communication interface 23. The processing circuit 20 stores the received voice data of the conversation partner in the main memory 21 or the auxiliary storage device 22.

[0096] The processing circuit 20 determines whether the conversation partner is the specific user (ACT15). ACT15 may be processing by the determination unit 204. In ACT15, for example, the processing circuit 20 determines whether the conversation partner is the specific user based on the voice data of the conversation partner and the registered voice data of the specific user. If the voice data of the conversation partner is not generated based on the fake voice data of the specific user, the processing circuit 20 can determine that the conversation partner is the specific user. If the voice data of the conversation partner is generated based on the fake voice data of the specific user, the processing circuit 20 can determine that the conversation partner is not the specific user.

[0097] If the processing circuit 20 determines that the conversation partner is the specific user himself / herself (ACT16, YES), the processing transitions from ACT16 to ACT 17. If the processing circuit 20 determines that the conversation partner is not the specific user himself / herself (ACT16, NO), the processing transitions from ACT16 to ACT18.

[0098] The processing circuit 20 executes processing based on the fact that the conversation partner is the specific user himself / herself (ACT17). ACT17 may be processing by the notification unit 205. In ACT17, for example, the processing circuit 20 notifies information that the conversation partner is the specific user as a determination result. This allows user B to easily know that the conversation partner is the specific user himself / herself. User B can continue the conversation with the conversation partner after knowing that the conversation partner is the specific user himself / herself.

[0099] The processing circuit 20 executes processing based on the fact that the conversation partner is not the specific user (ACT18). ACT18 may be processing by the notification unit 205. In ACT18, for example, the processing circuit 20 notifies information that the conversation partner is not the specific user as a determination result. This allows user B to easily know that the conversation partner is not the specific user. After knowing that the conversation partner is not the specific user, user B can stop the conversation with the conversation partner.

[0100] If the processing circuit 20 determines that the conversation partner is not the specific user, the processing circuit 20 may terminate the running conversation app. This process may be performed by the application processing unit 201. This process may be performed after or before the notification of the determination result. By terminating the running conversation app, the processing circuit 20 interrupts the conversation between user B using the second terminal 2 and the conversation partner. This allows user B to end the conversation with the conversation partner without performing an operation to terminate the conversation app using the input device 29.

[0101] As described above, the second terminal 2 can determine whether the conversation partner is the specific user himself or herself. This allows the second terminal 2 to detect deepfake audio using fake audio data and prevent audio fraud. In this way, the second terminal 2 can deal with impersonation of other people in conversations using chat apps.

[0102] The flow up to when the second terminal 2 determines that the conversation partner is not the specific user will be described. FIG. 7 is a diagram for explaining an example in which the second terminal 2 determines that the conversation partner is not the specific user himself / herself.

[0103] First, user A uses first terminal 1 to transmit user A's raw voice data to second terminal 2 used by user B. User B uses second terminal 2 to register the received raw voice data of user A in second terminal 2 as user A's registered voice data.

[0104] Next, User A uses the voice fraud prevention app on the first terminal 1 to have a conversation with User C. During the conversation with User A, User C uses the third terminal 3 to obtain fake voice data of User A.

[0105] Next, User C impersonates User A and uses a third terminal 3 to converse with User B using voice data based on User A's fake voice data. User B uses a voice fraud prevention app on a second terminal 2 to determine whether the person he is conversing with is User A. The second terminal 2 can determine that the person he is conversing with is not User A.

[0106] The flow up to when the second terminal 2 determines that the conversation partner is the specific user will be described. FIG. 8 is a diagram for explaining an example in which the second terminal 2 determines that the conversation partner is the specific user himself / herself.

[0107] First, user A uses first terminal 1 to transmit user A's raw voice data to second terminal 2 used by user B. User B uses second terminal 2 to register the received raw voice data of user A in second terminal 2 as user A's registered voice data.

[0108] Next, user A converses with user B on first terminal 1 without using the voice fraud prevention app. During the conversation with user B, first terminal 1 transmits unprocessed input voice data after passing through the conversation app to second terminal 2. User B uses the voice fraud prevention app on second terminal 2 to determine whether the conversation partner is user A. Second terminal 2 can determine that the conversation partner is user A.

[0109] [Other embodiments] In the above-described embodiment, for convenience, the processing realized by executing the voice fraud prevention app has been described as being different between the first terminal 1 and the second terminal 2, but it is not different for each terminal. By executing the voice fraud prevention app, the processing circuit 10 of the first terminal 1 can realize the selection unit 202, determination unit 204, and notification unit 205 described for the second terminal 2 in addition to the acquisition unit 102, voice processing unit 103, and communication processing unit 104. By executing the voice fraud prevention app, the processing circuit 10 of the first terminal 1 can realize the acquisition unit 102 and voice processing unit 103 described for the first terminal 1 in addition to the selection unit 202, communication processing unit 203, determination unit 204, and notification unit 205. Note that when the terminal executes the function of the determination unit, the function of the voice processing unit may be switched on or off based on a user operation.

[0110] The above-described embodiments may be applied to a method executed by an apparatus, a program that can cause a computer of the apparatus to execute each function, or a recording medium that stores the program.

[0111] Each of the one or more circuits constituting the processing circuit performs one or more of the multiple processes. When the processing circuit is composed of a single circuit, the single circuit performs all of the multiple processes. When the processing circuit is composed of multiple circuits, each of the multiple circuits performs a part of the multiple processes. The part of the multiple processes may be one of the multiple processes, or two or more of the multiple processes. When the processing circuit is composed of multiple circuits, the multiple circuits may be included in a single device or may be distributed across multiple devices.

[0112] The program may be transferred in a state where it is stored in the device according to the embodiment, or may be transferred in a state where it is not stored in the device. In the latter case, the program may be transferred via a network, or may be transferred in a state where it is recorded on a recording medium. The recording medium is a non-transitory tangible medium. The recording medium is a computer-readable medium. The recording medium may be in any form, such as a CD-ROM or a memory card, as long as it is capable of storing the program and is computer-readable.

[0113] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims.

[0114] Some of the above-described embodiments may be expressed as follows. [C1] An information processing device for a user to have a conversation with a conversation partner using a conversation app, an acquisition unit that acquires voice data after passing through the conversation application based on the user's voice; a voice processing unit that processes the voice data acquired by the acquisition unit; a communication processing unit that transmits voice data generated based on the processing of the voice data by the voice processing unit to the conversation partner; An information processing device comprising: [C2] The information processing device according to [C1], wherein processing the voice data acquired by the acquisition unit includes processing the pronunciation of words indicated in the voice data acquired by the acquisition unit. [C3] The information processing device according to [C1], wherein processing the voice data acquired by the acquisition unit includes synthesizing other voice data with the voice data acquired by the acquisition unit. [C4] An information processing device for having a conversation with a conversation partner using a conversation app, a storage unit that stores unprocessed voice data of a specific user after passing through a conversation app, the unprocessed voice data being acquired based on the voice of the specific user; a communication processing unit for receiving voice data of the conversation partner; a determination unit that determines whether the conversation partner is the specific user based on the voice data of the conversation partner and unprocessed voice data of the specific user; An information processing device comprising: [C5] The information processing device according to [C4], further comprising a notification unit that notifies a result of the determination by the determination unit. [C6] A computer of an information processing device for a user to have a conversation with a conversation partner using a conversation app, acquiring voice data after passing through the conversation app based on the user's voice; processing the acquired audio data; transmitting voice data generated based on the processing of the voice data to the conversation partner; A program that can be executed. [Explanation of symbols]

[0115] 1...first terminal, 2...second terminal, 3...third terminal, 4...server, 10...processing circuit, 11...main memory, 12...auxiliary storage device, 13...communication interface, 14...input / output interface, 15...display device, 16...audio output device, 17...audio input device, 18...imaging device, 19...input device, 20...processing circuit, 21...main memory, 22...auxiliary storage device, 23...communication interface, 24...input / output interface, 25...display device, 26...audio output device , 27...voice input device, 28...imaging device, 29...input device, 101...application processing unit, 102...acquisition unit, 103...audio processing unit, 104...communication processing unit, 121...conversation application storage area, 122...voice fraud prevention application storage area, 201...application processing unit, 202...selection unit, 203...communication processing unit, 204...determination unit, 205...notification unit, 221...conversation application storage area, 222...voice fraud prevention application storage area, 223...registered voice data storage area, NW...network, S...conversation processing system.

Claims

1. An information processing device for a user to have a conversation with a conversation partner using a conversation app, an acquisition unit that acquires voice data after passing through the conversation application based on the user's voice; a voice processing unit that processes the voice data acquired by the acquisition unit; a communication processing unit that transmits voice data generated based on the processing of the voice data by the voice processing unit to the conversation partner; An information processing device comprising:

2. The information processing device according to claim 1 , wherein processing the voice data acquired by the acquisition unit includes processing a manner of speaking of words indicated by the voice data acquired by the acquisition unit.

3. The information processing device according to claim 1 , wherein processing the voice data acquired by the acquisition unit includes synthesizing the voice data acquired by the acquisition unit with other voice data.

4. An information processing device for having a conversation with a conversation partner using a conversation app, a storage unit that stores unprocessed voice data of a specific user after passing through a conversation app, the unprocessed voice data being acquired based on the voice of the specific user; a communication processing unit for receiving voice data of the conversation partner; a determination unit that determines whether the conversation partner is the specific user based on the voice data of the conversation partner and unprocessed voice data of the specific user; An information processing device comprising:

5. The information processing device according to claim 4 , further comprising a notification unit that notifies a result of the determination made by the determination unit.

6. A computer of an information processing device for a user to have a conversation with a conversation partner using a conversation app, acquiring voice data after passing through the conversation app based on the user's voice; processing the acquired audio data; transmitting voice data generated based on the processing of the voice data to the conversation partner; A program that can be executed.

Citation Information

Patent Citations

  • Voice authentication apparatus, voice authentication method, and program

    JP2019204368A