Determination method, determination program, and information processing apparatus
By extracting and comparing feature information from low-frequency behaviors, the method enhances the detection accuracy of impersonation in remote conversations, addressing the limitations of conventional deepfake determination methods.
Patent Information
- Application Number
- JP2023573698
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-12
AI Technical Summary
Conventional deepfake determination methods struggle to accurately detect impersonation in remote conversations, especially when there is a large amount of training data, as they may not be able to confirm identity by simply comparing past and current behaviors.
The method extracts feature information from actions, voices, and states of a participant with a low extraction frequency and compares it with the feature information from current sensing data to determine impersonation based on the degree of coincidence.
This approach improves the detection accuracy of impersonation in remote conversations by focusing on behaviors with low frequencies, which are harder to replicate accurately.
Smart Images

Figure 0007697535000002 
Figure 0007697535000003 
Figure 0007697535000004
Abstract
Description
Technical Field
[0001] The present invention relates to a determination method, a determination program, and an information processing apparatus.
Background Art
[0002] In recent years, synthetic media using images and voices generated and edited using AI (Artificial Intelligence) has been developed and is expected to be used in various fields. On the other hand, synthetic media manipulated for improper purposes has become a social problem.
[0003] Synthetic media manipulated for improper purposes may be referred to as deepfakes. Also, a fake image generated by a deepfake may be referred to as a deepfake image, and a fake video generated by a deepfake may be referred to as a deepfake video.
[0004] With the technological evolution of AI and the improvement of computer resources, it has become technically possible to generate deepfake images and deepfake videos that do not actually exist, and fraud damages and the like caused by deepfake images and deepfake videos have occurred, becoming a social problem.
[0005] Moreover, there is a risk that the damage will become even greater due to the abuse of deepfake images and deepfake videos for forgery.
[0006] In order to detect deepfake videos using synthetic media, for example, at the time of a remote conversation via the Internet, a method is known in which the behavior in the past and the current time is compared, and if the behavior does not match, a warning is given that the participant is not the person himself / herself.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0008] However, in such a conventional deepfake determination method, there are cases where determination cannot be made only by comparing the past and current behaviors of the target person (participant).
[0009] For example, in an image generation model used for face conversion or a voice generation model used for voice conversion, generally, learning is performed so that the training data (= past behavior of the target person) matches the data to be generated.
[0010] Therefore, if there is a large amount of training data, an attacker can reproduce behaviors similar to those of the target person, and in particular, behaviors with high frequencies are easy to reproduce. Therefore, simply comparing the past and current behaviors may not be able to confirm identity.
[0011] In one aspect, the present invention makes it possible to improve the detection accuracy of impersonation in a remote conversation.
Means for Solving the Problems
[0012] For this reason, when this determination method receives first sensing data associated with the account of a participant in a remote conversation, it extracts from the past second sensing data of the participant, and the feature information of any one of the actions, voices, and states of the participant whose extraction frequency is less than a first reference value is obtained, and a determination regarding impersonation is made based on the degree of coincidence between the feature information extracted from the first sensing data and the feature information extracted from the second sensing data.
Effects of the Invention
[0013] According to one embodiment, the detection accuracy of impersonation in a remote conversation can be improved.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
[0015] Hereinafter, embodiments of the present determination method, determination program, and information processing apparatus will be described with reference to the drawings. However, the embodiments shown below are merely examples, and there is no intention of excluding various modifications and applications of technologies not explicitly shown in the embodiments. That is, the present embodiments can be implemented with various modifications (such as combining the embodiments and various modification examples) without departing from the gist thereof. Also, each figure is not intended to include only the components shown in the figure, but can include other functions and the like.
[0016] (I) Description of the First Embodiment (A) Configuration FIG. 1 schematically shows the hardware configuration of a computer system 1 as an example of the first embodiment, and FIG. 2 illustrates its functional configuration.
[0017] The computer system 1 illustrated in FIG. 1 includes an information processing device 10, a host terminal 3, and a plurality of participant terminals 3. These information processing device 10, host terminal 3, and plurality of participant terminals 3 are communicably connected to each other via a network 20.
[0018] The computer system 1 realizes a remote conversation via the network 20 among users of the plurality of participant terminals 3. In FIG. 1, for the sake of convenience, three participant terminals 2 and one host terminal 3 are shown, but it is not limited thereto. It may be provided with two or less or four or more participant terminals 2, or may be provided with a plurality of host terminals 3.
[0019] The remote conversation is carried out between two or more accounts among a plurality of accounts set to be able to participate in the remote conversation. Hereinafter, the participants in the remote conversation may be simply referred to as participants. The users of the participant terminals 2 all correspond to participants. Hereinafter, the user himself / herself of the participant terminal 2 may be referred to as a participant. The remote conversation may be, for example, an online meeting.
[0020] In this computer system 1, in a remote conversation carried out between a plurality of participant terminals 2, a spoofing detection process is realized to detect whether the video transmitted from each participant terminal 2 is that of the user himself / herself of the participant terminal 2 or a fake video (deepfake video) generated by an attacker using synthetic media.
[0021] In this computer system 1, when a remote conversation is carried out among a plurality of participants, it is assumed that an attacker may impersonate a participant in the remote conversation (participant). The participant impersonated by the attacker may be referred to as an attack target.
[0022] Also, it is assumed that an attacker can obtain information such as videos and voices of the target of the attack in advance for impersonation.
[0023] Furthermore, based on the information of the above-mentioned target of the attack, the attacker can impersonate the target of the attack using a known person generation tool (face conversion tool) or voice generation tool (voice conversion tool). That is, it is assumed that the attacker can participate in the meeting with the same face or the same voice as the target of the attack.
[0024] The attacker impersonates the target of the attack and conducts a remote conversation with other recipients using the account (the first account) of the target of the attack. When the attacker conducts impersonation using a deepfake video, the actual attacker is the target of the attack. The attacker who impersonates the target of the attack participates in the remote conversation using the account (the first account) of the target of the attack.
[0025] The plurality of participant terminals 2 are each a computer and have a similar configuration to each other. Each participant terminal 2 includes a processor, a memory, a display, a camera, a microphone, and a speaker (not shown).
[0026] Note that in each participant terminal 2, the processor, the memory, and the display are the same as the processor 11, the memory 12, and the monitor 14a in the information processing apparatus 10 described later with reference to FIG. 1, and detailed descriptions thereof are omitted.
[0027] In the participant terminal 2, the participant uses a camera to capture a video of his / her face or the like, and transmits the video data to the other participant terminal 3 and the information processing apparatus 10 during a remote conversation.
[0028] The video data transmitted from the participant terminal 2 is associated with the account of the participant who uses the participant terminal 2.
[0029] At each participant terminal 2, the participant acquires their own voice using a microphone and transmits the voice data to other participant terminals 3 and the information processing apparatus 10 during a remote conversation. At each participant terminal 2, the participant plays back the voice data transmitted from other participant terminals 2 using a speaker.
[0030] The video data transmitted from the participant terminal 2 is associated with the account of the participant using the participant terminal 2.
[0031] On the display of each participant terminal 2, the video of the participant transmitted from other participant terminals 3 is displayed. In the embodiment shown below, an example where the video is a moving image (video image) is shown. Also, hereinafter, the video data may be simply referred to as video. The video includes audio.
[0032] The host terminal 3 is a computer used by the host of a remote conversation (online meeting) and includes a processor, a memory, a display, a camera, a microphone, and a speaker (not shown).
[0033] Note that in the host terminal 3, the processor, the memory, and the display are the same as the processor 11, the memory 12, and the monitor 14a in the information processing apparatus 10 described later with reference to FIG. 1, respectively, and detailed descriptions thereof are omitted.
[0034] On the display of the host terminal 3, presentation information (message) output from the notification unit 107 of the information processing apparatus 10 described later is displayed.
[0035] The information processing apparatus 10 is a computer and includes, for example, as shown in FIG. 1, a processor 11, a memory 12, a storage device 13, a graphic processing device 14, an input interface 15, an optical drive device 16, a device connection interface 17, and a network interface 18 as components. These components 11 to 18 are configured to be communicable with each other via a bus 19.
[0036] The processor (control unit) 11 controls the entire information processing apparatus 10. The processor 11 may be a multi-processor. The processor 11 may be, for example, any one of a CPU, MPU (Micro Processing Unit), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), FPGA (Field Programmable Gate Array), and GPU (Graphics Processing Unit). Further, the processor 11 may be a combination of two or more types of elements among the CPU, MPU, DSP, ASIC, PLD, FPGA, and GPU.
[0037] Then, by the processor 11 executing a control program (determination program, OS program) for the information processing apparatus 10, functions as the first behavior detection unit 101, first behavior extraction unit 102, second behavior detection unit 104, second behavior extraction unit 105, identity determination unit 106, and notification unit 107, which will be described later with reference to FIG. 2, are realized. OS is an abbreviation for Operating System.
[0038] Programs describing the processing content to be executed by the information processing apparatus 10 can be recorded on various recording media. For example, a program to be executed by the information processing apparatus 10 can be stored in the storage device 13. The processor 11 loads at least a part of the program in the storage device 13 into the memory 12 and executes the loaded program.
[0039] Also, programs to be executed by the information processing apparatus 10 (processor 11) can be recorded on non-temporary portable recording media such as an optical disk 16a, a memory device 17a, and a memory card 17c. The program stored in the portable recording media becomes executable after being installed in the storage device 13, for example, under the control of the processor 11. Further, the processor 11 can also directly read and execute the program from the portable recording media.
[0040] The memory 12 is a storage memory including a ROM (Read Only Memory) and a RAM (Random Access Memory). The RAM of the memory 12 is used as the main memory device of the information processing apparatus 10. At least a part of the program to be executed by the processor 11 is temporarily stored in the RAM. Also, various data necessary for the processing by the processor 11 are stored in the memory 12.
[0041] The storage device 13 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a storage class memory (SCM), and stores various data. The storage device 13 is used as an auxiliary storage device of the information processing apparatus 10.
[0042] The storage device 13 stores an OS program, a control program, and various data. The control program includes a determination program. Also, the storage device 13 may store information constituting the database group 103. The database group 103 includes a plurality of databases.
[0043] Note that as the auxiliary storage device, a semiconductor storage device such as an SCM or a flash memory can also be used. Also, a RAID (Redundant Arrays of Inexpensive Disks) may be configured using a plurality of storage devices 13.
[0044] FIG. 3 is a diagram illustrating a plurality of databases included in the database group 103 in the computer system 1 as an example of the first embodiment.
[0045] In the example shown in FIG. 3, the database group 103 includes a first phrase corresponding text storage database 1031, a first face position information storage database 1032, a first skeleton position information storage database 1033, and a first behavior database 1034. Further, the database group 103 includes a second phrase corresponding text storage database 1035, a second face position information storage database 1036, a second skeleton position information storage database 1037, and a second behavior database 1038. The database may be represented as DB. DB is an abbreviation of Data Base.
[0046] Details of these, the first phrase corresponding text storage database 1031, the first face position information storage database 1032, the first skeleton position information storage database 1033, the first behavior database 1034, the second phrase corresponding text storage database 1035, the second face position information storage database 1036, the second skeleton position information storage database 1037, and the second behavior database 1038 will be described later.
[0047] The memory 12 and the storage device 13 may store data and the like generated in the process of the first behavior detection unit 101, the first behavior extraction unit 102, the second behavior detection unit 104, the second behavior extraction unit 105, the identity determination unit 106, and the notification unit 107 executing their respective processes.
[0048] A monitor 14a is connected to the graphic processing device 14. The graphic processing device 14 displays an image on the screen of the monitor 14a according to an instruction from the processor 11. Examples of the monitor 14a include a display device using a CRT (Cathode Ray Tube) and a liquid crystal display device.
[0049] The input interface 15 is connected to a keyboard 15a and a mouse 15b. The input interface 15 transmits signals sent from the keyboard 15a or the mouse 15b to the processor 11. Note that the mouse 15b is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include touch panels, tablets, touch pads, trackballs, etc.
[0050] The optical drive device 16 reads data recorded on the optical disc 16a using laser light or the like. The optical disc 16a is a portable non - temporary recording medium on which data is recorded in a readable manner by light reflection. Examples of the optical disc 16a include DVD (Digital Versatile Disc), DVD - RAM, CD - ROM (Compact Disc Read Only Memory), CD - R (Recordable) / RW (ReWritable), etc.
[0051] The device connection interface 17 is a communication interface for connecting peripheral devices to the information processing apparatus 10. For example, a memory device 17a and a memory reader / writer 17b can be connected to the device connection interface 17. The memory device 17a is a non - temporary recording medium equipped with a communication function with the device connection interface 17, such as a USB (Universal Serial Bus) memory. The memory reader / writer 17b writes data to or reads data from the memory card 17c. The memory card 17c is a card - type non - temporary recording medium.
[0052] The network interface 18 is connected to the network 20. The network interface 18 transmits and receives data via the network 20. The network 20 is connected to each participant terminal 2 and the host terminal 3. Note that other information processing apparatuses, communication devices, etc. may be connected to the network 20.
[0053] As shown in FIG. 2, the information processing apparatus 10 has functions as a first behavior detection unit 101, a first behavior extraction unit 102, a database group 103, a second behavior detection unit 104, a second behavior extraction unit 105, an identity determination unit 106, and a notification unit 107.
[0054] Among these, the first behavior detection unit 101 and the first behavior extraction unit 102 perform preprocessing using the video (video data) of remote conversations that have been held in the past among two or more participants. Hereinafter, the video data may be simply referred to as video. The video data includes audio data. Also, the audio data may be simply referred to as audio.
[0055] In addition, the second behavior detection unit 104, the second behavior extraction unit 105, the identity determination unit 106, and the notification unit 107 perform real-time processing using the video of an ongoing remote conversation (during the remote conversation) among two or more participants.
[0056] The first behavior detection unit 101 receives the video of past remote conversations held among two or more participants. This video includes the videos of the participants. The first behavior detection unit 101 may obtain it, for example, by reading out the video data of past remote conversations stored in the storage device 13.
[0057] Based on the video data of past remote conferences, the first behavior detection unit 101 detects phrases from the voices spoken by the participants, for example, by voice recognition processing. A phrase is a collection of multiple words (clause), and is a continuous string of words that represents a complete meaning. A phrase corresponds to the feature information of the actions or voices of the participants.
[0058] For the voice recognition processing, for example, feature quantity extraction processing is performed on the voices of the participants, and phrases are detected from the voices of the participants based on the extracted feature quantities. Note that the process of detecting phrases from the voices of the participants can be realized using various known methods, and the description thereof is omitted.
[0059] The first behavior detection unit 101 registers information regarding the extracted phrase in the first phrase-corresponding text storage database 1031.
[0060] FIG. 4 is a diagram illustrating the first phrase-corresponding text storage database 1031, the first face position information storage database 1032, and the first skeleton position information storage database 1033 in the computer system 1 as an example of the first embodiment.
[0061] In the first phrase-corresponding text storage database 1031 illustrated in FIG. 4, the start time, the end time, and the text (phrase) are associated with each other.
[0062] When the first behavior detection unit 101 detects that a participant has uttered some phrase in the video, it reads out the timestamps from the first frame and the last frame of the period in which the phrase is detected in the video, respectively. The timestamp read from the first frame may be the start time, and the timestamp read from the last frame may be the end time.
[0063] The first behavior detection unit 101 associates these start time and end time with the text representing the phrase and stores them in the first phrase-corresponding text storage database 1031. Note that the time zone (time frame) specified by the combination of these start time and end time may be referred to as the phrase detection time zone.
[0064] Further, the first behavior detection unit 101 performs, for example, an image recognition process (face detection process) on the video in the phrase detection time zone to detect the face of the participant and extracts the behavior in the face image. The behavior in the face image corresponds to the feature information of the action or state of the participant.
[0065] The first behavior detection unit 101 extracts the position information (coordinates) of a plurality (for example, 68) of feature points (Face Landmark) indicating eyes, nose, mouth, face contour, etc. for the detected face image, and detects the behavior in the face image by performing matching of these Face Landmarks. The detection of behavior in the face image can be realized using known methods, and detailed description thereof is omitted.
[0066] The first behavior detection unit 101 associates the coordinates of one or more feature points (Face Landmark) in the video with the time stamp of the frame in which the feature point is extracted in the video and records them in the first face position information storage database 1032.
[0067] The first face position information storage database 1032 illustrated in FIG. 4 associates time stamps with the coordinates (coordinate groups) of 68 feature points in the face image. By referring to this first face position information storage database 1032, the movement of the face (expression) in the video of past remote conversations can be detected as behavior. In the first face position information storage database 1032 illustrated in FIG. 4, the coordinate groups of feature points acquired every 0.1 second are registered as entries.
[0068] In addition, the first behavior detection unit 101 detects the skeletal structure of the participant by performing, for example, image recognition processing (gesture detection processing) on the video in the phrase detection time zone, and extracts the position information (coordinates) of the detected skeleton. The skeletal structure of the participant corresponds to the feature information of the participant's movement or state.
[0069] The detection of behavior in the skeletal structure can be realized by known methods, and detailed description thereof is omitted.
[0070] The first behavior detection unit 101 associates the coordinates of one or more feature points (skeletal positions) in the video with the time stamp of the frame in which the feature point is extracted in the video and records them in the first skeletal position information storage database 1033.
[0071] The first skeleton position information storage database 1033 illustrated in FIG. 4 associates time stamps with the coordinates of 15 feature points (skeleton positions) in the image. By referring to this first skeleton position information storage database 1033 and performing matching of the position changes of the feature points, the movement (gesture) of the skeleton can be detected as behavior. In the first skeleton position information storage database 1033 illustrated in FIG. 4, a coordinate group of feature points acquired every 0.1 second is registered as an entry.
[0072] Further, the first behavior detection unit 101 may extract, as feature amounts, voice channel characteristics and pitch corresponding to the phrases spoken or uttered by the participant by performing, for example, voice recognition processing (voice detection processing) on the video in the phrase detection time zone.
[0073] The first behavior detection unit 101 can detect voice as behavior by performing matching of the position changes of the time changes of one or more feature points (voice channel characteristics, pitch) in the voice included in the video. Detection of behavior in voice can be realized by a known method, and a detailed description thereof is omitted.
[0074] The first behavior detection unit 101 performs detection of phrases and detection of behaviors (for example, face movement, skeleton position movement) in the phrase detection time zone based on all the videos of the participant.
[0075] The first phrase correspondence text storage database 1031, the first face position information storage database 1032, and the first skeleton position information storage database 1033 are created for each participant.
[0076] Further, the first behavior detection unit 101 creates the first phrase correspondence text storage database 1031, the first face position information storage database 1032, and the first skeleton position information storage database 1033 for all the participants.
[0077] The first-phase corresponding text storage database 1031, the first face position information storage database 1032, and the first skeleton position information storage database 1033 for all participants may be collectively referred to as the full-behavior database. The full-behavior database may store the video (audio) data of the participants and the metadata that can be extracted from the video (audio) data.
[0078] Based on the full-behavior database generated by the first-behavior detection unit 101, the first-behavior extraction unit 102 extracts behaviors with low occurrence frequencies for each participant.
[0079] For the participant to be determined (hereinafter also referred to as the determination target participant), the first-behavior extraction unit 102 selects one phrase (determination target phrase) from among a plurality of phrases registered in the first-phase corresponding text storage database 1031 of the determination target participant, and reads out the text constituting this determination target phrase.
[0080] Then, the first-behavior extraction unit 102 extracts one or more words from the text of this determination target phrase. The words extracted from the determination target phrase may be referred to as extracted words. Note that the process of extracting words (extracted words) from the text can be realized using various known methods, and the description thereof is omitted.
[0081] The first-behavior extraction unit 102 calculates the occurrence frequency of the extracted words from among all the words spoken by the determination target participant in all the videos of the determination target participant. The first-behavior extraction unit 102 calculates the occurrence frequency in all the words for each of all the extracted words included in the determination target phrase.
[0082] Then, the first-behavior extraction unit 102 calculates the average value of the frequencies of the extracted words for the determination target phrase by calculating the average of the logarithmic sums of the frequencies of the plurality of extracted words included in the determination target phrase. The average value of the frequencies of the extracted words included in the determination target phrase may be referred to as the frequency average value of the determination target phrase. The first-behavior extraction unit 102 calculates the frequency in units of phrases.
[0083] When the average frequency value of the phrase to be determined calculated by the first behavior extraction unit 102 is smaller than the threshold value T0 (the first reference value), the phrase to be determined is registered in the first behavior database 1034 as a behavior with a low frequency for the participant. The first behavior database 1034 stores the feature information (behavior, phrase) of the participant whose appearance frequency (extraction frequency) is less than the threshold value T0 (the first reference value).
[0084] Specific phrases uttered by participants detected based on the video data of past remote meetings may be referred to as past phrases. Also, among the past phrases, a determination target phrase whose average frequency value is smaller than the threshold value T0 may be referred to as a past low-frequency phrase.
[0085] The first behavior database 1034 stores past low-frequency phrases for each participant. The first behavior database 1034 may, for example, associate information for identifying a participant with a determination target phrase determined as a behavior with a low frequency for the participant. Also, a first behavior database 1034 may be provided for each participant, and the determination target phrase determined as a behavior with a low frequency for the participant may be stored in this first behavior database 1034, and it can be implemented with appropriate modifications.
[0086] The first behavior extraction unit 102 sequentially switches the participants to be determined, and extracts behaviors with low appearance frequencies for each participant to be determined. Thereby, the first behavior extraction unit 102 extracts behaviors with low appearance frequencies for all participants. The appearance frequency may simply be referred to as the frequency.
[0087] The first behavior extraction unit 102 may determine the frequency based on the statistical quantity of ordinary people + the statistical quantity of the participant.
[0088] For example, in the case of voice, greetings such as "Good morning, everyone" or words that participants often say such as "What do you think about XX?" may be regarded as phrases with high frequencies.
[0089] Also, phrases including foreign words, foreign names, technical terms, etc. may be regarded as phrases with low frequency.
[0090] For example, in Japanese, words and phrases containing "ja", "rya", "bye", "mye", "jo", "cho" may be regarded as phrases with low frequency.
[0091] Also, in Japanese, phrases containing terms with consecutive "n" such as "two-thousand-yen bill", phrases containing words with voiceless "u" or "i", and phrases containing words with nasalized sounds (sounds that sound like "nga" or "ngi") may be regarded as phrases with low frequency.
[0092] Also, in English, words and phrases containing the sounds of the pronunciation symbols exemplified below may be regarded as phrases with low frequency.
Number
[0093] The second behavior detection unit 104 receives the video of a remote conversation (being executed in real time) among a plurality of participants. The video of the remote conversation (being executed in real time) among the plurality of participants corresponds to the first sensing data (video data) linked to the accounts of the participants in the remote conversation.
[0094] This video includes the video of each participant. The video of the remote conversation among the participants is generated, for example, by a program that realizes a remote conversation between two participant terminals and is transmitted to the information processing device 10. The program for realizing the remote conversation may operate on each participant terminal 2, or may also operate on the information processing device 10 or another information processing device having a server function.
[0095] The video of the remote conversation (being executed in real time) among the plurality of participants is stored in a predetermined storage area of, for example, the memory 12 or the storage device 13 of the information processing device 10. The second behavior detection unit 104 may obtain it by reading the stored video data of the remote conversation.
[0096] The second behavior detection unit 104 detects a specific phrase from the voices of the participants by performing speech recognition processing based on the input video of the ongoing (currently ongoing) remote conversation in real time.
[0097] A specific phrase uttered by a participant detected from the video of the ongoing (currently ongoing) remote conversation in real time may be referred to as the current phrase.
[0098] The second behavior detection unit 104 detects the current phrase from the voices of the participants using the same method as the first behavior detection unit 101.
[0099] The second behavior detection unit 104 registers information regarding the extracted phrase in the second phrase corresponding text storage database 1035. The second phrase corresponding text storage database 1035 has the same configuration as the first phrase corresponding text storage database 1031, and its description is omitted.
[0100] Also, the second behavior detection unit 104 performs, in the same manner as the first behavior detection unit 101, for example, image recognition processing (face detection processing) on the video in the phrase detection time zone in the video of the ongoing (currently ongoing) remote conversation in real time. As a result, the second behavior detection unit 104 detects the faces of the participants in the video of the ongoing (currently ongoing) remote conversation in real time, and extracts the position information (coordinates) of the feature points (Face Landmark) for the detected face images.
[0101] The second behavior detection unit 104 associates the coordinates of one or more feature points (Face Landmark) in the video of the ongoing (currently ongoing) remote conversation with the time stamp of the frame in which the feature point was extracted in the video and records them in the second face position information storage database 1036.
[0102] The second face position information storage database 1036 has the same configuration as the first face position information storage database 1032 illustrated in FIG. 4, and the description thereof is omitted.
[0103] By referring to the second face position information storage database 1036, in the video of the ongoing (currently ongoing) remote conversation in real time, the movement of the face (expression) can be detected as behavior.
[0104] Also, the second behavior extraction unit 105 performs image recognition processing (gesture detection processing) on the video in the phrase detection time zone in the video of the ongoing (currently ongoing) remote conversation in real time in the same manner as the first behavior detection unit 101. As a result, the second behavior extraction unit 105 detects the skeletal structure of the participant in the video of the ongoing (currently ongoing) remote conversation in real time and extracts the position information (coordinates) of the detected skeleton.
[0105] The second behavior extraction unit 105 associates the coordinates of one or more feature points (skeleton positions) in the video with the time stamp of the frame in which the feature point is extracted in the video and records them in the second skeleton position information storage database 1037.
[0106] The second skeleton position information storage database 1037 has the same configuration as the first skeleton position information storage database 1033 illustrated in FIG. 4, and the description thereof is omitted.
[0107] By referring to the second skeleton position information storage database 1037, in the video of the ongoing (currently ongoing) remote conversation in real time, the movement of the skeleton (gesture) can be detected as behavior.
[0108] The second behavior extraction unit 105 extracts behaviors with a low appearance frequency from the phrases (current phrases) detected by the second behavior detection unit 104 in the ongoing (currently ongoing) remote conversation in real time.
[0109] The second behavior extraction unit 105 checks whether a phrase (a past low-frequency phrase) that matches a phrase detected in a real-time ongoing remote conversation is registered as a low-frequency phrase of the same participant in the first behavior database 1034. As a result of this check, when a phrase identical to the current phrase is registered in the first behavior database 1034, a pair of these current phrases and the past low-frequency phrases is generated.
[0110] When the second behavior extraction unit 105 receives a video (first sensing data) of a remote conversation being conducted among a plurality of participants (in real-time execution), it acquires the feature information (behavior, phrase) of the participant that is extracted from the video of a past remote conversation (second sensing data) conducted among the participants and whose appearance frequency (extraction frequency) is less than a threshold value T0 (first reference value).
[0111] The pair of the current phrase and the past low-frequency phrase generated by the second behavior extraction unit 105 is generated on the premise that the speakers of each phrase have the same account.
[0112] The second behavior extraction unit 105 desirably generates a plurality (N) of pairs of the current phrase and the past low-frequency phrase.
[0113] The information on the pair of the current phrase and the past low-frequency phrase generated in this way may be stored, for example, in a predetermined area of the memory 12 or the storage device 13.
[0114] Based on the pair of the current phrase and the past low-frequency phrase generated by the second behavior extraction unit 105 using the same account, the identity determination unit 106 determines whether the participant who spoke the current phrase and the participant who spoke the past low-frequency phrase are the same.
[0115] The identity determination unit 106 obtains the behavior for the current phrase and the behavior for the past low-frequency phrase, for each pair of the current phrase and the past low-frequency phrase generated by the second behavior extraction unit 105. Here, the behavior for the current phrase may be referred to as the current behavior. Also, the behavior for the past low-frequency phrase may be referred to as the past behavior.
[0116] In the following, an example in which the behavior for the current phrase and the behavior for the past low-frequency phrase are voice signals corresponding to the phrases is shown.
[0117] The identity determination unit 106 obtains the past behavior (voice signal corresponding to the phrase) from the video data of the remote conversation conducted in the past, and obtains the current behavior (voice signal corresponding to the current phrase) from the video data of the remote conversation in progress (currently in progress) in real time.
[0118] The identity determination unit 106 performs matching between the current behavior (voice signal corresponding to the current phrase) and the past behavior (voice signal corresponding to the past low-frequency phrase) related to these same accounts.
[0119] FIG. 5 is a diagram for explaining a behavior matching method by the identity determination unit 106 in the computer system 1 as an example of the embodiment.
[0120] In this FIG. 5, an example is shown in which the identity determination unit 106 corrects the time-series deviation of the behavior using DTM (Dynamic Time Warping) and performs matching.
[0121] In FIG. 5, the past behavior (voice signal of the phrase) and the current behavior (voice signal of the phrase) are input to the DTW.
[0122] Also, as the output of the DTW, a graph is shown in which the vertical axis is the past behavior (voice signal of the phrase) and the horizontal axis is the current behavior (voice signal of the phrase). This graph shows which parts of the time-series signals correspond to each other.
[0123] In the method using DTM, a value obtained by dividing the distance (the magnitude of deviation), which is the output of DTW, by the past and current time series lengths may be used as a matching score. The minimum value of the matching score may be set to 0.0 and the maximum value may be set to 1.0. The matching score when completely matching (coinciding) is 0, and the matching score when not matching at all (not coinciding) is 1.
[0124] For each of a plurality (N) of pairs of the current phrase generated by the second behavior extraction unit 105 and the past low-frequency phrases, the identity determination unit 106 obtains matching scores D1 to Dn between the current behavior (the audio signal corresponding to the current phrase) and the past behavior (the audio signal corresponding to the past low-frequency phrase).
[0125] That is, the identity determination unit 106 calculates the degree of coincidence (matching score) for each of a plurality (N) of pairs of the phrase (feature information) extracted from the video (first sensing data) of the remote conversation being conducted (executed in real time) among the participants and the low-frequency phrase (feature information) extracted from the video (second sensing data) of the past remote conversation conducted among the participants.
[0126] Then, the identity determination unit 106 compares each of the obtained matching scores D1 to Dn with a predetermined threshold value T1 (second reference value), and obtains the number of matching scores that are less than the threshold value T1, that is, the number of pairs of the current phrase and the past low-frequency phrase.
[0127] The identity determination unit 106 compares the number of pairs of the current phrase and the past low-frequency phrase that are less than the threshold value T1 with a predetermined threshold value T2 (third reference value).
[0128] When the number of pairs of the current phrase and the past low-frequency phrase where the matching score is less than the threshold T1 is equal to or greater than the threshold T2, the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrase are the same for the pair of the current phrase and the past low-frequency phrase.
[0129] On the other hand, when the number of pairs of the current phrase and the past low-frequency phrase where the matching score is less than the threshold T1 is less than the threshold T2, the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrase are not the same for the pair of the current phrase and the past low-frequency phrase.
[0130] When the number of pairs where the degree of match (matching score) is less than the threshold T1 (second reference value) is less than the threshold T2 (third reference value), the identity determination unit 106 determines that impersonation has occurred.
[0131] The identity determination unit 106 determines the participant who uttered the current phrase, who has been determined not to be the same as the participant who uttered the past low-frequency phrase associated with the same account, as an impersonating participant.
[0132] The identity determination unit 106 makes a determination regarding impersonation based on the degree of match (matching score) between the phrases (feature information) extracted from the video (first sensing data) of a remote conversation (being executed in real time) conducted among a plurality of participants and the phrases (feature information) extracted from the video (second sensing data) of a past remote conversation conducted among the participants.
[0133] When the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrase associated with the same account are not the same for the pair of the current phrase and the past low-frequency phrase, the notification unit 107 notifies the organizer.
[0134] The notification unit 107 may transmit a message (notification information) indicating that "there may be a possibility of impersonation by a participant" to the organizer terminal 3. Further, the notification unit 107 may notify the organizer terminal 3 of information (for example, account information; notification information) identifying the impersonating participant, together with the message, which is determined by the identity determination unit 106.
[0135] The notification unit 107 may, for example, cause information (message; notification information) indicating that "there may be a possibility of impersonation by a participant" to be displayed on the display of the organizer terminal 3.
[0136] At the organizer terminal 3, the organizer may, for example, remove a participant determined to be an impersonating participant from the remote conversation. Further, the organizer may confirm whether the determination by the identity determination unit 106 is correct by asking some question (for example, a question that only the correct participant can answer) to the participant determined to be an impersonating participant.
[0137] (B) Operations The processing of the first behavior detection unit 101 in the computer system 1 as an example of the first embodiment configured as described above will be described according to the flowchart (steps A1 to A4) shown in FIG. 6.
[0138] Video data of a remote conference held by a participant in the past is input to the first behavior detection unit 101.
[0139] Based on the video data of the remote conference held in the past, the first behavior detection unit 101 detects phrases from the voice spoken by the participant by voice recognition processing (step A1).
[0140] Also, based on the video data of the remote conference held in the past, the first behavior detection unit 101 performs face detection of the participant by performing image recognition processing. Further, the first behavior detection unit 101 extracts position information (coordinates) of feature points (Face Landmark) for the detected face image.
[0141] Furthermore, the first behavior detection unit 101 performs gesture detection processing by performing image recognition processing based on video data of a remote conference held in the past (step A3). In addition, the first behavior detection unit 101 detects the skeletal structure of the detected participants and extracts the position information (coordinates) of the detected skeleton.
[0142] The processes of steps A1 to A3 described above may be performed in parallel, or for example, the processes of steps A2 and A3 may be performed after the process of step A1, and can be appropriately changed and implemented.
[0143] Thereafter, in step A4, the first behavior detection unit 101 associates the start time and end time of a phrase in the video data of the remote conference held in the past with the text representing the phrase and stores it in the first phrase correspondence text storage database 1031.
[0144] In addition, the first behavior detection unit 101 associates the position information (coordinates of Face Landmark) of the facial parts (feature points) of the participants in the video with the time stamp and records it in the first face position information storage database 1032.
[0145] Furthermore, the first behavior detection unit 101 associates the coordinates (skeletal position information) of one or more skeletal positions (feature points) in the video with the time stamp and records it in the first skeletal position information storage database 1033. Then, the process ends.
[0146] Next, the processing of the first behavior extraction unit 102 in the computer system 1 as an example of the first embodiment will be described according to the flowchart (steps B1 to B4) shown in FIG. 7.
[0147] The first behavior extraction unit 102 is input with a full behavior database for all participants generated by the first behavior detection unit 101.
[0148] In step B1, the first behavior extraction unit 102 acquires the text corresponding to the phrase (phrase to be determined) from the first phrase-corresponding text storage database 1031.
[0149] In step B2, the first behavior extraction unit 102 calculates the appearance frequency of the extracted words from all the words spoken by the participant to be determined in all the videos of the participant to be determined. The first behavior detection unit 101 calculates the appearance frequency of each of all the extracted words included in the phrase to be determined in all the words.
[0150] The first behavior extraction unit 102 calculates the average value of the frequencies of the extracted words for the phrase to be determined by calculating the average of the logarithmic sums of the frequencies of the plurality of extracted words included in the phrase to be determined.
[0151] In step B3, the first behavior extraction unit 102 checks whether the calculated average frequency value of the phrase to be determined is smaller than the threshold value T0. As a result of the check, if the calculated average frequency value of the phrase to be determined is smaller than the threshold value T0 (refer to the YES route in step B3), the process proceeds to step B4.
[0152] In step B4, the first behavior extraction unit 102 registers the phrase to be determined in the first behavior database 1034 as a behavior with a low frequency for the participant. Then, the process ends.
[0153] Also, as a result of the check in step B3, if the calculated average frequency value of the phrase to be determined is equal to or greater than the threshold value T0 (refer to the NO route in step B3), step B4 is skipped and the process ends.
[0154] Next, the processing of the second behavior detection unit 104 in the computer system 1 as an example of the first embodiment will be described according to the flowchart (steps C1 to C4) shown in FIG. 8.
[0155] The second behavior detection unit 104 receives the video of a remote conversation (being executed in real time) conducted among a plurality of participants.
[0156] Based on the video data of the remote conversation being conducted in real time among a plurality of participants, the second behavior detection unit 104 detects phrases from the voices uttered by the participants through speech recognition processing (step C1).
[0157] Also, based on the video data of the remote conversation being conducted in real time among a plurality of participants, the second behavior detection unit 104 performs face detection of the participants by performing image recognition processing (step C2). Further, based on the video data of past remote conferences, the second behavior detection unit 104 extracts the position information (coordinates) of feature points (Face Landmark) for the detected face images.
[0158] Furthermore, based on the video data of the remote conversation being conducted in real time among a plurality of participants, the second behavior detection unit 104 performs gesture detection processing by performing image recognition processing (step C3). Also, the second behavior detection unit 104 detects the skeletal structure of the detected participants and extracts the position information (coordinates) of the detected skeleton.
[0159] The processes of steps C1 to C3 described above may be performed in parallel, or for example, the processes of steps C2 and C3 may be performed after the process of step C1, and can be appropriately changed and implemented.
[0160] Thereafter, in step C4, the second behavior detection unit 104 associates the start time and end time of the phrases in the video data of the remote conversation being conducted in real time among a plurality of participants with the text representing the phrases and stores them in the second phrase-corresponding text storage database 1035.
[0161] Also, the second behavior detection unit 104 records the position information (coordinates of Face Landmark) of the facial parts of the participants in the video in association with the timestamps in the second face position information storage database 1036.
[0162] Furthermore, the second behavior detection unit 104 records the coordinates of one or more skeleton positions (skeleton position information) in the video, in association with the time stamp, in the second skeleton position information storage database 1037. Then, the process ends.
[0163] Next, the processing of the second behavior extraction unit 105 in the computer system 1 as an example of the first embodiment will be described according to the flowchart (steps D1 to D4) shown in FIG. 9.
[0164] In step D1, the second behavior detection unit 104 acquires (extracts) the text corresponding to the phrase detected by the second behavior detection unit 104 from the second phrase corresponding text storage database 1035. The phrase detected by the second behavior detection unit 104 from the video data of the remote conversation being performed in real time among a plurality of participants may be referred to as phrase X.
[0165] In step D2, the second behavior extraction unit 105 checks whether a phrase (past low-frequency phrase) that matches the phrase X detected in step D1 is registered as a low-frequency phrase of the same participant (same account) in the first behavior database 1034.
[0166] As a result of the check, if a phrase (past low-frequency phrase) that matches phrase X is not registered as a low-frequency phrase of the same participant (same account) in the first behavior database 1034 (see the NO route in step D2), the process returns to step D1.
[0167] If a phrase (past low-frequency phrase) that matches phrase X is registered as a low-frequency phrase of the same participant (same account) in the first behavior database 1034 (see the YES route in step D2), the process proceeds to step D3. Note that the same low-frequency phrase of the same participant (same account) registered in the first behavior database 1034 may be referred to as past phrase Y.
[0168] In step D3, the second behavior extraction unit 105 stores the phrase X and the phrase Y as a pair in a predetermined area of, for example, the memory 12 or the storage device 13.
[0169] In step D4, the second behavior extraction unit 105 checks whether the number of pairs of the phrase X and the phrase Y stored in a predetermined area of the memory 12 or the storage device 13 is equal to or greater than a predetermined number (N).
[0170] As a result of the check, if the number of pairs of the phrase X and the phrase Y is less than the predetermined number (N) (see the NO route in step D4), the process returns to step D1.
[0171] On the other hand, if the number of pairs of the phrase X and the phrase Y is equal to or greater than the predetermined number (N) (see the YES route in step D4), the process ends.
[0172] Next, the processing of the identity determination unit 106 in the computer system 1 as an example of the first embodiment will be described according to the flowchart shown in FIG. 10 (steps E1 to E6).
[0173] In step E1, N pairs of the current phrase generated by the second behavior extraction unit 105 and the past low-frequency phrase by the same account are input to the identity determination unit 106.
[0174] In step E2, the identity determination unit 106 acquires the behavior for the current phrase and the behavior for the past low-frequency phrase, respectively.
[0175] In step E3, for each of the plurality of (N) pairs of the current phrase and the past low-frequency phrase, the identity determination unit 106 acquires the matching scores D1 to Dn between the current behavior (the voice signal corresponding to the current phrase) and the past behavior (the voice signal corresponding to the past low-frequency phrase).
[0176] In step E4, the identity determination unit 106 compares each of the acquired matching scores D1 to Dn with a predetermined threshold value T1, and checks whether the number of matching scores less than the threshold value T1 is equal to or greater than a threshold value T2. For example, the threshold value T1 may be 0.25, and the threshold value T2 may be 2.
[0177] As a result of the check, if the number of matching scores less than the threshold value T1 is equal to or greater than the threshold value T2 (see the YES route in step E4), the process proceeds to step E5.
[0178] In step E5, the identity determination unit 106 determines that the participants who uttered the current phrase and the participants who uttered the past low-frequency phrase are the same for the pair of the current phrase and the past low-frequency phrase. Then, the process ends.
[0179] On the other hand, if the number of matching scores less than the threshold value T1 is less than the threshold value T2 (see the NO route in step E4), the process proceeds to step E6.
[0180] In step E6, the identity determination unit 106 determines that the participants who uttered the current phrase and the participants who uttered the past low-frequency phrase are not the same for the pair of the current phrase and the past low-frequency phrase. Then, the process ends.
[0181] Next, the processing of the notification unit 107 in the computer system 1 as an example of the first embodiment will be described according to the flowchart shown in FIG. 11 (steps F1 to F2).
[0182] In step F1, the notification unit 107 checks whether the identity determination unit 106 has determined that the participants who uttered the current phrase and the participants who uttered the past low-frequency phrase are the same for the pair of the current phrase and the past low-frequency phrase associated with the same account.
[0183] When the identity determination unit 106 does not determine that the participant who uttered the current phrase is the same as the participant who uttered the past low-frequency phrase (refer to the NO route in step F1), the process proceeds to step F2.
[0184] In step F2, the notification unit 107 notifies the organizer that "there is a possibility that a participant is impersonating". Then, the process ends.
[0185] Also, as a result of the confirmation in step F1, when the identity determination unit 106 determines that the participant who uttered the current phrase is the same as the participant who uttered the past low-frequency phrase (refer to the YES route in step F1), the process ends as it is.
[0186] Next, FIG. 12 shows an example of applying the impersonation determination method in the computer system 1 as an example of the first embodiment to a remote conference system.
[0187] In the example shown in this FIG. 12, an example is shown in which three participants A, B, and C participate in a remote conference hosted by the organizer.
[0188] First, based on the video data of the remote conferences previously attended by participants A, B, and C, preprocessing by the first behavior detection unit 101 and the first behavior extraction unit 102 is performed. Note that the video data of the remote conferences previously attended by participants A, B, and C does not necessarily have to be the video data of the remote conferences in which all of participants A, B, and C participated. Video data of a plurality of remote conferences individually participated in by participants A, B, and C may be used.
[0189] The first behavior detection unit 101 detects phrases for each of the participants A, B, and C and acquires the text corresponding to the detected phrases based on the video data when the participants A, B, and C participated in past remote conferences.
[0190] In addition, based on the video data when participants A, B, and C participated in past remote meetings, the first behavior detection unit 101 extracts the feature points (Face Landmark, skeletal position information) of the skeletal position information storage database 1033 structure for the face images of each of the participants A, B, and C, and generates a full behavior database.
[0191] Then, based on the full behavior database generated by the first behavior detection unit 101, the first behavior extraction unit 102 extracts behaviors with low occurrence frequencies for each participant (see reference symbol P1 in FIG. 12).
[0192] Next, based on the remote conversation being conducted in real time among a plurality of participants A, B, and C, real-time processing is performed by the second behavior detection unit 104, the second behavior extraction unit 105, the identity determination unit 106, and the notification unit 107.
[0193] The second behavior detection unit 104 detects phrases for each of the participants A, B, and C and obtains the text corresponding to the detected phrases based on the video data when participating in the remote meeting being conducted in real time among the participants A, B, and C.
[0194] In addition, based on the video data when participating in the remote meeting being conducted in real time among the participants A, B, and C, the second behavior detection unit 104 extracts the feature points (Face Landmark, skeletal position information) of the skeletal position information storage database 1033 structure for the face images of each of the participants A, B, and C, and generates a full behavior database. The second behavior extraction unit 105 generates a plurality of pairs of the current phrases detected by the second behavior detection unit 104 and the past low-frequency phrases for each of the participants A, B, and C.
[0195] Thereafter, based on the pairs of the current phrases and the past low-frequency phrases generated by the second behavior extraction unit 105 for each of the participants A, B, and C, the identity determination unit 106 determines whether the participant who uttered the current phrase and the participant who uttered the past low-frequency phrase are the same (see reference symbol P2).
[0196] In the example shown in FIG. 12, Participant C is the target of the attack, and the video to be transmitted associated with the account of this Participant C is a fake video generated by the attacker using deep fake.
[0197] For example, in voice synthesis that generates spoofing data from scratch, a large amount of data is used to create a generation model from scratch. However, when attempting to generate less frequent data, there is a characteristic that the quality deteriorates.
[0198] Also, for example, in voice conversion that generates spoofing data using a standard model, a generation model (accurately, a difference model of the standard model) is created using a pre-created standard model and a small amount of data. When generating the behavior of a target person with low frequency using such a voice conversion method, the quality is less likely to deteriorate, but there is a characteristic that the authenticity (person-specific behavior) decreases. Therefore, the reproducibility of low-frequency phrases in the fake video becomes low.
[0199] When the number of pairs of the current phrase and the past low-frequency phrases for which the matching score is less than the threshold T1 is less than the threshold T2, the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrases are not the same for the pair of the current phrase and the past low-frequency phrases (see reference sign P3).
[0200] When the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrases are not the same, the notification unit 107 notifies the meeting organizer (see reference sign P4).
[0201] (C) Effect As described above, according to the computer system 1 as an example of the first embodiment, the first behavior extraction unit 102 extracts behaviors with low appearance frequency for the participants based on the video data of the remote conversations conducted in the past. The first behavior extraction unit 102 registers the determination target phrase in the first behavior database 1034 as a behavior (feature information) with low frequency for the participants.
[0202] Further, the second behavior extraction unit 105 generates a plurality (N) of pairs of the current phrase and past low-frequency phrases.
[0203] Then, for each of the plurality (N) of pairs of the current phrase and past low-frequency phrases generated by the second behavior extraction unit 105, the identity determination unit 106 obtains matching scores D1 to Dn between the current behavior (audio signal corresponding to the current phrase) and the past behavior (audio signal corresponding to the past low-frequency phrase).
[0204] When the number of pairs of the current phrase and the past low-frequency phrase is less than the threshold value T2, the identity determination unit 106 determines that the participant who uttered the current phrase and the participant who uttered the past low-frequency phrase are not the same for the pair of the current phrase and the past low-frequency phrase.
[0205] Thereby, it is possible to easily determine whether a participant in a remote conversation is an impersonation by an attacker.
[0206] (II) Description of the Second Embodiment (A) Configuration FIG. 13 is a diagram illustrating a functional configuration of a computer system 1 as an example of the second embodiment.
[0207] As shown in this FIG. 13, the computer system 1 of the second embodiment is provided with an authority change unit 108 instead of the notification unit 107 of the computer system 1 of the first embodiment, and other parts are configured in the same manner as the computer system 1 of the first embodiment.
[0208] In this second embodiment, by the processor 11 executing the determination program, the functions as the first behavior detection unit 101, the first behavior extraction unit 102, the second behavior detection unit 104, the second behavior extraction unit 105, the identity determination unit 106, and the authority change unit 108 are realized.
[0209] In the figure, since the same reference numerals as the above-described reference numerals indicate the same parts, the description thereof will be omitted.
[0210] The permission change unit 108 has a function of changing the participation permission of a participant (account) for a remote conversation. For example, the permission change unit 108 deprives a participant of the participation permission to participate in a remote conversation and makes the participant leave the remote conversation.
[0211] When the identity determination unit 106 determines that the participants who uttered the current phrase and the past low-frequency phrase related to the same account are not the same, the permission change unit 108 deprives the participant (account) of the participation permission for the remote conversation.
[0212] In order to allow a participant whose participation permission for a remote conversation has been deprived to re-participate in the remote conversation, for example, some penalty may be imposed on the participant, such as not being able to re-participate in the remote conversation until a predetermined time (for example, 30 minutes) has elapsed after the participation permission for the remote conversation has been deprived.
[0213] (B) Operation The processing of the permission change unit 108 in the computer system 1 as an example of the second embodiment will be described according to the flowchart (steps G1 to G2) shown in FIG. 14.
[0214] This processing is started when the identity determination unit 106 determines whether the participants who uttered the current phrase and the past low-frequency phrase are the same.
[0215] In step G1, the permission change unit 108 checks whether the identity determination unit 106 has determined that the participants who uttered the current phrase and the past low-frequency phrase are the same.
[0216] As a result of the confirmation, when the identity determination unit 106 determines that the participant who uttered the current phrase is not the same as the participant who uttered the past low-frequency phrase (refer to the NO route in step G1), the process proceeds to step G2.
[0217] In step G2, the authority change unit 108 revokes the participation authority of the participant (account) for the remote conversation and makes the participant leave the remote conversation. Then, the process ends.
[0218] Also, as a result of the confirmation, when the identity determination unit 106 determines that the participant who uttered the current phrase is the same as the participant who uttered the past low-frequency phrase (refer to the YES route in step G1), the process ends as it is.
[0219] (C) Effect As described above, according to the computer system 1 as an example of the second embodiment, the same operational effects as those of the first embodiment described above can be obtained.
[0220] Also, when the identity determination unit 106 determines that the participant who uttered the current phrase is not the same as the participant who uttered the past low-frequency phrase, the authority change unit 108 revokes the participation authority of the participant (account) for the remote conversation and makes the participant leave the remote conversation.
[0221] Thereby, it is highly convenient because the organizer does not need to take any action against a participant who may be impersonating. Also, by promptly making a participant who is highly likely to be impersonating leave the remote conversation, the security of the remote conversation can be improved.
[0222] (III) Description of the Third Embodiment (A) Configuration FIG. 15 is a diagram illustrating the functional configuration of the computer system 1 as an example of the third embodiment.
[0223] As shown in FIG. 15, the computer system 1 of the third embodiment includes a first behavior extraction unit 102a instead of the first behavior extraction unit 102 of the computer system 1 of the first embodiment, a second behavior extraction unit 105a instead of the second behavior extraction unit 105, and an identity determination unit 106a instead of the identity determination unit 106. Other parts are configured in the same manner as the computer system 1 of the first embodiment.
[0224] In this third embodiment, when the processor 11 executes the determination program, functions as the first behavior detection unit 101, the first behavior extraction unit 102a, the second behavior detection unit 104, the second behavior extraction unit 105a, the identity determination unit 106a, and the notification unit 107 are realized.
[0225] In the figure, the same reference numerals as the above-mentioned ones indicate the same parts, so the description thereof is omitted.
[0226] Based on the entire behavior database generated by the first behavior detection unit 101, the first behavior extraction unit 102a extracts the behaviors with high and low occurrence frequencies for each participant, respectively.
[0227] The first behavior extraction unit 102a calculates the occurrence frequency of the extracted words from all the words spoken by the participant to be determined in all the videos of the participant to be determined. The first behavior extraction unit 102a calculates the occurrence frequency in all the words for each of the extracted words included in the phrase to be determined.
[0228] Then, the first behavior extraction unit 102a calculates the average value of the frequencies of the extracted words for the phrase to be determined by calculating the average of the logarithmic sums of the frequencies of the plurality of extracted words included in the phrase to be determined.
[0229] When the calculated average frequency value of the phrase to be determined is smaller than the threshold value T01, the first behavior extraction unit 102a registers the phrase to be determined in the first behavior database 1034 as a behavior with a low frequency for the participant.
[0230] In addition, when the average frequency value of the phrase to be determined calculated by the first behavior extraction unit 102a is greater than the threshold value T02, the phrase to be determined is registered in the first behavior database 1034 as a behavior with a high frequency for the participant.
[0231] The second behavior extraction unit 105a extracts, from among the phrases (current phrases) detected by the second behavior detection unit 104 in a remote conversation that is in progress (currently in progress) in real time, the behaviors with a low appearance frequency and the behaviors with a high appearance frequency, respectively.
[0232] The second behavior extraction unit 105a checks whether a phrase that matches the phrase detected in the remote conversation in progress (currently in progress) in real time is registered as a low-frequency phrase or a high-frequency phrase of the same participant in the first behavior database 1034.
[0233] As a result of this check, when a phrase identical to the current phrase is registered as a low-frequency phrase in the first behavior database 1034, a pair (low-frequency pair) of these current phrases and the past low-frequency phrases is generated.
[0234] In addition, when a phrase identical to the current phrase is registered as a high-frequency phrase in the first behavior database 1034, a pair (high-frequency pair) of these current phrases and the past high-frequency phrases is generated.
[0235] The low-frequency pairs and high-frequency pairs generated by the second behavior extraction unit 105 are generated on the premise that the speakers of the respective phrases are the same account.
[0236] The second behavior extraction unit 105 preferably generates a plurality (N) of high-frequency pairs and low-frequency pairs, respectively.
[0237] The information on the high-frequency pairs and low-frequency pairs generated in this way may be stored, for example, in a predetermined area of the memory 12 or the storage device 13.
[0238] The identity determination unit 106a determines whether the participant who uttered the current phrase is the same as the participant who uttered the past low-frequency phrases based on the high-frequency pairs and low-frequency pairs generated by the second behavior extraction unit 105 using the same account.
[0239] In the computer system 1 as an example of the third embodiment, the identity determination unit 106a determines that there is a possibility of impersonation when the following determination conditions 1 and 2 are not satisfied.
[0240] Condition 1: Degree of match of frequently occurring behavior < threshold Th, degree of match of rarely occurring behavior < threshold Tl Condition 2: (Degree of match of rarely occurring behavior) - (Degree of match of frequently occurring behavior) > threshold Td FIG. 16 is a diagram for explaining a method for determining the possibility of impersonation by the identity determination unit 106a in the computer system 1 as an example of the third embodiment.
[0241] In this FIG. 16, the degree of match (matching score) of frequently occurring behavior and the degree of match (matching score) of rarely occurring behavior are shown on two-dimensional coordinates with the horizontal axis being frequency and the vertical axis being the matching score.
[0242] The degree of match of frequently occurring behavior is less than the threshold Th, and the degree of match of rarely occurring behavior is less than the threshold Tl, satisfying the above condition 1.
[0243] When the difference between the degree of match of rarely occurring behavior and the degree of match of frequently occurring behavior is large in the same participant, the possibility of impersonation is high. Therefore, when the difference between the degree of match of rarely occurring behavior (degree of match of low-frequency pairs) and the degree of match of frequently occurring behavior (degree of match of high-frequency pairs) is greater than a predetermined threshold Td (condition 2), the identity determination unit 106a determines that the participant who uttered the current phrase and the participant who uttered the past phrase are not the same.
[0244] The identity determination unit 106a obtains the degree of match (matching scores L1 to Ln) between the second feature information (behavior with low frequency) extracted from the video of the remote conversation being executed in real time among a plurality of participants and having a frequency less than the threshold value Tl (the fourth reference value), and the second feature information (behavior with low frequency) extracted from the video of the past remote conversation (the second sensing data) conducted among the participants.
[0245] Also, the identity determination unit 106 obtains the degree of match (matching scores H1 to Hn) between the first feature information (behavior with high frequency) extracted from the video of the remote conversation being executed in real time among a plurality of participants and having a frequency greater than the threshold value Th (the fifth reference value), and the first feature information (behavior with high frequency) extracted from the video of the past remote conversation (the second sensing data) conducted among the participants.
[0246] Then, when the number of pairs in which the difference (L1 - H1, L2 - H2, ··· Ln - Hn) between these degrees of match is less than the threshold value Td (the sixth reference value) is equal to or greater than the threshold value Tn (the seventh reference value), the identity determination unit 106 determines that impersonation has occurred.
[0247] (B) Operations The processing of the first behavior extraction unit 102a in the computer system 1 as an example of the third embodiment will be described according to the flowchart shown in FIG. 17 (steps H1 to H6).
[0248] The entire behavior database for all participants generated by the first behavior detection unit 101 is input to the first behavior extraction unit 102a.
[0249] In step H1, the first behavior extraction unit 102a acquires the text corresponding to the phrase (phrase to be determined) from the first phrase - corresponding text storage database 1031.
[0250] In step H2, the first behavior extraction unit 102a calculates the appearance frequency of the extracted words from all the words spoken by the participant to be determined in all the videos of the participant to be determined. The first behavior detection unit 101 calculates the appearance frequency of each of the extracted words included in the phrase to be determined in all the words.
[0251] The first behavior extraction unit 102a calculates the average value of the frequencies of the extracted words for the phrase to be determined by calculating the average of the logarithmic sums of the frequencies of the plurality of extracted words included in the phrase to be determined.
[0252] In step H3, the first behavior extraction unit 102a checks whether the calculated average frequency value of the phrase to be determined is less than the threshold value Tl. For example, the threshold value Tl may be -1000. As a result of the check, if the calculated average frequency value of the phrase to be determined is less than the threshold value Tl (see the YES route in step H3), the process proceeds to step H4.
[0253] In step H4, the first behavior extraction unit 102a registers the phrase to be determined in the first behavior database 1034 as a behavior with a low frequency for the participant. Then, the process ends.
[0254] Also, as a result of the check in step H3, if the calculated average frequency value of the phrase to be determined is greater than or equal to the threshold value Tl (see the NO route in step H3), step H4 is skipped and the process ends.
[0255] Also, in step H5, the first behavior extraction unit 102a checks whether the calculated average frequency value of the phrase to be determined is greater than the threshold value Th. For example, the threshold value Th may be -100. As a result of the check, if the calculated average frequency value of the phrase to be determined is greater than the threshold value Th (see the YES route in step H5), the process proceeds to step H6.
[0256] In step H6, the first behavior extraction unit 102a registers the phrase to be determined as first behavior data in the first behavior database 1034 as a behavior with a high frequency for the participant. Then, the process ends.
[0257] Also, as a result of the confirmation in step H5, if the calculated average frequency value of the phrase to be determined is equal to or less than the threshold Th (refer to the NO route in step H5), step H6 is skipped and the process ends.
[0258] Next, the processing of the identity determination unit 106a in the computer system 1 as an example of the third embodiment will be described according to the flowchart shown in FIG. 18 (steps J1 to J7).
[0259] In step J1, N pairs of the current phrase generated by the second behavior extraction unit 105a and the past low-frequency phrases by the same account are input to the identity determination unit 106a.
[0260] In step J2, the identity determination unit 106a acquires N pairs each of the pair of the current phrase and the past low-frequency phrase (low-frequency pair) and the pair of the current phrase and the past high-frequency phrase (high-frequency pair).
[0261] In step J3, for each of the N pairs of the current phrase and the past high-frequency phrases (high-frequency pairs), the identity determination unit 106a acquires the matching scores H1 to Hn between the current behavior (the voice signal corresponding to the current phrase) and the past behavior (the voice signal corresponding to the past high-frequency phrase).
[0262] In step J4, for each of the N pairs of the current phrase and the past low-frequency phrases (low-frequency pairs), the identity determination unit 106a acquires the matching scores L1 to Ln between the current behavior (the voice signal corresponding to the current phrase) and the past behavior (the voice signal corresponding to the past low-frequency phrase).
[0263] In step J5, the identity determination unit 106a compares each of the acquired matching scores H1 to Hn with a threshold Th, and checks whether each of the matching scores H1 to Hn is less than the threshold Th (condition A). For example, the threshold Th may be 0.25.
[0264] Also, the identity determination unit 106a compares each of the acquired matching scores L1 to Ln with a threshold Tl, and checks whether each of the matching scores L1 to Ln is less than the threshold Tl (condition B). For example, the threshold Tl may be 0.25.
[0265] Furthermore, the identity determination unit 106a calculates the differences between the matching scores, L1 - H1, L2 - H2, ··· Ln - Hn, and checks whether there are at least a threshold number Tn of pairs for which these differences between the matching scores are less than a threshold Td (condition C). For example, the threshold Td may be 0.1, and the threshold Tn may be 2.
[0266] As a result of the check, if all of conditions A, B, and C are satisfied (see the YES route in step J5), the process proceeds to step J6.
[0267] In step J6, the identity determination unit 106a determines that the participant who uttered the current phrase and the participant who uttered the past phrase are the same. Then, the process ends.
[0268] On the other hand, as a result of the check in step J5, if at least one of conditions A, B, and C is not satisfied (see the NO route in step J5), the process proceeds to step J7.
[0269] In step J7, the identity determination unit 106a determines that the participant who uttered the current phrase and the participant who uttered the past phrase are not the same. Then, the process ends.
[0270] (C) Effect Thus, according to the computer system 1 as an example of the third embodiment, the same operational effects as those of the above-described first embodiment can be obtained.
[0271] In addition, when the identity determination unit 106 determines that the participant who uttered the current phrase is not the same as the participant who uttered the past low-frequency phrase, the authority change unit 108 deprives the participant (account) of the participation right for the remote conversation and makes the participant withdraw from the remote conversation.
[0272] Thereby, it is highly convenient because the organizer does not need to take any action with respect to the participant who may be impersonating. In addition, by promptly making the participant who is highly likely to be impersonating withdraw from the remote conversation, the security of the remote conversation can be improved.
[0273] (IV) Description of the Fourth Embodiment (A) Configuration FIG. 19 is a diagram illustrating the functional configuration of the computer system 1 as an example of the fourth embodiment.
[0274] As shown in this FIG. 19, the computer system 1 of the fourth embodiment includes an authority change unit 108 instead of the notification unit 107 of the computer system 1 of the third embodiment, and the other parts are configured in the same manner as the computer system 1 of the third embodiment.
[0275] In this fourth embodiment, by the processor 11 executing the determination program, the functions as the first behavior detection unit 101, the first behavior extraction unit 102a, the second behavior detection unit 104, the second behavior extraction unit 105a, the identity determination unit 106a, and the authority change unit 108 are realized.
[0276] In the figure, the same reference numerals as the above-mentioned ones indicate the same parts, and thus the description thereof is omitted.
[0277] (B) Effect Thus, according to the computer system 1 as an example of the fourth embodiment, the same operational effects as those of the above-described third embodiment can be obtained.
[0278] In addition, when the identity determination unit 106 determines that the participant who uttered the current phrase is not the same as the participant who uttered the past low-frequency phrase, the authority change unit 108 revokes the participation authority of the participant (account) for the remote conversation and makes the participant leave the remote conversation.
[0279] As a result, it is highly convenient because the organizer does not need to take any action against a participant who may be impersonating. In addition, by promptly making a participant who is highly likely to be impersonating leave the remote conversation, the security of the remote conversation can be improved.
[0280] (V) Others The disclosed technology is not limited to the above-described embodiments, and can be variously modified and implemented without departing from the spirit of the present embodiment. Each configuration and each process of the present embodiment can be selected as needed, or may be appropriately combined.
[0281] In each of the above-described embodiments, an example of detecting impersonation in a remote conversation conducted among users (participants) of the participant terminal 2 has been shown, but the present invention is not limited thereto. The user (organizer) of the organizer terminal 3 may participate in the remote conversation. In that case, the organizer also corresponds to a participant.
[0282] Also, in each of the first embodiments, the first behavior extraction unit 102 calculates the appearance frequency of each extraction word included in the determination target phrase in all words, and calculates the frequency average value of the determination target phrase, but the present invention is not limited thereto. For example, the first behavior extraction unit 102 may use tf-idf (term frequency - inverse document frequency).
[0283] In each of the above-described embodiments, the first behavior extraction unit 102 calculates the appearance frequency of extracted words from all the words spoken by the participant being determined in all the videos of the participant being determined. However, the present invention is not limited to this. For example, the first behavior extraction unit 102 may calculate the appearance frequency of extracted words from all the words spoken by all the participants in all the videos of all the participants.
[0284] In each of the above-described embodiments, either the notification unit 107 or the authority change unit 108 is provided. However, the present invention is not limited to this, and both the notification unit 107 and the authority change unit 108 may be provided.
[0285] Also, based on the above disclosure, those skilled in the art can implement and manufacture the present embodiment.
Description of Reference Numerals
[0286] 1 Computer system 2 Participant terminal 3 Organizer terminal 11 Processor (control unit) 12 Memory 13 Storage device 14 Graphics processing unit 14a Monitor 15 Input interface 15a Keyboard 15b Mouse 16 Optical drive device 16a Optical disk 17 Device connection interface 17a Memory device 17b Memory reader / writer 17c Memory card 18 Network interface 19 Bus 20 Network 101 First behavior detection unit 102, 102a First behavior extraction unit 103 Database group 104 Second behavior detection unit 105, 105a Second Behavior Extraction Unit 106, 106a Identity Determination Unit 107 Notification Unit 108 Authority Change Unit 1031 First Phrase Corresponding Text Storage Database 1032 First Face Position Information Storage Database 1033 First Skeleton Position Information Storage Database 1034 First Behavior Database 1035 Second Phrase Corresponding Text Storage Database 1036 Second Face Position Information Storage Database 1037 Second Skeleton Position Information Storage Database 1038 Second Behavior Database
Claims
1. When receiving first sensing data associated with the account of a participant in a remote conversation, obtaining feature information of any one of the actions, voices, and states of the participant, which is extracted from the participant's past second sensing data and has an extraction frequency less than a first reference value; Based on the degree of coincidence between the feature information extracted from the first sensing data and the feature information extracted from the second sensing data, performing a determination regarding impersonation A determination method characterized in that a computer executes the process.
2. The process of making the determination regarding impersonation is Calculating the degree of coincidence for each of a plurality of pairs of the feature information extracted from the first sensing data and the feature information extracted from the second sensing data; Including a process of determining that impersonation has occurred when the number of pairs for which the degree of coincidence is less than a second reference value is less than a third reference value The determination method according to claim 1, characterized in that.
3. The feature information is a phrase spoken by the participant, The process of obtaining the feature information is Including a process of comparing the extraction frequency of the phrase, calculated based on the appearance frequency of each of a plurality of words included in the phrase spoken by the participant, among all the words spoken by the participant in all the videos of the participant, with the first reference value The determination method according to claim 1 or 2, characterized in that.
4. The first sensing data includes a video of the participant taken during an ongoing remote conversation with the participant, The second sensing data includes a video of the participant taken during a past remote conversation with the participant The determination method according to any one of claims 1 to 3, characterized in that.
5. The process of making the determination regarding impersonation is When the number of pairs for which the difference between the degree of coincidence between the second feature information extracted from the first sensing data and having a frequency less than a fourth reference value and the second feature information extracted from the second sensing data, and the degree of coincidence between the first feature information extracted from the first sensing data and having a frequency greater than a fifth reference value and the first feature information extracted from the second sensing data is less than a sixth reference value is greater than or equal to a seventh reference value, including a process of determining that impersonation has occurred The determination method according to any one of claims 1 to 4, characterized in that.
6. When it is determined that the forgery has occurred, including a process of outputting notification information indicating that the forgery has occurred The determination method according to any one of claims 1 to 5, characterized by the above.
7. When it is determined that the forgery has occurred, including a process of depriving the participant of the target of the forgery of the participation right for the remote conversation from the account The determination method according to any one of claims 1 to 6, characterized by the above.
8. When receiving first sensing data associated with the account of a participant in a remote conversation, obtaining characteristic information of any one of the actions, voices, and states of the participant, which is extracted from the participant's past second sensing data and the extraction frequency of which is less than a first reference value, Based on the degree of coincidence between the characteristic information extracted from the first sensing data and the characteristic information extracted from the second sensing data, performing a determination regarding forgery A determination program characterized by causing a computer to execute the process.
9. When receiving first sensing data associated with the account of a participant in a remote conversation, obtaining characteristic information of any one of the actions, voices, and states of the participant, which is extracted from the participant's past second sensing data and the extraction frequency of which is less than a first reference value, Based on the degree of coincidence between the characteristic information extracted from the first sensing data and the characteristic information extracted from the second sensing data, performing a determination regarding forgery An information processing apparatus characterized by including a control unit.
Citation Information
Patent Citations
Specific conversation detection device, method and program
JP2018013529A
Remote dialogue system, remote dialogue method, and remote dialogue program
JP6901190B1
JPP6901190B
Method and apparatus for detecting abnormality of caller
US20200228648A1
Detecting robocalls using biometric voice fingerprints
US20210136200A1