Information processing apparatus, control method, and program
The information processing apparatus enhances speaker identification in online meetings by using a combination of emphasized icons, voiceprint, and face recognition to automate and improve the accuracy of speaker detection.
Patent Information
- Application Number
- JP2023215486
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-03
AI Technical Summary
Existing speaker identification technologies in online meetings are inefficient and prone to errors due to reliance on video data of the mouth and body, and manual correction by administrators is time-consuming.
An information processing apparatus that acquires video data, segments conversations, and uses a combination of emphasized icons, voiceprint, and face image recognition to accurately identify speakers, integrating voice and face determination processing to confirm speaker identities.
Efficient and accurate speaker identification in online meetings, reducing the need for manual correction and improving the reliability of speaker detection.
Smart Images

Figure 2025099098000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, a control method, and a program.
Background Art
[0002] Many current online meeting applications have a function of saving meetings in videos or documents, indicating that it is necessary to review the meeting content.
[0003] To review the meeting content, it is necessary to analyze who said what, and for this purpose, there is a technology for discriminating speakers. As methods for discriminating speakers, there are methods such as comparing the voice obtained from the video with the voiceprint information of the pre-registered attendees, and identifying the speaker from the video data of the movements of the mouth and body of the attendees.
[0004] Patent Document 1 discloses speaker discrimination by the above method. It is a system that identifies a speaker based on the image obtained from the video data and the voice data of each attendee, and outputs the voice data of each attendee as a timeline in the time series of the speech.
[0005] In the video of the meeting, there is a background that the automatic judgment of the speaker cannot be correctly judged only by information such as voice, video data of the mouth and body. In addition, there is a problem that it is time-consuming for the administrator who registers the video to manually correct the speaker.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Disclosure of the Invention
Problems to be Solved by the Invention
[0007] Patent Document 1 describes a function for identifying a speaker based on an image obtained from video data and voice data of each participant.
[0008] However, there is a problem that the automatic determination of the speaker cannot be correctly determined only by video data of the mouth or body and voice information.
[0009] Therefore, an object of the present invention is to efficiently identify a speaker from video data of an online meeting.
Means for Solving the Problem
[0010] The present invention an acquisition means for acquiring an image related to a meeting in which an object corresponding to a participant in an online meeting is displayed; a registration means for specifying a person related to the object among the participants from the display form of the object and registering the person as a speaker; An information processing apparatus comprising:
Effects of the Invention
[0011] According to the present invention, it becomes possible to efficiently identify a speaker from video data of an online meeting.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0014] FIG. 1 is a diagram showing an example of the system configuration of a speaker discrimination system for a video in the present invention. The speaker discrimination system 100 for a video includes a voiceprint registration device 110, a voiceprint DB 120, a face image registration device 130, a face image DB 140, a speaker discrimination device 150, a speaker DB 160, and a video DB 170.
[0015] The voiceprint registration device 110 is a device for registering the voiceprint of a person to be discriminated, and includes a voiceprint reception unit 111 and a voiceprint registration processing unit 112. The user can transmit his / her voiceprint to the voiceprint reception unit 111 through a web browser or the like.
[0016] The voiceprint registration processing unit 112 is a device for associating the voiceprint received by the voiceprint reception unit 111 with the name of the person of that voiceprint and storing it in the voiceprint DB 120.
[0017] The face image registration device 130 is a device for registering the face image of a person to be discriminated, and includes a face image reception unit 131 and a face image registration processing unit 132. The user can transmit his / her face image to the face image reception unit 131 through a web browser or the like.
[0018] The face image registration processing unit 132 is a device that associates the face image received by the face image receiving unit 131 with the name of the person in the face image and stores it in the face image DB 140.
[0019] The speaker discrimination device 150 is a device for discriminating the speaker in a video, and is composed of a video receiving unit 151, an audio and video segmentation processing unit 152, an emphasized icon determination processing unit 153, a voice and face determination processing unit 154, a speaker determination processing unit 155, a speaker registration processing unit 156, and a video registration processing unit 157.
[0020] The video receiving unit 151 is a device for receiving the transmitted video. The video to be received may be a video file that an online meeting application records all or part of an online meeting and stores in a specific database, or may be time-series video data such as real-time video or streaming during an online meeting. In the latter case, the video data may be accumulated for a certain period of time and then processed in the same way as a video file.
[0021] The audio and video segmentation processing unit 152 is a device that segments the video received by the video receiving unit 151 for each conversation.
[0022] The emphasized icon determination processing unit 153 is a device that determines the speaker of the conversation when there is an emphasized icon in the video segmented by the audio and video segmentation processing unit 152 and there is a name below it, or when the face in the icon matches the one registered in the face image DB 140.
[0023] The voice and face determination processing unit 154 is a device that determines the speaker of the conversation when the voice segmented by the audio and video segmentation processing unit 152 matches the one registered in the voiceprint DB 120, or when the face of the person speaking shown in the segmented video matches the one registered in the face image DB 140.
[0024] The speaker determination processing unit 155 is a device that receives the determination from both the emphasized icon determination processing unit 153 and the voice and face determination processing unit 154, or from one of them, and makes a final speaker determination.
[0025] Upon receiving the determination from the speaker determination processing unit 155, the speaker registration processing unit 156 registers the input voice, the determined speaker name, the icon image (including the name under the icon), and the name under the icon in the speaker DB 160.
[0026] After the speaker determination of all conversations is completed, the video registration processing unit 157 registers the video associated with the speaker determination result in the video DB 170.
[0027] FIG. 2 is a block diagram showing an example of the hardware configuration of a speaker discrimination system for videos and a client terminal in an embodiment of the present invention.
[0028] As shown in FIG. 2, an information processing apparatus has a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 connected via a system bus 204.
[0029] The CPU 201 comprehensively controls each device and controller connected to the system bus 204.
[0030] The ROM 202 or the external memory 211 holds a BIOS (Basic Input / Output System), an OS (Operating System), which are control programs executed by the CPU 201, a computer-readable and executable program for realizing this information processing method, and various necessary data (including data tables).
[0031] The RAM 203 functions as the main memory, work area, etc. of the CPU 201. The CPU 201 loads programs and the like necessary for executing processing from the ROM 202 or the external memory 211 into the RAM 203, and realizes various operations by executing the loaded programs.
[0032] The input controller 205 controls the input from input devices such as a keyboard 209 and a pointing device such as a mouse (not shown). When the input device is a touch panel, it is assumed that various instructions can be given by the user pressing (touching with a finger or the like) in accordance with the icons, cursors, and buttons displayed on the touch panel.
[0033] Further, the touch panel may be a touch panel capable of detecting the positions touched by a plurality of fingers, such as a multi-touch screen.
[0034] The video controller 206 controls the display to an external output device such as a display 210. The display is assumed to include the display of a notebook personal computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. Also, for a device capable of receiving the above-described touch operation, an input device is also provided.
[0035] Note that the video controller 206 can control a video memory (VRAM) for performing display control, and can use a part of the RAM 203 as a video memory area, or can separately provide a dedicated video memory.
[0036] The memory controller 207 controls access to an external memory 211. As the external memory, an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a compact flash (registered trademark) memory connected via an adapter to a PCMCIA card slot can be used.
[0037] The communication I / F controller 208 is for connecting and communicating with an external device via a network, and executes communication control processing on the network. For example, communication using TCP / IP, a telephone line such as ISDN, and communication using a 3G line of a mobile phone are possible.
[0038] Also, the CPU 201 enables display on the display 210 by executing an outline font expansion (rasterization) process on, for example, a display information area in the RAM 203. Further, the CPU 201 enables user instructions using a mouse cursor (not shown) or the like on the display 210.
[0039] The speaker discrimination process, which is the main process in the embodiment of the present invention performed by the speaker discrimination device 150 when receiving a video of an online conference, will be described with reference to the flowchart of FIG. 3.
[0040] First, in step S301, the video receiving unit 151 receives a video of an online conference.
[0041] In step S302, the audio and video splitting unit 152 performs audio and video splitting processing for each conversation in the video. Here, splitting is performed based on the break of the audio, the height of the voice, and the change in the voiceprint. From step S303 onwards, sequential processing is performed on the split audio and video.
[0042] In step S303, the emphasized icon determination processing unit 153 performs speaker determination processing using the emphasized icon. The details of this processing flow will be described in FIG. 4.
[0043] In step S304, if there is a determination result in the emphasized icon determination processing of step S303, the process proceeds to step S305. If not, the process proceeds to step S306.
[0044] In step S305, the voice and face determination processing unit 154 performs speaker determination processing using the voice and face video. Even though the speaker is identified in step S303, this additional determination is performed because there is a possibility that the person corresponding to the emphasized icon is different from the person actually speaking. For example, when multiple people participate in an online conference using a single user ID. The details of this processing flow will be described in FIG. 5. After the processing, the process proceeds to step S307.
[0045] In step S306, the voice and face determination processing unit 154 performs speaker determination processing based on voice and face images. This processing is the same as the processing in step S305. However, here, since speaker determination cannot be performed using the emphasis icon, it is determined by another method. After the processing, the process proceeds to step S312.
[0046] In step S307, if there is a determination result of the voice and face determination processing in step S305, the process proceeds to step S308. If not, the process proceeds to step S309.
[0047] In step S308, it is confirmed whether the determination result of the emphasis icon determination processing in step S303 is the same as the determination result of the voice and face determination processing in step S305. If they are the same, the process proceeds to step S311. If they are different, the process proceeds to step S310.
[0048] In step S309, speaker registration is performed using the speaker name of the determination result of the emphasis icon determination processing in step S303. After the processing, the process proceeds to step S315.
[0049] In step S310, speaker registration is performed in the speaker determination table 900 of the speaker DB160 using the speaker name of the determination result of the voice and face determination processing in step S305 (the same applies hereinafter). After the processing, the process proceeds to step S315.
[0050] In step S311, speaker registration is performed using the speaker name of the determination result that matched in step S308 (either the speaker name of the determination result of the emphasis icon determination processing in step S303 or the speaker name of the determination result of the voice and face determination processing in step S305). After the processing, the process proceeds to step S315.
[0051] In step S312, if there is a determination result of the voice and face determination processing in step S306, the process proceeds to step S313. If not, the process proceeds to step S314.
[0052] In step S313, perform speaker registration using the speaker name as the determination result of the voice and face determination process. After the process, proceed to step S315.
[0053] In step S314, perform speaker registration with the speaker name set to "unknown". After the process, proceed to step S315.
[0054] In step S315, check whether all conversations have been speaker-determined. If completed, the main process ends. If there are still remaining conversations, return to step S303 and continue the process.
[0055] Using the flowchart of FIG. 4, the icon-based speaker determination process executed by the emphasized icon determination unit 153 in the embodiment of the present invention when receiving the video segmented by the audio and video segmentation unit will be described.
[0056] First, in step S401, check whether an icon exists in the input video. If an icon exists, proceed to step S402. If not, proceed to step S408.
[0057] In step S402, check whether there is an emphasized icon. In an online meeting application such as Teams (registered trademark), the icon of the person speaking is emphasized with a thick frame. In the example of FIG. 6, the thick frame 601 corresponds. If an emphasized icon exists, proceed to step S403. If not, speaker determination cannot be made from the icon, so proceed to step S408.
[0058] In step S403, check whether a name is displayed around (such as below) the icon. If it is displayed, since it corresponds to the speaker's name, proceed to step S406. In the example of FIG. 6, "Mr. A" (603) corresponds. If not, proceed to step S404.
[0059] In step S404, it is checked whether a face is displayed on the icon. In the example of FIG. 6, the face 602 of person A in the icon 601 corresponds. If no face is displayed on the icon, speaker determination based on the face cannot be performed, so the process proceeds to step S408. If a face is displayed, the process proceeds to step S405.
[0060] In step S405, if the displayed face matches a face registered in the face image DB140, the process proceeds to step S406. If there is no match, speaker determination based on the face cannot be performed, so the process proceeds to step S408.
[0061] In step S406, the determination result is set to yes, and the speaker name is set as the name around the icon (or the name of the person who matched in step S405) confirmed in step S403, and this flow ends.
[0062] In step S407, the determination result is set to yes, and the speaker name is set as the name of the person of the face in the face image DB140 determined to match in step S404, and this flow ends.
[0063] In step S408, since speaker determination cannot be performed, the determination result is no, the speaker name is set to "unknown", and this flow ends.
[0064] The speaker determination process using voice and video executed by the voice and face determination processing unit 154 when receiving the voice and video divided by the voice and video division processing unit in the embodiment of the present invention will be described with reference to the flowchart of FIG. 5.
[0065] First, in step S501, if the input voice matches a voice registered in the voiceprint DB120, in step S503, the determination result is set to yes, and the speaker name is set as the name of the person of that voiceprint. If there is no match, since speaker determination based on the voice cannot be performed, the process proceeds to step S502.
[0066] In step S502, if the face of the person speaking shown in the input video matches a face registered in the face image DB140, the determination result is "yes" in step S504, and the speaker name is set to the name of the person in that face image. If there is no match, the process proceeds to step S505.
[0067] In step S505, since the speaker cannot be determined from the input voice and the face shown in the video, the determination result is "no", the speaker name is set to "unknown", and the determination process ends.
[0068] The voiceprint determination table 700 in FIG. 7 is a table that registers a set of a person's name and voiceprint and is used during voice and face determination processing. It consists of the items of person 701 and voiceprint 702. The voiceprint determination table 700 is stored in the voiceprint DB120.
[0069] Person 701 is an item for specifying the speaker name, and a name is registered.
[0070] Voiceprint 702 is an item for comparing whether the voice in the video matches the voiceprint, and a voiceprint is registered.
[0071] The face determination table 800 in FIG. 8 is a table that registers a set of a person's name and face image and is used during highlighted icon determination processing and face determination processing. It consists of the items of person 801 and face image 802. The face determination table 800 is stored in the face image DB140.
[0072] Person 801 is an item for specifying the speaker name, and a name is registered.
[0073] Face image 802 is an item for comparing whether the face of the person speaking shown in the video matches the face image, and a face image is registered.
[0074] The speaker determination table 900 in FIG. 9 is a table for registering the speaker determination result. It consists of the items of the input voice 901, speaker 902, icon (face image including name) 903, and name extracted from the icon 904.
[0075] The input voice 901 has the voices of the segmented audio for each conversation registered.
[0076] The speaker 902 is an item for registering the speaker name of the determination result of the input voice, and the determined name is registered.
[0077] The icon (face image including name) 903 has an image including the icon and the name under the icon registered.
[0078] The name 904 extracted from the icon has the name extracted from the icon registered.
[0079] Note that in this example, the results of speaker determination are mainly registered, but in order to be available when searching for online meeting videos, the type of meeting, the meeting name, the elapsed time and time from the start of the meeting, etc. may be registered.
[0080] As described above, it becomes possible to efficiently identify speakers from the video data of online meetings.
[0081] As described above, the embodiments according to the present invention have been shown. However, the present invention can take an embodiment as, for example, a system, an apparatus, a method, a program, or a recording medium. Specifically, it may be applied to a system composed of a plurality of devices, or may be applied to an apparatus composed of a single device.
[0082] Also, the program in the present invention is a program that a computer can execute the processing method of the flowchart shown in each figure, and the storage medium of the present invention stores a program that a computer can execute the processing method of each figure. Note that the program in the present invention may be a program for each processing method of each device in each figure.
[0083] As described above, it goes without saying that the object of the present invention can also be achieved by supplying a recording medium storing a program for realizing the functions of the above-described embodiments to a system or an apparatus, and causing a computer (or a CPU or MPU) of the system or apparatus to read and execute the program stored in the recording medium.
[0084] In this case, the program itself read from the recording medium realizes the novel functions of the present invention, and the recording medium storing the program constitutes the present invention.
[0085] As the recording medium for supplying the program, for example, a flexible disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a DVD-ROM, a magnetic tape, a non-volatile memory card, a ROM, a non-volatile memory chip, a silicon disk, etc. can be used.
[0086] Further, by executing the program read by the computer, not only the functions of the above-described embodiments are realized, but also based on the instructions of the program, an OS (operating system) or the like operating on the computer performs part or all of the actual processing, and the functions of the above-described embodiments are realized by the processing. This goes without saying.
[0087] Furthermore, after the program read from the recording medium is written into a memory provided in a function expansion board inserted into the computer or a function expansion unit connected to the computer, based on the instructions of the program code, a CPU or the like provided in the function expansion board or the function expansion unit performs part or all of the actual processing, and the functions of the above-described embodiments are realized by the processing. This goes without saying.
[0088] Furthermore, the present invention may be applied to a system composed of a plurality of devices or to an apparatus composed of a single device. Needless to say, the present invention is also applicable when it is achieved by supplying a program to a system or an apparatus. In this case, by reading out a recording medium storing a program for achieving the present invention to the system or the apparatus, the system or the apparatus can enjoy the effects of the present invention.
[0089] Furthermore, by downloading and reading out a program for achieving the present invention from a server, a database, etc. on a network by means of a communication program, the system or the apparatus can enjoy the effects of the present invention. Note that all configurations combining the above-described respective embodiments and their modified examples are also included in the present invention.
Explanation of Reference Numerals
[0090] 100 Speaker discrimination system for video
Claims
1. An acquisition means for acquiring an image related to a meeting in which an object corresponding to a participant in an online meeting is displayed; A registration means for identifying, from the display form of the object, a person among the participants related to the object and registering the person as a speaker; An information processing apparatus characterized by comprising the above.
2. The acquisition means acquires the video of the online meeting; The registration means identifies a person speaking from the video and registers the person as a speaker. The information processing apparatus according to claim 1, characterized by the above.
3. The registration means determines which of the person identified from the display form of the object and the person identified from the video is to be registered as the speaker. The information processing apparatus according to claim 2, characterized by the above.
4. The acquisition means acquires the audio of the online meeting; The registration means identifies a person speaking from the audio and registers the person as a speaker. The information processing apparatus according to claim 1, characterized by the above.
5. The registration means determines which of the person identified from the display form of the object and the person identified from the audio is to be registered as the speaker. The information processing apparatus according to claim 4, characterized by the above.
6. An acquisition step of acquiring an image related to a meeting in which an object corresponding to a participant in an online meeting is displayed; A registration step of identifying, from the display form of the object, a person among the participants related to the object and registering the person as a speaker; A control method for an information processing apparatus, characterized by comprising the above.
7. A program for causing an information processing apparatus to function as: An acquisition means for acquiring an image related to a meeting in which an object corresponding to a participant in an online meeting is displayed; A registration means for identifying, from the display form of the object, a person among the participants related to the object and registering the person as a speaker.
Citation Information
Patent Citations
Conference support system and conference support program
JP2019061594A