Information processing apparatus, information processing program, information processing system, and information processing method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2026-04-21
AI Technical Summary
Conventional conference systems fail to identify participants who may not understand or be interested in the speaker's content, making it difficult for speakers to prioritize their attention effectively.
An information processing system that includes video analysis units to estimate facial expressions and movements of participants, allowing speakers to be notified of individuals who may be struggling to follow or show disinterest, enabling dynamic display adjustments based on these detections.
Speakers can identify and prioritize their interactions with participants who may be having trouble understanding or showing disengagement, enhancing engagement and participation in remote conferences.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing program, an information processing system, and an information processing method. [Background technology]
[0002] Conventionally, there has been known a conference system for holding a remote conference among a plurality of remote locations, which includes a technique for setting information indicating the level of a participant's desire to speak in the conference video when a predetermined movement that is assumed to indicate a desire to speak is detected from video data of the participant. Summary of the Invention [Problem to be solved by the invention]
[0003] The conventional technology described above does not allow a speaker to find people who should be given importance, such as participants who do not understand what the speaker is saying or participants who are not interested in what the speaker is saying.
[0004] The disclosed technology aims to notify a speaker of a party that should be given priority. [Means for solving the problem]
[0005] The disclosed technology includes an overall processing unit that receives setting contents indicating the facial expression of the detection target, and a network processing unit that notifies other information processing devices of the setting contents, and the overall processing unit receives a notification indicating that the facial expression of the detection target has been detected from image data acquired by the other information processing devices, and outputs the notification to a display unit. [Effects of the Invention]
[0006] The speaker can be notified of the person to whom he or she should attach importance. [Brief explanation of the drawings]
[0007] [Figure 1]FIG. 1 illustrates an example of a system configuration of an information processing system according to a first embodiment. [Figure 2] FIG. 2 illustrates an example of a hardware configuration of a server according to the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a communication terminal. [Figure 4] FIG. 2 is a diagram illustrating functions of a communication terminal according to the first embodiment. [Figure 5] 10 is a flowchart illustrating an operation of the communication terminal according to the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of a display layout of the communication terminal according to the first embodiment. [Figure 7] FIG. 10 is a diagram illustrating another example of a display layout of the communication terminal according to the first embodiment. [Figure 8] FIG. 2 is a diagram illustrating a hardware configuration of an electronic whiteboard. [Figure 9] FIG. 1 is a diagram illustrating an example of a hardware configuration of a smartphone. [Figure 10] FIG. 10 is a diagram illustrating a system configuration of an information processing system according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] (First embodiment) The first embodiment will be described below with reference to the drawings: Fig. 1 is a diagram showing an example of the system configuration of an information processing system according to the first embodiment.
[0009] The information processing system 100 of this embodiment includes a server 200 and a plurality of communication terminals 300-1, 300-2, ..., 300-N. In the information processing system 100, the server 200 and the communication terminals 300-1, 300-2, ..., 300-N are connected via a network N such as the Internet, an intranet, or a LAN (Local Area Network). In the following description, when there is no need to distinguish between the communication terminals 300-1, 300-2, ..., 300-N, they will be referred to as communication terminals 300. The communication terminal 300 of this embodiment is an example of an information processing device having a CPU (Central Processing Unit) and a memory.
[0010] In the information processing system 100 of this embodiment, a so-called video conference is held between a plurality of locations using these multiple devices.
[0011] The server 200 of this embodiment performs various controls related to a video conference held by the communication terminals 300. For example, at the start of a video conference, the server 200 monitors the communication connection status between each communication terminal 300 and the server 200, calls each communication terminal 300, etc. During the video conference, the server 200 also performs transfer processing of various data (e.g., video data, audio data, drawing data, etc.) between the multiple communication terminals 300.
[0012] The communication terminal 300 of this embodiment is an example of a video processing device or an imaging device. The communication terminal 300 is installed at each location where a video conference is held and used by participants in the video conference. For example, the communication terminal 300 transmits various data (e.g., video data, audio data, drawing data, etc.) input during the video conference to other communication terminals 300 via the network N and the server 200.
[0013] Furthermore, for example, the communication terminal 300 outputs various data received from other communication terminals 300 using an output method (for example, display, audio output, etc.) according to the type of data, thereby presenting the data to participants in the video conference.
[0014] In addition, each of the communication terminals 300-1, 300-2, ..., 300-N of this embodiment has video analysis units 310-1, 310-2, ..., 310-N and video generation units 320-1, 320-2, ..., 320-N as functional units that realize the main processing of this embodiment.
[0015] In this embodiment, the video analysis units 310-1, 310-2, ..., 310-N each realize the same function. In the following description, when there is no need to distinguish between the video analysis units 310-1, 310-2, ..., 310-N, they will be referred to as video analysis unit 310. In addition, in this embodiment, the video generation units 320-1, 320-2, ..., 320-N each realize the same function. In the following description, when there is no need to distinguish between the video generation units 320-1, 320-2, ..., 320-N, they will be referred to as video generation unit 320.
[0016] In the communication terminal 300 of this embodiment, the video analysis unit 310 analyzes image data acquired from the imaging unit of each of the multiple communication terminals 300, and estimates the facial expression of a person from a face image of the person recognized from the image data. The person recognized from the image data is a participant in a conference at the location where the communication terminal 300 is installed.
[0017] Furthermore, the video analysis unit 310 receives a request to detect a specific facial expression from another communication terminal 300. Here, the other communication terminal 300 refers to a communication terminal 300 installed at the location where the participant is speaking. In the following description, the communication terminal 300 installed at the location where the participant is speaking may be referred to as the speaker terminal 300.
[0018] Then, when a specific facial expression is estimated from the image data, the video analysis unit 310 transmits information indicating that the specific facial expression has been estimated, together with the image data acquired by the imaging unit, to the speaker terminal 300. The specific facial expression in this embodiment may be, for example, an anxious expression, an indifferent expression, a sleepy expression, etc.
[0019] In this embodiment, when the video generation unit 320 is the speaker terminal 300 and receives image data from another communication terminal 300 together with information indicating that a specific facial expression has been estimated, the video generation unit 320 generates image data including this image data and information notifying the facial expression of the participant at the location from which the image data was sent, and displays the image data on the display unit.
[0020] In this way, when the communication terminal 300 of this embodiment is not the speaker terminal 300, it estimates the facial expression of the participant, and when it estimates the facial expression that the speaker terminal 300 has requested to be detected, it transmits the estimation result to the speaker terminal 300.
[0021] In addition, when the communication terminal 300 of this embodiment is the speaker terminal 300, it receives a notification from another communication terminal 300 indicating that a specific facial expression has been estimated from a participant, and outputs image data of the participant and a notification indicating that it is the specific facial expression.
[0022] Therefore, according to this embodiment, it is possible to find listeners that a speaker should pay attention to, such as participants who are having trouble understanding what the speaker is saying, or participants who the speaker wants to get interested in what is being said, and notify the speaker. Also, in this embodiment, by displaying on the screen the listeners that the speaker should pay attention to, the speaker can understand the state of these listeners.
[0023] The hardware configuration of each device included in the information processing system 100 of this embodiment will be described below.
[0024] 2 is a diagram showing an example of the hardware configuration of a server according to the first embodiment. The server 200 according to this embodiment is constructed by a computer, and includes a CPU 231, a ROM 232, a RAM 233, an HD 234, an HDD (Hard Disk Drive) controller 235, a display 236, an external device connection I / F (Interface) 238, a network I / F 239, a data bus B, a keyboard 241, a pointing device 242, a DVD-RW (Digital Versatile Disk Rewritable) drive 244, and a media I / F 246.
[0025] Of these, the CPU 231 controls the overall operation of the server 5. The ROM 232 stores programs used to drive the CPU 231, such as an IPL (Initial Program Loader). The RAM 233 is used as a work area for the CPU 231. The HD 234 stores various data such as programs. The HDD controller 235 controls the reading and writing of various data from and to the HD 234 under the control of the CPU 231. The display 236 displays various information such as a cursor, menu, window, text, or image.
[0026] The external device connection I / F 238 is an interface for connecting various external devices. In this case, the external devices are, for example, a USB (Universal Serial Bus) memory or a printer. The network I / F 239 is an interface for data communication using the network N. The bus line B is an address bus, a data bus, or the like for electrically connecting the components such as the CPU 231 shown in FIG. 3.
[0027] The keyboard 241 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 242 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The DVD-RW drive 244 controls reading and writing of various data from a DVD-RW 243, which is an example of a removable recording medium. Note that this is not limited to DVD-RW, and may be DVD-R, etc. The media I / F 246 controls reading and writing (storing) of data from a recording medium 245, such as a flash memory.
[0028] Fig. 3 is a diagram showing an example of the hardware configuration of a communication terminal 300. Fig. 3 shows the hardware configuration of communication terminal 300 when communication terminal 300 is an example of a video conference terminal.
[0029] The video conference terminal 7 is an example of the communication terminal 300, and the communication terminal 300 is not limited to the video conference terminal 7. Other examples of the communication terminal 300 will be described later.
[0030] The video conference terminal 7 includes a CPU 701, a ROM 702, a RAM 703, a flash memory 704, an SSD 705, a media I / F 707, an operation button 708, a power switch 709, a bus line 710, a network I / F 711, a CMOS (Complementary Metal Oxide Semiconductor) sensor 712, an image sensor I / F 713, a microphone 714, a speaker 715, an audio input / output I / F 716, a display I / F 717, an external device connection I / F (Interface) 718, a short-range communication circuit 719, and an antenna 719a of the short-range communication circuit 719.
[0031] Of these, the CPU 701 controls the overall operation of the videoconference terminal 7. The ROM 702 stores programs used to drive the CPU 701, such as an IPL. The RAM 703 is used as a work area for the CPU 701. The flash memory 704 stores various data, such as communication programs, image data, and sound data. The flash memory 704 may be a flash memory mounted inside the SSD 705.
[0032] The SSD 705 controls reading and writing of various data from and to the flash memory 704 under the control of the CPU 701. Note that an HDD may be used instead of an SSD. The media I / F 707 controls reading and writing (storing) of data from and to a recording medium 706 such as a flash memory. The operation button 708 is a button that is operated when selecting a destination for the video conference terminal 7, for example. The power switch 709 is a switch for turning the power of the video conference terminal 7 ON / OFF.
[0033] Furthermore, the network I / F 711 is an interface for data communication using a network N such as the Internet. The CMOS sensor 712 is a type of built-in imaging means that captures an image of a subject and obtains image data under the control of the CPU 701. Note that instead of a CMOS sensor, an imaging means such as a CCD (Charge Coupled Device) sensor may also be used.
[0034] The image sensor I / F 713 is a circuit that controls the driving of the CMOS sensor 712. The microphone 714 is a built-in circuit that converts sound into an electrical signal. The speaker 715 is a built-in circuit that converts the electrical signal into physical vibrations to produce sound such as music or voice. The sound input / output I / F 716 is a circuit that processes the input and output of sound signals between the microphone 714 and the speaker 715 under the control of the CPU 701.
[0035] The display I / F 717 is a circuit that transmits image data to an external display under the control of the CPU 701. The external device connection I / F 718 is an interface for connecting various external devices. The short-range communication circuit 719 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark).
[0036] 3. The bus line 710 is an address bus, a data bus, or the like for electrically connecting the components such as the CPU 701 shown in FIG.
[0037] The display connected to the display I / F 717 is a type of display means configured with a liquid crystal display (LCD) or organic EL (Electro Luminescence) display that displays an image of a subject, operation icons, etc. The display is connected to the display I / F 717 via a cable. This cable may be a cable for analog RGB (VGA) signals, a cable for component video, or a cable for HDMI (High-Definition Multimedia Interface) (registered trademark) or DVI (Digital Video Interactive) signals.
[0038] The CMOS (Complementary Metal Oxide Semiconductor) sensor 712 is a type of built-in imaging means that captures an image of a subject and obtains image data under the control of the CPU 701. Instead of a CMOS sensor, an imaging means such as a CCD (Charge Coupled Device) sensor may be used. External devices such as an external camera, an external microphone, and an external speaker can be connected to the external device connection I / F 718 via a USB cable or the like.
[0039] When an external camera is connected, the external camera is driven under the control of the CPU 701, prioritizing the built-in CMOS sensor 712. Similarly, when an external microphone or external speaker is connected, the external microphone or external speaker is driven under the control of the CPU 701, prioritizing the built-in microphone 714 or built-in speaker 715, respectively.
[0040] The recording medium 706 is detachable from the video conference terminal 7. The storage medium 706 is not limited to the flash memory 704, and may be an EEPROM or the like, as long as it is a non-volatile memory that reads or writes data under the control of the CPU 701.
[0041] Next, functions of the communication terminal 300 of this embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram for explaining functions of the communication terminal of the first embodiment.
[0042] The communication terminal 300 of this embodiment includes a video analysis unit 310, a video generation unit 320, a video editing unit 330, an audio processing unit 340, an overall processing unit 350, an imaging unit 361, a sound collection unit 362, an audio output unit 363, a network processing unit 364, a codec unit 365, an operation unit 366, and a recording unit 367. Each of the above-mentioned units is realized by the CPU 701 reading and executing a program stored in the ROM 702 or the like. The communication terminal 300 of this embodiment also has a storage unit 368. The storage unit 368 is, for example, a storage area provided in a RAM or the like.
[0043] The video analysis unit 310 recognizes face images contained in the image data and estimates facial expressions. Details of the video analysis unit 310 will be described later.
[0044] In this embodiment, the image includes a still image and a video, and the image data includes a still image data and a video data. In this embodiment, in the information processing system 100, the image data captured by the imaging unit 361 during a video conference is video data. In the following description, data in which video data and audio data are synchronized may be referred to as video data.
[0045] The video generation unit 320 generates image data according to the processing result of the video analysis unit 310. The video editing unit 330 takes in, via the network processing unit 364, video data transferred from other communication terminals 300 installed at other locations participating in the video conference, and combines it with the image data generated by the video generation unit 320 to display on the display unit 370.
[0046] The display unit 370 may be, for example, a monitor device connected to the communication terminal 300. The display unit 370 may also be included in the communication terminal 300.
[0047] When the audio processing unit 340 acquires audio data received via the network processing unit 364, it performs processing that is generally used in audio data processing, such as codec processing and noise cancellation, and transfers the data to the audio output unit 363. In addition, the audio processing unit 340 performs echo cancellation (EC) processing on the audio data that is input to the sound collection unit 362.
[0048] Furthermore, the audio processing unit 340 of this embodiment has a speaker tracking detection unit 341. The speaker tracking detection unit 341 detects and tracks a speaker based on audio data collected by the audio collection unit 362 and a face image of the person detected by the video analysis unit 310. While tracking the speaker, the speaker tracking detection unit 341 of this embodiment may transmit information identifying the speaker to the communication terminal 300 at another location via the network processing unit 364.
[0049] The overall processing unit 350 is responsible for overall control of the communication terminal 300. The overall processing unit 350 also performs mode setting and status management for each module and block according to instructions from the conference participants and the like.
[0050] Specifically, when audio data is input from the sound collection unit 362 to the audio processing unit 340, the overall processing unit 350 determines that the device itself has become the speaker terminal 300.
[0051] Furthermore, when the overall processing unit 350 becomes the speaker terminal 300, for example, it accepts the settings of the facial expressions of the participants that are requested to be detected at the other communication terminals 300, and when the overall processing unit 350 becomes the speaker terminal 300, it requests the other communication terminals 300 to detect the set facial expressions.
[0052] In addition, the overall processing unit 350 of this embodiment performs layout settings and instructions related to the display on the display unit 370 on the video generation unit 320, and generates and selects messages to be sent to other communication terminals 300 according to the screen layout control situation.
[0053] Specifically, when the overall processing unit 350 receives a notification from another communication terminal 300 indicating that the requested facial expression has been detected, it controls the layout so that the image data sent from the sender of the notification and the notification are displayed on the display unit 370.
[0054] The imaging unit 361 is a camera module, and acquires image data of an image captured by the CMOS sensor 712, the imaging element I / F, etc. The imaging unit 361 inputs image data (video data) of a conference scene. The imaging unit 361 includes, for example, a lens, an image sensor that converts an image collected through the lens into an electrical signal, and a DSP (digital signal processor) that applies various known processes to the RAW data transferred from the image sensor to generate YUV data.
[0055] The sound collection unit 362 acquires audio data of the audio input to the microphone. When the sound collection unit 362 collects the audio data of the speaker in the conference, it converts the collected audio data into digital data and transfers it to the audio processing unit 340. Note that the sound collection unit 362 may also be configured to collect audio from a plurality of microphones in an array format.
[0056] The audio output unit 363 converts audio data received from another communication terminal 300 installed at another base into an analog signal and outputs it to a speaker.
[0057] For image data to be transmitted, the network processing unit 364 transfers the coded data transferred from the codec unit 365 to the destination communication terminal 300 via the network.
[0058] Furthermore, the network processing unit 364 acquires encoded data transferred from other communication terminals 300 via the network and transfers the encoded data to the codec unit 365. The network processing unit 364 may have a function of monitoring the network bandwidth in order to determine encoding parameters (such as a QP value). The communication terminal 300 may also have a function of inputting information about the functions and performance of other communication terminals 300 in order to optimize the settings of encoding parameters and transmission modes.
[0059] The codec unit 365 is realized by a codec circuit or software for encoding / decoding image data to be transmitted and received.
[0060] For image data to be transmitted, the codec unit 365 performs encoding processing on the image data input from the video analysis unit 310 and transfers the encoded image data to the network processing unit 364. For image data to be received, the codec unit 365 receives encoded image data from another communication terminal 300 via the network processing unit 364, decodes the encoded image data, and transfers it to the video generation unit 320.
[0061] The operation unit 366 accepts pan / tilt operations by conference participants, etc. The operation unit 366 also performs various settings, calls to conference participants, and other operations.
[0062] The recording unit 367 acquires audio data and video data during the conference from the video generation unit 320 and audio processing unit 340, and records video of the conference scene. In this embodiment, the recording data is output to the audio processing unit 340 and video generation unit 320, and the conference scene can be played back.
[0063] The storage unit 368 is realized by, for example, a RAM or the like, and temporarily stores the processing results of the video analysis unit 310.
[0064] Next, a further description will be given of the video analysis unit 310 of this embodiment. The video analysis unit 310 of this embodiment has a face detection unit 311, a movement detection unit 312, a facial expression estimation unit 313, and a determination unit 314.
[0065] The face detection unit 311 of this embodiment detects a person's face from image data (video data) captured by the imaging unit 361. The face detection unit 311 also provides the movement detection unit 312 with information indicating the position of the area where the person's face is detected.
[0066] The movement detection unit 312 acquires image data of a person and analyzes the movement based on the position information provided by the face detection unit 311. Specifically, the movement detection unit 312 detects movements such as raising a hand, nodding, looking at or not looking at a monitor (display unit 370), sleeping, etc., and stores the detection results in the storage unit 368.
[0067] The facial expression estimation unit 313 estimates the facial expression of the person based on the acquired image data and stores the estimation result in the storage unit 368. Specifically, the facial expression estimation unit 313 may estimate facial expressions such as joy, surprise, anger, sadness, and anxiety from changes in the facial image of the person, for example.
[0068] The determination unit 314 refers to the memory unit 368 and determines whether the estimation result by the facial expression estimation unit 313 or the detection result by the action detection unit 312 is the facial expression requested by the speaker terminal 300, and notifies the overall processing unit 350 of the determination result.
[0069] Specifically, for example, if the facial expression requested by speaker terminal 300 is "anxiety," determination unit 314 determines whether or not the facial expression estimated by facial expression estimation unit 313 is "anxiety." If the estimated facial expression is "anxiety," determination unit 314 notifies overall processing unit 350 that the requested facial expression has been detected.
[0070] Next, the operation of the communication terminal 300 of this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart illustrating the operation of the communication terminal of the first embodiment.
[0071] In the communication terminal 300 of this embodiment, the overall processing unit 350 performs initial settings for each module at startup, and sets the imaging unit 361 to a state where it is possible to capture an image (step S501).
[0072] Next, the communication terminal 300 acquires information about the participants participating in the remote conference (step S502).
[0073] Specifically, the communication terminal 300 may have a face authentication function. In this case, the communication terminal 300 may acquire image data in which the participant's name is added to the face image of the participant from a communication terminal 300 installed at another base.
[0074] In this embodiment, the server 200 may perform facial authentication of the participants in the remote conference, and the server 200 may distribute image data in which the participant's name is added to the facial image of the participant to the communication terminal 300 at each location.
[0075] Furthermore, if the communication terminal 300 or the server 200 does not have a face authentication function, participants in the remote conference at each location may enter their own names as participant information and send it to the communication terminals 300 at other locations.
[0076] Next, communication terminal 300 starts the remote conference, initializes the timer, and starts counting (step S503). The count value of the timer indicates the time since the screen layout of display unit 370 was updated (changed). Also, for example, the timer may be included as part of the functions of overall processing unit 350.
[0077] Next, communication terminal 300 determines whether to change the settings related to the display of display unit 370 (step S504). It is assumed that the display layout of display unit 370 immediately after the start of the remote conference remains the default setting or the display layout last set is maintained.
[0078] If the setting is not to be changed in step S504, communication terminal 300 proceeds to step S507, which will be described later.
[0079] In step S504, if the settings regarding the display layout of display unit 370 are to be changed, communication terminal 300 causes operation unit 366 to display a screen for accepting changes to the settings, and overall processing unit 350 performs the accepted settings (step S505).
[0080] The settings related to the display layout in this embodiment include information indicating the facial expression to be detected (the facial expression of the detection target) and the movement to be detected.
[0081] The setting contents also include information indicating the number of locations whose screens are to be displayed on the display unit 370, the size (number of pixels) of the image for each location, and the layout of the image for each location. The setting contents also include information indicating how to assign priority to each location.
[0082] Specifically, for example, the following setting contents can be considered: Example 1) The location where a participant is currently speaking is given the highest priority, followed by the location where the participant who spoke previously, and priorities are assigned to locations based on the order in which previous participants spoke. Example 2) The location where a participant is speaking is given the highest priority, followed by the location where the participant has been speaking for the longest time. Example 3) The facial expression estimation unit 313 of the video analysis unit 310 extracts locations where the participant's facial expression is estimated to be "anxious" and assigns a priority to those locations. In this case, for example, the frequency with which the participant's facial expression is estimated to be "anxious" for each location is stored as log information in the storage unit 368, and the locations are assigned a priority in descending order of frequency. Example 4) The location where a participant is speaking is given the highest priority, and the location where the participant's facial expression is estimated to be "anxious" is given the next highest priority.
[0083] The settings regarding the display layout are not limited to the above example, and may be arbitrarily set by the user (participant) of the communication terminal 300 for each location.
[0084] Next, the communication terminal 300 notifies the other communication terminals 300 at each base of the setting contents via the network processing unit 364, and initializes the timer again to start counting (step S506).
[0085] In communication terminal 300, for example, it is assumed that the setting contents regarding the display layout are the contents shown in Example 1. In this case, the other communication terminal 300 refers to memory unit 368 by its own determination unit 314 and determines whether the detection result of action detection unit 312 is “utterance.”
[0086] If the detection result of the action detection unit 312 is "utterance," the other communication terminal 300 transmits, to the communication terminal 300, information indicating that "utterance" has been detected, together with image data of the participant.
[0087] Also, assume that the settings of Example 3 are made as the settings related to the display layout in the communication terminal 300. In this case, the communication terminal 300 notifies other communication terminals 300 installed at other bases that the participant's facial expression of "anxiety" is information that should be detected.
[0088] The other communication terminal 300 that has received this notification refers to the storage unit 368 by the determination unit 314 of its own device, and determines whether the estimation result of the facial expression estimation unit 313 is "anxious" or not.
[0089] If the estimation result is "anxious," the other communication terminal 300 transmits to the communication terminal 300 information indicating that an "anxious" facial expression has been detected, along with image data of the participant.
[0090] Next, communication terminal 300 determines whether or not time Tm has elapsed based on the count value of the timer (step S507). If time Tm has not elapsed in step S507, communication terminal 300 proceeds to step S513, which will be described later.
[0091] If the time Tm has elapsed in step S507, the communication terminal 300 determines whether to change the layout (step S508).
[0092] If it is determined in step S508 that the layout is to be changed, communication terminal 300 changes the display layout of display unit 370 in accordance with the setting (step S509), and proceeds to step S512, which will be described later.
[0093] The processes in steps S508 and S509 will be described below.
[0094] In step S508, the communication terminal 300 of this embodiment determines whether or not information that should be detected has been detected at each base.
[0095] For example, when the setting of Example 1 is made as the setting content regarding the display layout, the communication terminal 300 determines whether or not it has received, from each location, information indicating that the action of "speaking" has been detected along with image data.
[0096] Specifically, the communication terminal 300 determines that a participant at a location where the action of "speaking" has been detected a predetermined number of times or more within a predetermined time period is "speaking." The overall processing unit 350 then assigns the highest priority to the location where the participant has been determined to be "speaking," and changes the display layout of the display unit 370 so that the image data transmitted from this location is displayed largest.
[0097] Furthermore, for example, when the setting of Example 4 is made as the setting content for the display layout, the communication terminal 300 determines whether there is a location that has sent information indicating that the action of "speaking" has been detected together with the image data, and whether there is a location that has sent information indicating that an "anxious" expression has been detected together with the image data.
[0098] Specifically, the communication terminal 300 determines that a participant at a location where the action of "speaking" is detected a predetermined number of times or more within a predetermined time is "speaking."
[0099] Furthermore, for a location where an "anxious" expression is detected a predetermined number of times or more within a predetermined period of time, communication terminal 300 determines that the participant at this location has an "anxious" expression.
[0100] When there is a location where a participant is speaking and another location where a participant has an anxious expression, the communication terminal 300 assigns the highest priority to the location where the participant is speaking and the second highest priority to the location where the participant has an anxious expression.
[0101] Then, communication terminal 300 changes the display layout of display unit 370 so that the image data to be transmitted is displayed larger in descending order of priority.
[0102] The communication terminal 300 may transmit a message corresponding to the display layout to the communication terminal 300 at the location to which the priority has been assigned, for example, by the overall processing unit 350.
[0103] If it is determined in step S508 that the layout is not to be changed, communication terminal 300 determines whether the display layout of display unit 370 is in the default state (step S510).
[0104] If the display layout is in the default state in step S510, communication terminal 300 proceeds to step S513, which will be described later. If the display layout is not in the default state in step S510, communication terminal 300 returns the display layout to the default state (step S511).
[0105] Next, the communication terminal 300 resets the timer and starts counting again (step S512). Next, the communication terminal 300 determines whether the remote conference is continuing (step S513). Specifically, the communication terminal 300 determines whether an instruction to end the remote conference has been received.
[0106] If the remote conference is continuing in step S513, the communication terminal 300 returns to step S504. If the remote conference is ending in step S513, the communication terminal 300 ends the process.
[0107] As described above, in this embodiment, the layout of the display unit can be changed according to the settings including the facial expressions of the participants. Also, in this embodiment, by setting a timer to count the time Tm, the display layout is frequently changed according to the behavior of the participants at each location, which prevents the participants from feeling uncomfortable.
[0108] Next, the display layout of the communication terminal 300 of this embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the display layout of the communication terminal of the first embodiment.
[0109] In the example of FIG. 6, communication terminals 300-A, 300-B, 300-C, and 300-D are set up at locations A to D, respectively, and a remote conference is being held.
[0110] In the example of FIG. 6, the participant at location A is the speaker, and the cumulative speaking time up to now is the participants at location A, location C, location B, and location D.
[0111] In the example of Figure 6, the communication terminal 300-A at location A has a display layout setting that gives the highest priority to the location where the participant's facial expression is estimated to be ``anxious,'' and is set to display images of the two locations.
[0112] At locations B, C, and D, the display layout is set so that the location where a participant is speaking has the highest priority and images from the two locations are displayed.
[0113] In this case, the communication terminal 300-A at the site A requests the communication terminals 300-B, 300-C, and 300-D at the sites B to D to notify that it has detected that the participant's facial expression shows "anxiety."
[0114] Then, when communication terminal 300-A at point A receives, from communication terminal 300-D at point D, image data and information indicating that it has detected an "anxious" expression from a participant a predetermined number of times or more within a predetermined time, it changes the display layout of display unit 370A as shown in Figure 6.
[0115] Specifically, the communication terminal 300-A displays an image 371 of the participant at the location D on the display unit 370A.
[0116] Furthermore, the communication terminals 300 at the locations B, C, and D request the communication terminals 300 at the other locations to notify them that they have detected a participant's "speech." Therefore, the communication terminals 300 at the locations B, C, and D display an image of the participant at location A, and then preferentially display images of the location with the longest cumulative speaking time.
[0117] In this way, in this embodiment, even if a participant is not speaking or does not have a desire to speak during a remote conference, the image of that participant can be preferentially displayed on the display unit 370.
[0118] In other words, in this embodiment, in a remote conference, participants who are not actively participating in the conversation or who do not appear to understand the content of the conversation can be detected from the participants' facial expressions and notified to the speaker.
[0119] Fig. 7 is a diagram showing another example of the display layout of the communication terminal of the first embodiment. In the example of Fig. 7, it is assumed that a participant at location A is speaking. In addition, the example of Fig. 7 shows a case where the display layout settings of communication terminal 300-A are set to assign priority in order of cumulative speaking time, and to notify a location where a participant's facial expression is detected as "anxious" if such a location exists.
[0120] In this case, the display unit 370A at base A displays an image of the participant at base B and an image of the participant at base C in descending order of cumulative speaking time. Also, the display unit 370A displays a message 372A indicating that the facial expression of the participant at base residence D is estimated to be "anxious."
[0121] In this embodiment, the presence of a location where the participant's facial expression is estimated to be "anxious" can be notified to the participant at location A who is currently speaking. This allows the participant at location A to ask the participant at location D if they have any questions or if they have any opinions about the content of the discussion, for example, thereby livening up the meeting.
[0122] In the present embodiment, the communication terminal 300 has been described as the video conference terminal 7, but is not limited to this. The communication terminal 300 may be, for example, an electronic whiteboard or a smartphone.
[0123] If the communication terminal 300 is an electronic whiteboard (Interactive Whiteboard: a whiteboard with electronic whiteboard functions that allow mutual communication) or a smartphone, the communication terminal 300 will include a display unit 370 (see Figure 4).
[0124] The following describes the hardware configuration of an electronic whiteboard, which is an example of the communication terminal 300. Fig. 8 is a diagram illustrating the hardware configuration of the electronic whiteboard.
[0125] The electronic whiteboard 2 includes a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an SSD (Solid State Drive) 204, a network I / F 205, and an external device connection I / F (Interface) 206.
[0126] Of these, the CPU 201 controls the overall operation of the electronic whiteboard 2. The ROM 202 stores programs used to drive the CPU 201, such as the CPU 201 and an IPL (Initial Program Loader). The RAM 203 is used as a work area for the CPU 201. The SSD 204 stores various data, such as programs for the electronic whiteboard. The network I / F 205 controls communication with the network N. The external device connection I / F 206 is an interface for connecting various external devices. In this case, the external devices are, for example, a USB (Universal Serial Bus) memory 230 and external devices (a microphone 240, a speaker 250, and a camera 260).
[0127] The electronic whiteboard 2 also includes a capture device 211, a GPU 212, a display controller 213, a contact sensor 214, a sensor controller 215, an electronic pen controller 216, a short-range communication circuit 219, an antenna 219a of the short-range communication circuit 219, a power switch 222, and selection switches 223.
[0128] Of these, the capture device 211 displays video information as still images or moving images on the display of an external PC (Personal Computer) 270. The GPU (Graphics Processing Unit) 212 is a semiconductor chip that specializes in graphics. The display controller 213 controls and manages the screen display to output the output image from the GPU 212 to a display 280 or the like.
[0129] The contact sensor 214 detects contact of the electronic pen 290, the user's hand H, or the like on the display 280. The sensor controller 215 controls the processing of the contact sensor 214. The contact sensor 214 inputs and detects coordinates using an infrared blocking method.
[0130] The method for inputting and detecting these coordinates is as follows: two light receiving and emitting devices installed at both ends of the upper side of the display 280 emit multiple infrared rays parallel to the display 280, and receive the light that is reflected by a reflecting member installed around the display 280 and returns along the same optical path as the light emitted by the light receiving element.
[0131] The contact sensor 214 outputs the ID of the infrared light emitted by the two light receiving and emitting devices that has been blocked by an object to the sensor controller 215, and the sensor controller 215 identifies the coordinate position of the contact position of the object. The electronic pen controller 216 determines whether the pen tip or the pen tail has touched the display 280 by communicating with the electronic pen 290. The short-range communication circuit 219 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The power switch 222 is a switch for switching the power of the electronic whiteboard 2 on and off. The selection switches 223 are a group of switches for adjusting, for example, the brightness and color of the display 280.
[0132] Furthermore, the electronic whiteboard 2 includes a bus line 210. The bus line 210 is an address bus, a data bus, or the like for electrically connecting the components such as the CPU 201 shown in FIG.
[0133] The contact sensor 214 is not limited to an infrared blocking type, and various detection means may be used, such as a capacitive touch panel that identifies the contact position by detecting a change in capacitance, a resistive film touch panel that identifies the contact position by a change in voltage between two opposing resistive films, or an electromagnetic induction touch panel that identifies the contact position by detecting electromagnetic induction caused by an object touching the display unit. Also, the electronic pen controller 216 may determine whether or not the part of the electronic pen 290 that the user holds or other parts of the electronic pen have been touched, in addition to the pen tip and pen butt.
[0134] Next, the hardware configuration of a smartphone, which is an example of the communication terminal 300 of this embodiment, will be described with reference to Fig. 9. Fig. 9 is a diagram showing an example of the hardware configuration of a smartphone.
[0135] The smartphone 4 includes a CPU 401 , a ROM 402 , a RAM 403 , an EEPROM 404 , a CMOS sensor 405 , an image sensor I / F 406 , an acceleration / direction sensor 407 , a media I / F 409 , and a GPS receiving unit 411 .
[0136] Of these, the CPU 401 controls the overall operation of the smartphone 4. The ROM 402 stores the CPU 401 and programs used to drive the CPU 401, such as the IPL. The RAM 403 is used as a work area for the CPU 401. The EEPROM 404 reads and writes various data, such as smartphone programs, under the control of the CPU 401.
[0137] The CMOS (Complementary Metal Oxide Semiconductor) sensor 405 is a type of built-in imaging means that captures an image of a subject (mainly a self-portrait) and obtains image data under the control of the CPU 401. Note that instead of a CMOS sensor, an imaging means such as a CCD (Charge Coupled Device) sensor may also be used. The imaging element I / F 406 is a circuit that controls the driving of the CMOS sensor 405.
[0138] The acceleration / azimuth sensor 407 is an electronic magnetic compass that detects geomagnetism, a gyrocompass, an acceleration sensor, or other such sensors. The media I / F 409 controls the reading and writing (storage) of data from and to a recording medium 408 such as a flash memory. The GPS receiving unit 411 receives GPS signals from GPS satellites.
[0139] The smartphone 4 also includes a long-distance communication circuit 412, a CMOS sensor 413, an image sensor I / F 414, a microphone 415, a speaker 416, an audio input / output I / F 417, a display 418, an external device connection I / F (Interface) 419, a short-distance communication circuit 420, an antenna 420a of the short-distance communication circuit 420, and a touch panel 421.
[0140] Of these, the long-distance communication circuit 412 is a circuit that communicates with other devices via the network N. The CMOS sensor 413 is a type of built-in imaging means that captures an image of a subject and obtains image data under the control of the CPU 401. The imaging element I / F 414 is a circuit that controls the driving of the CMOS sensor 413.
[0141] The microphone 415 is a built-in circuit that converts sound into an electrical signal. The speaker 416 is a built-in circuit that converts the electrical signal into physical vibrations to produce sound such as music or voice. The sound input / output I / F 417 is a circuit that processes input and output of sound signals between the microphone 415 and the speaker 416 under the control of the CPU 401.
[0142] The display 418 is a type of display means such as a liquid crystal display or organic electroluminescence (EL) display that displays an image of a subject, various icons, etc. The external device connection I / F 419 is an interface for connecting various external devices. The short-range communication circuit 420 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 421 is a type of input means that allows a user to operate the smartphone 4 by pressing the display 418.
[0143] The smartphone 4 also includes a bus line 410. The bus line 410 is an address bus, a data bus, or the like for electrically connecting the components such as the CPU 401 shown in FIG.
[0144] Furthermore, the communication terminal 300 of this embodiment may be any device equipped with a communication function. The communication terminal 300 may be, for example, an output device such as a PJ (Projector), digital signage, a HUD (Head Up Display) device, industrial machinery, medical equipment, network home appliances, connected cars, notebook PCs (Personal Computers), mobile phones, tablet terminals, game consoles, PDAs (Personal Digital Assistants), digital cameras, wearable PCs, or desktop PCs.
[0145] (Second embodiment) The second embodiment will be described below with reference to the drawings. The second embodiment differs from the first embodiment in that the server is provided with the functionality of a video analysis unit. Therefore, in the following description of the second embodiment, only the differences from the first embodiment will be described, and components having the same functional configuration as the first embodiment will be assigned the same reference numerals as those used in the description of the first embodiment, and their description will be omitted.
[0146] 10 is a diagram illustrating the system configuration of an information processing system according to the second embodiment. An information processing system 100A according to this embodiment includes a server 200A and a communication terminal 300A.
[0147] The server 200A of this embodiment includes a video analysis unit 310 and a video generation instruction unit 320A. The communication terminal 300A of this embodiment does not include the video analysis unit 310.
[0148] The video analysis unit 310 of the server 200A of this embodiment holds information indicating the settings made on each communication terminal 300A regarding the display layout.
[0149] Then, the video analysis unit 310 analyzes the image data sent from each communication terminal 300A, and the video generation instruction unit 320A instructs each communication terminal 300A to generate image data including images of the location selected in accordance with the settings regarding the display layout.
[0150] In this embodiment, by providing the video analysis unit 310 in the server 200A, it is possible to reduce the processing load on the communication terminal 300A. Furthermore, since the analysis results of image data transmitted from a plurality of communication terminals 300A are accumulated in the server 200A, it is possible to improve the accuracy of facial expression estimation, for example.
[0151] The communication terminal in each of the above-described embodiments may be any device equipped with a communication function. The communication terminal 300 may be, for example, a projector (PJ), an output device such as digital signage, a head-up display (HUD), industrial machinery, medical equipment, network appliances, connected cars, personal computers (laptop PCs), mobile phones, tablet terminals, game consoles, personal digital assistants (PDAs), digital cameras, wearable PCs, or desktop PCs.
[0152] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or a conventional circuit module designed to perform each function described above.
[0153] Although the present invention has been described above based on the embodiments, the present invention is not limited to the requirements shown in the above embodiments. These requirements can be changed without departing from the spirit of the present invention, and can be appropriately determined depending on the application form. [Explanation of symbols]
[0154] 100, 100A Information Processing System 200, 200A Server 300, 300A communication terminal 310 Video Analysis Department 311 Face detection unit 312 Motion detection unit 313 Facial expression estimation section 314 Judgment section 320 Image Generation Unit 350 Overall Processing Unit 370 Display section [Prior art documents] [Patent documents]
[0155] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-244285
Claims
1. An information processing device for conducting video conferences between multiple locations, An estimation unit that estimates the facial expressions of participants in the video conference captured at the base and assigns a higher priority to the base that detects a specific facial expression, The modification section changes the layout of the display to show image data of locations with higher priority, An information processing device having a network processing unit that notifies other information processing devices of the settings related to the aforementioned layout.
2. The information processing apparatus according to Claim 1, wherein the estimation unit stores in the memory the frequency at which a participant's facial expression is estimated to be a specific facial expression for each location, and assigns priority to locations in order from those with the highest frequency.
3. The information processing apparatus according to claim 1, wherein a message indicating that the specific facial expression has been detected is displayed on the display unit.
4. Information processing device for conducting video conferences between multiple locations, The facial expressions of the participants in the video conference captured at each location are estimated, and locations that detect specific facial expressions are assigned a higher priority. The display layout has been changed to show image data of locations with higher priority. An information processing program that notifies other information processing devices of the settings related to the aforementioned layout and causes them to execute processing.
5. An information processing system comprising a plurality of information processing devices and a server device, which performs video conferencing between multiple locations, An estimation unit that estimates the facial expressions of participants in the video conference captured at the base and assigns a higher priority to the base that detects a specific facial expression, The modification section changes the layout of the display to show image data of locations with higher priority, An information processing system comprising a network processing unit that notifies other information processing devices of the settings related to the aforementioned layout.
6. An information processing method using an information processing system that has a plurality of information processing devices and a server device and performs video conferencing between a plurality of locations, wherein the information processing system is The facial expressions of the participants in the video conference captured at each location are estimated, and locations that detect specific facial expressions are assigned a higher priority. The display layout has been changed to show image data of locations with higher priority. An information processing method for notifying another information processing device of the settings related to the aforementioned layout.