Remote Conference Execution Program, Remote Conference Execution Method, and Remote Conference Execution Device

The remote conferencing system addresses the challenge of conveying expressions of participants with shielded faces by deforming and superimposing partial images based on speech content and emotion, thereby improving communication in remote meetings.

JP7694709B2Active Publication Date: 2025-06-18NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023566347
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-10
Filing Date
2022-12-07
Publication Date
2025-06-18
Estimated Expiration
2042-12-07

AI Technical Summary

Technical Problem

Existing remote conferencing methods cannot accurately convey the expressions of participants wearing masks or with shielded lips, limiting the ability of other participants to understand their emotions and reactions during a remote meeting.

Method used

A remote conferencing system that acquires a participant's face image and voice, detects shielded portions, estimates the content and emotion of the speech, deforms a partial image based on this information, and generates a superimposed image to display the participant's intended expression.

Benefits of technology

Enables other participants to accurately grasp the expressions and emotions of participants with shielded faces, enhancing communication and understanding in remote meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694709000001
    Figure 0007694709000001
  • Figure 0007694709000002
    Figure 0007694709000002
  • Figure 0007694709000003
    Figure 0007694709000003
Patent Text Reader

Abstract

In the present invention, an image of the face of a participant and voice of the participant are acquired, a masked portion of the participant in which a portion of the face is masked is detected from the image. The content of speech of the participant and an emotion thereof are inferred from the image or the voice. A partial image, which is an image of the portion of the face of the participant, is deformed according to the content of the speech of the participant and the emotion thereof. A superimposed image is generated in which the deformed partial image is superimposed on an area corresponding to the masked portion in the image of the participant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a recording medium, a remote conferencing method, and a remote conferencing apparatus.

Background Art

[0002] It is assumed that participants who participate in a remote conference from a public place participate in the conference while wearing a mask.

[0003] Patent Document 1 describes that in a communication conference system that conducts conversations between terminals, a receiving terminal displays a face image corresponding to a communication partner acquired in advance as a still image on a monitor. In the method described in Patent Document 1, the receiving terminal deforms the mouth area of the face image according to the vowel of the conversation sound transmitted from the communication partner.

[0004] Patent Document 2 describes that when a speaker detection system cannot detect a speaker from the movement of the lips and can detect a person with an appearance in which the lips are shielded, the person with an appearance in which the lips are shielded is detected as the speaker.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the method described in Patent Document 1, the receiving terminal deforms the mouth part of the face image according to the vowel of the conversation voice of the communication partner. Therefore, even if the communication partner has different expressions such as laughing or being angry, as long as the vowel is the same, the receiving terminal will display the face image with the same mouth part on the display. Even if the receiving terminal has previously acquired the face image of a participant who is wearing a mask and participating in the meeting without a mask, the method described in Patent Document 1 has a problem that it is impossible for the other participants in the remote meeting to grasp the expression of the participant who is participating in the meeting while wearing a mask.

[0007] In the method described in Patent Document 2, when a person with an appearance where the lips are shielded speaks, it is possible to detect that the person is the speaker. However, in the method described in Patent Document 2, it is impossible for others to grasp the expression of a person whose lips are shielded.

[0008] As described above, the methods described in Patent Document 1 to Patent Document 2 have a problem that it is impossible for other meeting participants to grasp the expressions of participants who are participating in a remote meeting with a part of their face shielded.

[0009] An example of the object of the present invention is to provide a recording medium, a remote meeting execution method, and a remote meeting execution device that enable other participants to grasp the expressions of participants who are participating in a remote meeting with a part of their face shielded.

Means for Solving the Problems

[0010] In one aspect of the present invention, a remote meeting execution program recorded on a computer-readable non-transitory recording medium causes a computer to have an acquisition function for acquiring an image of a participant's face and the participant's voice, a detection function for detecting a shielded part of a participant with a part of the face shielded from the image, an estimation function for estimating the content and emotion of the participant's speech from the image or voice, an image deformation function for deforming a partial image that is an image of a part of the participant's face according to the content and emotion of the participant's speech, and a superimposition function for generating a superimposed image in which the deformed partial image is superimposed on a range corresponding to the shielded part in the participant's image.

[0011] Also, in another aspect of the present invention, a remote conferencing execution method includes: acquiring an image of a participant's face and the participant's voice; detecting a blocked portion of a participant whose face is partially blocked from the image; estimating the content and emotion of the participant's speech from the image or the voice; deforming a partial image, which is an image of a part of the participant's face, according to the content and emotion of the participant's speech; and generating a superimposed image by superimposing the deformed partial image on a range corresponding to the blocked portion in the participant's image.

[0012] Also, in another aspect of the present invention, a remote conferencing execution apparatus includes: an acquisition unit that acquires an image of a participant's face and the participant's voice; a detection unit that detects a blocked portion of a participant whose face is partially blocked from the image; an estimation unit that estimates the content and emotion of the participant's speech from the image or the voice; an image deformation unit that deforms a partial image, which is an image of a part of the participant's face, according to the content and emotion of the participant's speech; and a superimposing unit that generates a superimposed image by superimposing the deformed partial image on a range corresponding to the blocked portion in the participant's image.

Advantages of the Invention

[0013] With the recording medium, remote conferencing execution method, and remote conferencing execution apparatus of the present invention, it becomes possible for other participants to grasp the expression of a participant who is participating in a remote conference with a part of their face blocked.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0015] [First Embodiment] The first embodiment of the present invention will be described.

[0016] FIG. 1 is a block diagram showing a configuration example of the remote conferencing execution apparatus 1 of the present embodiment.

[0017] The remote conferencing execution apparatus 1 of the present embodiment includes an acquisition unit 11, a detection unit 12, an estimation unit 13, an image deformation unit 14, and a superimposition unit 15.

[0018] For example, the remote conferencing execution apparatus 1 is realized by using a computer. The acquisition unit 11, the detection unit 12, the estimation unit 13, the image deformation unit 14, and the superimposition unit 15 of the remote conferencing execution apparatus 1 are realized by causing a computer to execute processing according to a remote conferencing execution program that realizes an acquisition function, a detection function, an estimation function, an image deformation function, and a superimposition function. That is, the remote conferencing execution program causes a computer to realize an acquisition function, a detection function, an estimation function, an image deformation function, and a superimposition function.

[0019] The acquisition unit 11 is an example of acquisition means. The acquisition unit 11 acquires an image of the face of a participant and the voice of the participant.

[0020] For example, the remote conferencing execution apparatus 1 is a server that controls a remote conference in which voice data and image data are received from a transmission terminal used by a participant, and the received voice data and image data are associated and transmitted to a reception terminal used by another participant. The transmission terminal includes a voice input device that receives an input of the voice of a participant and generates voice data corresponding to the voice, and a photographing device that photographs the participant and generates image data corresponding to the face of the participant. The acquisition unit 11 receives the image data and the voice data transmitted from the transmission terminal, and acquires an image of the face of the participant and the voice of the participant. The remote conferencing execution apparatus 1 may be a reception terminal. When the remote conferencing execution apparatus 1 is a reception terminal, the acquisition unit 11 may acquire an image of the face of the participant and the voice by receiving the voice data and the image data transmitted from the transmission terminal via a server that controls the remote conference.

[0021] Alternatively, the remote conference execution device 1 may be a transmission terminal including a voice input device that receives the input of the voice of a participant and generates voice data corresponding to the voice, and a photographing device that photographs the participant and generates image data of an image of the face of the participant. When the remote conference execution device 1 is a transmission terminal, voice data corresponding to the voice is input from the voice input device to the acquisition unit 11 of the remote conference execution device 1, and image data corresponding to the face of the participant is input from the photographing device. In this way, the acquisition unit 11 acquires the image of the face of the participant and the voice of the participant.

[0022] The detection unit 12 is an example of a detection means. The detection unit 12 detects a shielded portion of a participant whose part of the face is shielded from the image acquired by the acquisition unit 11. For example, it is assumed that a participant who participates in the remote conference from a public place participates in the conference while wearing a mask.

[0023] The estimation unit 13 is an example of an estimation means. The estimation unit 13 estimates the content and emotion of the speech of the participant from the image acquired by the acquisition unit 11 or the voice acquired by the acquisition unit 11. The estimation unit 13 may estimate the content and emotion of the speech of the participant using both the image of the participant and the voice of the participant.

[0024] The image deformation unit 14 is an example of an image deformation means. The image deformation unit 14 deforms a partial image, which is an image of a shielded part of the face of the participant, according to the content and emotion of the speech of the participant estimated by the estimation unit 13. The result of the estimation in which the content and emotion of the speech of the participant are estimated is input from the estimation unit 13 to the image deformation unit 14. The estimation result is data representing the result of estimating the content and emotion of the speech of the participant based on the voice of the participant or the image of the participant.

[0025] The superimposing unit 15 is an example of a superimposing means. The superimposing unit 15 generates a superimposed image in which the deformed partial image is superimposed on a range corresponding to the shielded portion in the image of the participant.

[0026] In this way, the remote conferencing device 1 acquires images of the faces of the participants and the voices of the participants, and detects the obscured portions of the participants whose faces are partially obscured from the images. The remote conferencing device 1 estimates the content and emotions of the participants' speech from the images or voices. The remote conferencing device 1 deforms a partial image, which is an image of a partially obscured portion of a participant's face, according to the content and emotions of the participant's speech, and generates a superimposed image in which the deformed partial image is superimposed on a range corresponding to the obscured portion in the participant's image. Since it is possible to show other participants the superimposed image in which the partial image deformed according to the content and emotions of the participant's speech is superimposed, it becomes possible for other participants to grasp the expressions of the participants who are participating in the remote conference with a part of their faces still obscured.

[0027] Next, with reference to FIG. 2, an operation example of the remote conferencing device 1 of the present embodiment will be described. FIG. 2 is a flowchart showing an operation example of the remote conferencing device 1.

[0028] The acquisition unit 11 acquires images of the faces of the participants and the voices of the participants (step S101).

[0029] The detection unit 12 detects the obscured portions of the participants whose faces are partially obscured from the images acquired in step S101 (step S102).

[0030] The estimation unit 13 estimates the content and emotions of the participants' speech from the images or voices acquired in step S101 (step S103).

[0031] The image deformation unit 14 deforms a partial image, which is an image of a partially obscured portion of a participant's face, according to the content and emotions of the participant's speech (step S104). The image deformation unit 14 uses the result of the estimation in step S103 in step S104.

[0032] The superimposition unit 15 generates a superimposed image in which the partial image deformed in step S104 is superimposed on a range corresponding to the obscured portion in the participant's image (step S105).

[0033] As described above, the remote conferencing device 1 enables the display of a superimposed image in which a partial image deformed according to the content and emotion of a participant's speech is superimposed to other participants. Thereby, it becomes possible for other participants to grasp the expression of a participant who is participating in the remote conference with a part of the face shielded.

[0034] [Second Embodiment] Next, the remote conferencing device 3 in the second embodiment of the present invention will be specifically described.

[0035] FIG. 3 is a block diagram showing a configuration example of a remote conferencing system according to the second embodiment of the present invention. As shown in FIG. 3, the remote conferencing system includes a transmission terminal 2, a remote conferencing device 3, and a reception terminal 4. In the second embodiment, the remote conferencing device 3 basically includes the configuration and functions of the remote conferencing device 1 of the first embodiment. Further, in the second embodiment, the remote conferencing device 3 is a server that receives audio data and image data from the transmission terminal 2 used by a participant, associates the received audio data and image data, and transmits them to the reception terminal 4 used by other participants to control a remote conference. For example, a remote conferencing execution program that realizes the functions of the remote conferencing device 3 is installed in a server that controls the remote conference.

[0036] FIG. 4 is a diagram for explaining the operation in the remote conferencing system according to the second embodiment. As shown in FIG. 4, a case will be described in which the transmission terminal 2 transmits the audio data and image data of a participant TP1 whose part of the face is shielded, and the reception terminal 4 receives the audio data and the superimposed image data of the superimposed image by the remote conferencing device 3 and displays the image on the display unit 44.

[0037] Referring to FIG. 3, the configuration of the transmission terminal 2 of the present embodiment will be described in detail. The transmission terminal 2 includes a photographing unit 21, an audio input unit 22, and a transmission unit 23. The transmission terminal 2 is, for example, any one of a smartphone, a notebook personal computer, and a desktop personal computer. In the example of FIG. 4, the transmission terminal 2 is a notebook personal computer.

[0038] The imaging unit 21 is provided at a position where it can image the faces of the participants. The imaging unit 21 performs imaging and outputs image data corresponding to the faces of the participants to the transmission unit 23. The imaging unit 21 is, for example, a camera built into the transmission terminal 2. The imaging unit 21 may be built into the transmission terminal 2, or may be another device other than the transmission terminal 2 that is connected to the transmission terminal 2 by wire or wirelessly.

[0039] The voice input unit 22 receives the input of the voices of the participants. The voice input unit 22 outputs voice data, which is data corresponding to the voice, to the transmission unit 23. The voice input unit 22 is, for example, a microphone built into the transmission terminal 2. The voice input unit 22 may be built into the transmission terminal 2, or may be another device other than the transmission terminal 2 that is connected to the transmission terminal 2 by wire or wirelessly.

[0040] Image data is input to the transmission unit 23 from the imaging unit 21. Voice data is input to the transmission unit 23 from the voice input unit 22. The transmission unit 23 associates the participant identification information, the image data, and the voice data of the participants using the transmission terminal 2 and transmits them to the remote conference execution device 3. The participant identification information is information that can identify each of the participants. Also, the image data indicates a still image or a moving image. The participant identification information is, for example, an ID (identifier) assigned by a server that controls the remote conference to the transmission terminal 2 or the participant.

[0041] With reference to FIG. 3, the configuration of the remote conference execution device 3 of the present embodiment will be described in detail. The remote conference execution device 3 includes an acquisition unit 31, a detection unit 32, an estimation unit 33, an image deformation unit 34, and a superimposition unit 35. The partial image storage unit 36 and the conference information storage unit 39 will be described later. The control information generation unit 37 receives an input from at least the superimposition unit 35. The transmission unit 38 receives an input from at least the control information generation unit 37.

[0042] The acquisition unit 31 acquires image data corresponding to the faces of the participants and voice data corresponding to the voices of the participants. Specifically, the acquisition unit 31 receives the participant identification information, the image data, and the voice data of the participants using the transmission terminal 2 from the transmission terminal 2 and acquires the images of the faces of the participants and the voices of the participants.

[0043] The acquisition unit 31 outputs the acquired participant identification information and the image data to the detection unit 32 and the superimposition unit 35 in association with each other. The acquisition unit 31 inputs the data used by the estimation unit 33 for estimating the content and emotion of the speech to the estimation unit 33. Specifically, the acquisition unit 31 outputs at least one of the voice data and the image data and the acquired participant identification information to the estimation unit 33 in association with each other. The acquisition unit 31 outputs the acquired participant identification information, the image data, and the voice data of the participant in association with each other to the control information generation unit 37.

[0044] The detection unit 32 receives the image data and the participant identification information of the participant in the image data from the acquisition unit 31. The detection unit 32 detects the shielded part of the participant whose part of the face is shielded from the image. The detection unit 32 outputs the partial information indicating the shielded part, the information indicating the range of the shielded part, and the input participant identification information to the image deformation unit 34 in association with each other. When the shielded part cannot be detected, the detection unit 32 notifies the control information generation unit 37 that the shielded part cannot be detected.

[0045] For example, the detection unit 32 calculates a feature amount from the acquired image. The detection unit 32 determines the range of the shielded part based on the calculated feature amount. For example, the detection unit 32 determines, as the range of the shielded part, an area having a feature amount whose difference from the feature amount extracted from the image of the mask registered in the storage unit (not shown) in advance is within a predetermined threshold value. For example, the feature amount extracted from the image of the mask may be stored in the partial image storage unit 36 in advance. Alternatively, the detection unit 32 may perform edge detection of the image to determine the range of the shielded part.

[0046] In addition, the detection unit 32 specifies the shielded part of the face and generates partial information. For example, when the detection unit 32 cannot specify the mouth of the participant from the image, the detection unit 32 generates partial information indicating the mouth.

[0047] FIG. 5 is a schematic diagram for explaining a shielding portion detected by the remote conferencing apparatus 3 according to the second embodiment. In FIG. 5, an example of a shielding portion detected by the detection unit 32 of the remote conferencing apparatus 3 from an image IM1 taken by a participant TP1 using the transmission terminal 2 shown in FIG. 4 is illustrated by a thick line. In the examples of FIGS. 4 and 5, since the participant TP1 is wearing a mask, the mouth, which is a part of the face, of the participant is shielded.

[0048] In the second embodiment, as shown in the examples of FIGS. 4 and 5, processing of an image of a participant whose mouth is shielded will be described. The remote conferencing apparatus 3 processes an image of a participant whose mouth is covered by a face cover or an image of a participant whose eyes are covered by wearing sunglasses in the same manner as an image of a participant wearing a mask. For example, the detection unit 32 of the remote conferencing apparatus 3 detects a shielding portion of a participant whose mouth or eyes are shielded. The remote conferencing apparatus 3 processes an image of a participant wearing a face shield in the same manner as an image of a participant wearing a mask. When wearing a face shield, even if the face shield is composed of a transparent film, there may be cases where the conversation partner cannot grasp the facial expression. For example, it is assumed that due to the face shield reflecting light, a part of the participant's face cannot be visually recognized by the conversation partner due to the light reflection.

[0049] The estimation unit 33 estimates the content and emotion of a participant's speech from the image acquired by the acquisition unit 31 or the audio acquired by the acquisition unit 31. The estimation unit 33 associates the participant identification information, timing information, estimation emotion information indicating the result of emotion estimation, and estimation speech information indicating the result of estimation of the content of the speech of the participant to be estimated, and outputs them to the image deformation unit 34. The timing information indicates the timing of the image data or audio data for which the content and emotion of the speech are estimated.

[0050] The estimation unit 33 includes an emotion estimation unit 331, a speech estimation unit 332, and an output unit 333.

[0051] The emotion estimation unit 331 estimates emotions based on the analysis result of the voice or the analysis result of the changes in the unobscured part of the participant's face image. The emotion estimation unit 331 outputs the estimated emotion information indicating the emotion estimated based on the input participant identification information, timing information, and analysis result of the emotion to the output unit 333 in an associated manner. The estimated emotion information includes at least information indicating the emotion. Specifically, for the analysis of the voice or the analysis of the changes in the unobscured part, a learned model created by machine learning may be used. The learned model includes one or more models that can classify various emotions such as joy, anger, sorrow, and happiness. For machine learning, a learning engine using a neural network may be used.

[0052] A method for the emotion estimation unit 331 to estimate emotions by analyzing the voice will be described. The emotion estimation unit 331 estimates the participant's emotion by acoustic analysis of the voice data. Alternatively, the emotion estimation unit 331 estimates the participant's emotion by linguistic analysis of the voice data.

[0053] Next, a method for the emotion estimation unit 331 to estimate emotions based on the analysis result of the changes in the unobscured part of the face image will be described. For example, when analyzing an image of a participant wearing a mask, the emotion estimation unit 331 identifies the movement of the eyes from the changes in the images of the time-series image data of the participant using the transmission terminal 2. The emotion estimation unit 331 analyzes the identified movement of the eyes to estimate the participant's emotion. In addition to these methods, any method can be used for estimating emotions.

[0054] Note that the emotion estimation unit 331 may further estimate the degree of emotion. The degree of emotion is, for example, a value indicating the level of emotion. When estimating the degree of emotion, the estimated emotion information further includes information indicating the estimated degree of emotion. When the information indicating the degree of emotion is included in the estimated emotion information, the image deformation unit 34 described later deforms the partial image according to the degree of emotion based on the information indicating the degree of emotion. For example, when it is estimated that the degree of joy is high, the image deformation unit 34 deforms the partial image so that the image has an upturned mouth more than when the degree of joy is low. As a result, it becomes possible for other participants to more precisely grasp the expression of the participant who is participating in the remote conference with a part of the face shielded than when the expression is deformed for each emotion.

[0055] The speech estimation unit 332 estimates the content of the participant's speech based on the analysis result of the voice. The speech estimation unit 332 outputs to the output unit 333 the estimated speech information indicating the content of the speech estimated based on the input participant identification information, timing information, and the analysis result of the voice, in association with each other. The estimated speech information includes at least information indicating the vowel estimated to have been spoken. The estimated speech information may include information indicating the consonant estimated to have been spoken. Note that, in addition to this method, any method can be used to estimate the content of the speech.

[0056] For example, a pre-trained model created by machine learning may be used for the analysis of the voice. The pre-trained model includes one or more models that can recognize speech corresponding to the voice. A learning engine using a neural network may be used for machine learning.

[0057] The output unit 333 receives the participant identification information, timing information, and estimated emotion information of the participant to be estimated from the emotion estimation unit 331. The output unit 333 receives the participant identification information, timing information, and estimated speech information of the participant to be estimated from the speech estimation unit 332. The output unit 333 outputs the participant identification information, timing information, estimated emotion information, and estimated speech information of the participant to be estimated to the image deformation unit 34 in association with each other.

[0058] The image deformation unit 34 receives, from the detection unit 32, partial information indicating a shielded part, information indicating the range of the shielded part, and participant identification information of a participant whose part of the face is shielded. The image deformation unit 34 also receives, from the estimation unit 33, participant identification information of the participant to be estimated, timing information, estimated emotion information, and estimated speech information.

[0059] The image deformation unit 34 reads, from the partial image storage unit 36, a partial image that is an image of the shielded part of the face of the participant based on the participant identification information of the participant whose part of the face is shielded and the partial information indicating the shielded part.

[0060] The partial image storage unit 36 stores partial image information in advance. The partial image information includes participant identification information of the participants participating in the meeting, partial information indicating a part of the face, and partial image data that is data of the partial image.

[0061] FIG. 6 is a diagram showing an example of the partial image information stored in the partial image storage unit 36 of the remote conference execution device 3. In the example of FIG. 6, the partial image storage unit 36 stores "partial image data PIMD1" and "partial image data PIMD2" of a participant whose participant identification information is "ID1". "Partial image data PIMD1" is the image data of the partial image of the "mouth" of the participant as shown in the partial information. "Partial image data PIMD2" is the image data of the partial image of the "eye" of the participant as shown in the partial information. Also, in the example of FIG. 6, the partial image storage unit 36 stores "partial image data PIMD3" of a participant whose participant identification information is "ID2". "Partial image data PIMD3" is the image data of the partial image of the "mouth" of the participant as shown in the partial information.

[0062] Specifically, the image deformation unit 34 reads out from the partial image storage unit 36 a partial image stored in the partial image storage unit 36 in association with the partial information indicating a shielded part of the input and the input participant identification information. The image deformation unit 34 deforms the partial image read out from the partial image storage unit 36 as follows. The image deformation unit 34 deforms the partial image based on the estimated emotion information indicating the result of emotion estimation by the estimation unit 33 and the estimated speech information indicating the result of estimation of the content of the speech. The image deformation unit 34 associates the timing information, the partial image data of the deformed partial image, the information indicating the range of the shielded part, and the participant identification information of the participant whose part of the face is shielded, and outputs the result to the superimposing unit 35.

[0063] With reference to FIG. 7, the deformation process of the partial image of the image deformation unit 34 will be specifically described.

[0064] FIG. 7 is a schematic diagram for explaining the deformation process of the partial image by the remote conferencing apparatus 3 of the second embodiment. Assume that the participant identification information input to the image deformation unit 34 indicates "ID1", and the partial information indicating a shielded part indicates "mouth". When the partial image information shown in FIG. 6 is stored in the partial image storage unit 36, the image deformation unit 34 performs the following process. The image deformation unit 34 reads out the "partial image data PIMD1" (the partial image data in the first row and the third column shown in FIG. 6) associated with the participant identification information indicating "ID1" and the partial information indicating "mouth". As shown on the left side of FIG. 7, the partial image PIM1 shown in the partial image data PIMD1 is an image of the mouth of the participant whose participant identification information is "ID1". When the estimated emotion information input from the estimation unit 33 indicates "joy" and the estimated speech information indicates "yes", the image deformation unit 34 performs the following deformation process. The image deformation unit 34 creates a deformed partial image PIM1' (the right figure in the example of FIG. 7) by deforming the partial image PIM1 of the partial image data PIMD1 according to the emotion of the participant (in this example, "joy") and the content of the speech (in this example, "yes").

[0065] Note that the partial image storage unit 36 may store in advance the partial images of the participants for each emotion. Alternatively, the partial image storage unit 36 may store in advance the partial images of the participants for each utterance. The image deformation unit 34 may read out the partial images stored in the partial image storage unit 36 in association with the estimated emotion or utterance, and deform the read-out partial images according to the content and emotion of the utterance.

[0066] Note that when the estimated utterance information indicates that there is no utterance, the image deformation unit 34 deforms the partial image based on the estimated emotion information indicating the result of the emotion estimation by the estimation unit 33 and the estimated utterance information indicating the result of the estimation of the content of the utterance. That is, when the participant TP1 is not speaking, the image deformation unit 34 deforms the partial image read from the partial image storage unit 36 so that it becomes a partial image corresponding to the estimated emotion when the participant TP1 is not speaking.

[0067] The superimposing unit 35 receives the participant identification information and the image data acquired by the acquisition unit 31. The superimposing unit 35 receives the timing information, the partial image data indicating the deformed partial image, the information indicating the range of the shielded portion, and the participant identification information of the participant whose part of the face is shielded from the image deformation unit 34. The superimposing unit 35 generates a superimposed image by superimposing the partial image deformed by the image deformation unit 34 on a range corresponding to the shielded portion in the participant's image. Specifically, the superimposing unit 35 uses, for superimposition, the image data at the timing indicated by the timing information among the image data acquired by the acquisition unit 31. The superimposing unit 35 superimposes the partial image deformed by the image deformation unit 34 on a range corresponding to the shielded portion in the image indicated by the image data at that timing. For example, the superimposing unit 35 generates superimposed image data in which the image in the range corresponding to the shielded portion in the image is replaced with the deformed partial image. The superimposing unit 35 associates the timing information, the participant identification information of the participant whose part of the face is shielded, and the superimposed image data of the superimposed image, and outputs the result to the control information generation unit 37.

[0068] For example, when image data showing a moving image composed of a predetermined number of frames is acquired, the image data of the timing indicated in the timing information indicates an image of a frame to be displayed on the display unit at the timing of the voice for which the content and emotion of the speech are estimated.

[0069] FIG. 8 is a schematic diagram for explaining a process of superimposing a partial image PIM1' by the remote conference execution apparatus 3 of the second embodiment on an image IM1 shown by the image data photographed in the remote conference. FIG. 8 is an example in the case where timing information, partial image data showing the deformed partial image PIM1', and participant identification information indicated by "ID1" are input from the image deformation unit 34 to the superimposition unit 35. Also, it is assumed that the image IM1 shown in FIG. 5 is an image photographed by a participant TP1 whose participant identification information is "ID1" at the timing indicated in the timing information in the remote conference. The superimposition unit 35 generates a superimposed image IM1' in which the partial image PIM1' is superimposed on the image IM1 based on information indicating the range of the shielded portion of the image IM1.

[0070] The control information generation unit 37 receives the participant identification information, voice data, and image data of the participant acquired by the acquisition unit 31. The control information generation unit 37 receives the timing information, the participant identification information of the participant with a part of the face shielded, and the superimposed image data of the superimposed image from the superimposition unit 35. The control information generation unit 37 generates output control information for displaying the superimposed image on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing of the voice for which the content and emotion of the speech are estimated.

[0071] When notified that the shielded portion cannot be detected, the control information generation unit 37 performs the following operation. The control information generation unit 37 generates output control information for displaying the image acquired by the acquisition unit 31 on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing corresponding to the voice acquired by the acquisition unit 31.

[0072] The control information generation unit 37 outputs the generated output control information to the transmission unit 38.

[0073] Output control information is input from the control information generation unit 37 to the transmission unit 38. The transmission unit 38 reads out destination information indicating the destinations of the participants in the remote conference from the conference information storage unit 39. The destination of the output control information is, for example, a terminal used by other participants participating in the remote conference. The destinations indicated by the destination information include the receiving terminal 4. The transmission unit 38 transmits the output control information to the destinations indicated by the destination information. When a shielded portion is detected, the output control information includes the superimposed image data of the superimposed image and the audio data corresponding to the audio. When the shielded portion cannot be detected, the output control information includes the acquired image data and the audio data corresponding to the audio.

[0074] The conference information storage unit 39 stores the participant identification information of the participants participating in the remote conference and the destination information in association with each other. The destination information is, for example, an IP (Internet Protocol) address.

[0075] Referring to FIG. 3, the configuration of the receiving terminal 4 of the present embodiment will be described in detail. The receiving terminal 4 includes a receiving unit 41, an output control unit 42, an audio output unit 43, and a display unit 44. The receiving terminal 4 is, for example, any one of a smartphone, a notebook computer, and a desktop computer. In the example of FIG. 4, the receiving terminal 4 is a desktop computer. FIG. 4 shows an example in which an image in which a participant TP1 using the transmitting terminal 2 is photographed and a partial image is superimposed is displayed on the display unit 44 of the receiving terminal 4.

[0076] The receiving unit 41 receives the output control information from the remote conference execution device 3. The receiving unit 41 outputs the output control information to the output control unit 42.

[0077] The output control unit 42 controls the audio output unit 43 and the display unit 44 based on the output control information. The output control unit 42 causes the audio output unit 43 to output audio corresponding to the audio data based on the output control information. The output control unit 42 causes the display unit 44 to display an image based on the output control information so that the image is displayed at a timing corresponding to the audio output from the audio output unit 43.

[0078] The voice output unit 43 outputs voice under the control of the output control unit 42. The voice output unit 43 is, for example, a speaker built into the receiving terminal 4.

[0079] The display unit 44 displays an image under the control of the output control unit 42. The display unit 44 is, for example, a display built into the receiving terminal 4 or connected to the receiving terminal 4. As shown in FIG. 4, a superimposed image in which a partial image deformed according to the content and emotion of the speech of the participant TP1 using the transmitting terminal 2 is superimposed is displayed on the display unit 44.

[0080] In this way, the remote conference execution device 3 acquires the image of the participant's face and the participant's voice, and detects the shielded portion of the participant with a part of the face shielded from the image. The remote conference execution device 3 estimates the content and emotion of the participant's speech from the image or voice. The remote conference execution device 3 deforms a partial image, which is an image of a shielded part of the participant's face, according to the content and emotion of the participant's speech, and generates a superimposed image in which the deformed partial image is superimposed on a range corresponding to the shielded part in the participant's image. Since it is possible to show other participants a superimposed image in which a partial image deformed according to the content and emotion of the participant's speech is superimposed, it becomes possible for other participants to grasp the expression of the participant who is participating in the remote conference with a part of the face shielded.

[0081] Next, with reference to FIGS. 9 to 10, an operation example of the remote conference execution system of the present embodiment will be described. FIG. 9 is a sequence diagram showing an operation example of the remote conference execution system. FIG. 10 is a flowchart showing an operation example of the remote conference execution device 3.

[0082] First, with reference to FIG. 9, the operation of the remote conference execution system will be described. The operation of the remote conference execution device 3 when the shielding range cannot be detected will be described later with reference to FIG. 10. In FIG. 9, the operation of the remote conference execution system when the shielding range can be detected is shown.

[0083] The imaging unit 21 of the transmitting terminal 2 performs imaging. The voice input unit 22 receives the input of the voices of the participants (step S201). The imaging unit 21 outputs image data corresponding to the faces of the participants to the transmitting unit 23. The voice input unit 22 outputs voice data, which is data corresponding to the voice, to the transmitting unit 23.

[0084] The transmitting unit 23 associates the participant identification information, the image data, and the voice data of the participant using the transmitting terminal 2 and transmits them to the remote conference execution device 3 (step S202).

[0085] The acquisition unit 31 of the remote conference execution device 3 receives the participant identification information, the image data, and the voice data of the participant using the transmitting terminal 2 from the transmitting terminal 2. In this way, the acquisition unit 31 acquires the images of the faces of the participants and the voices of the participants.

[0086] The acquisition unit 31 outputs the acquired participant identification information and the image data in association with each other to the detection unit 32 and the superimposing unit 35. The acquisition unit 31 outputs at least one of the voice data and the image data and the acquired participant identification information in association with each other to the estimation unit 33. The acquisition unit 31 outputs the acquired participant identification information, the image data, and the voice data of the participant in association with each other to the control information generation unit 37.

[0087] The detection unit 32 detects the shielded portion of the participant whose face is partially shielded from the image (step S203). The detection unit 32 outputs the partial information indicating the shielded portion, the information indicating the range of the shielded portion, and the input participant identification information in association with each other to the image transformation unit 34.

[0088] The estimation unit 33 estimates the content and emotion of the speech of the participant from the image acquired by the acquisition unit 31 or the voice acquired by the acquisition unit 31 (step S204). The estimation unit 33 outputs the participant identification information of the participant to be estimated, the timing information, the estimated emotion information indicating the result of the emotion estimation, and the estimated speech information indicating the result of the estimation of the content of the speech in association with each other to the image transformation unit 34.

[0089] The image deformation unit 34 reads out a partial image from the partial image storage unit 36 based on the participant identification information and partial information of the participant with a part of the face shielded (step S205). The image deformation unit 34 deforms the partial image read from the partial image storage unit 36 according to the content and emotion of the participant's speech (step S206). The image deformation unit 34 outputs to the superimposing unit 35 by associating the timing information, the partial image data which is the data of the deformed partial image, the information indicating the range of the shielded part, and the participant identification information of the participant with a part of the face shielded.

[0090] The superimposing unit 35 generates a superimposed image by superimposing the partial image deformed by the image deformation unit 34 on a range corresponding to the shielded part in the participant's image (step S207). The superimposing unit 35 outputs to the control information generation unit 37 by associating the timing information, the participant identification information of the participant with a part of the face shielded, and the superimposed image data which is the data of the superimposed image.

[0091] The control information generation unit 37 generates output control information for displaying the superimposed image on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing of the voice for which the content and emotion of the speech have been estimated (step S208). The control information generation unit 37 outputs the generated output control information to the transmitting unit 38.

[0092] The transmitting unit 38 transmits the output control information to the communication destination indicated by the communication destination information (step S209). The communication destination includes the receiving terminal 4.

[0093] The receiving unit 41 of the receiving terminal 4 receives the output control information from the remote conference execution device 3. The receiving unit 41 outputs the output control information to the output control unit 42.

[0094] The output control unit 42 controls the voice output unit 43 and the display unit 44 based on the output control information (step S210). In step S210, the output control unit 42 causes the voice output unit 43 to output a voice corresponding to the voice data based on the output control information. In step S210, the output control unit 42 causes the display unit 44 to display the superimposed image based on the output control information so that an image is displayed at a timing corresponding to the voice output from the voice output unit 43.

[0095] The voice output unit 43 outputs a voice under the control of the output control unit 42. The display unit 44 displays an image under the control of the output control unit 42 (step S211). The image displayed in step S211 is a superimposed image.

[0096] Next, with reference to FIG. 10, the operation of the remote conference execution device 3 will be described. The operation in FIG. 10 details the operations from step S203 to step S209 in FIG. 9.

[0097] The acquisition unit 31 receives from the transmission terminal 2 the participant identification information of the participant using the transmission terminal 2, the image data corresponding to the face of the participant, and the voice data corresponding to the voice of the participant. In this way, the acquisition unit 31 acquires the image of the participant's face and the voice of the participant (step S301).

[0098] The acquisition unit 31 associates the acquired participant identification information and image data and outputs them to the detection unit 32 and the superimposition unit 35. The acquisition unit 31 associates at least one of the voice data and the image data with the acquired participant identification information and outputs it to the estimation unit 33. For example, in step S303 described later, when the estimation unit 33 uses the voice data for estimating the content and emotion of the speech, the acquisition unit 31 associates the voice data with the participant identification information and outputs it to the estimation unit 33. The acquisition unit 31 associates the acquired participant identification information, image data, and voice data of the participant and outputs them to the control information generation unit 37.

[0099] The detection unit 32 detects the occluded part of a participant whose part of the face is occluded from the image (step S302). When the occluded part can be detected (step S302, YES), the detection unit 32 outputs to the image deformation unit 34 by associating the part information indicating the occluded part, the information indicating the range of the occluded part, and the input participant identification information.

[0100] When the occluded part cannot be detected (step S302, NO), the detection unit 32 notifies the control information generation unit 37 that the occluded part cannot be detected. Also, the estimation unit 33 does not perform the operation of step S303. The image deformation unit 34 does not perform the operations from step S304 to step S305. The superimposition unit 35 does not perform the operation of step S306.

[0101] The estimation unit 33 estimates the content and emotion of the participant's speech from the image acquired in step S301 or the acquired voice (step S303). The estimation unit 33 outputs to the image deformation unit 34 by associating the participant identification information, the timing information, the estimated emotion information, and the estimated speech information of the participant to be estimated.

[0102] In step S304, the emotion estimation unit 331 of the estimation unit 33 estimates the emotion based on the analysis result of the voice or the analysis result of the change in the unoccluded part of the participant's face image. Also, in step S304, the speech estimation unit 332 estimates the content of the participant's speech based on the analysis result of the voice. The estimation unit 33 performs the estimation of the speech content and the estimation of the emotion in an arbitrary order. For example, the estimation of the emotion by the emotion estimation unit 331 and the estimation of the speech content by the speech estimation unit 332 may be performed in parallel. Also, the estimation of the emotion by the emotion estimation unit 331 may be performed after the estimation of the speech content by the speech estimation unit 332.

[0103] The image deformation unit 34 reads out the partial image of the participant from the partial image storage unit 36 based on the participant identification information and the partial information of the participant whose part of the face is occluded (step S304).

[0104] The image deformation unit 34 deforms the partial image read from the partial image storage unit 36 according to the content and emotion of the participant's speech (step S305). In step S305, the image deformation unit 34 deforms the partial image based on the estimated emotion information indicating the result of the emotion estimation by the estimation unit 33 and the estimated speech information indicating the result of the estimation of the speech content. The image deformation unit 34 outputs to the superimposition unit 35 in association with the timing information, the partial image data of the deformed partial image, the information indicating the range of the shielded portion, and the participant identification information of the participant whose part of the face is shielded.

[0105] The superimposition unit 35 generates a superimposed image in which the partial image deformed by the image deformation unit 34 in step S305 is superimposed on the range corresponding to the shielded portion in the participant's image (step S306). The superimposition unit 35 outputs to the control information generation unit 37 in association with the timing information, the participant identification information of the participant whose part of the face is shielded, and the superimposed image data of the superimposed image.

[0106] The control information generation unit 37 generates output control information (step S307). When the superimposed image data is input, the control information generation unit 37 generates output control information for displaying the superimposed image on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing of the voice for which the content and emotion are estimated.

[0107] When notified in step S302 that the shielded portion cannot be detected, the control information generation unit 37 performs the following operation in step S308. The control information generation unit 37 generates output control information for displaying the image acquired by the acquisition unit 31 on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing corresponding to the voice acquired by the acquisition unit 31. The control information generation unit 37 outputs the output control information to the transmission unit 38.

[0108] The transmission unit 38 transmits the output control information to the communication destination indicated by the communication destination information (step S308).

[0109] Note that the remote conferencing apparatus 3 can perform the operations from step S302 to step S307 in an arbitrary order. For example, the remote conferencing apparatus 3 may operate in the following order.

[0110] The detection unit 32 performs the operation of step S302. Next, the image deformation unit 34 performs the operation of step S304. Then, instead of the operation of step S306, the superimposition unit 35 superimposes the partial image before being deformed by the image deformation unit 34 on a range corresponding to the shielded portion in the participant's image. Instead of the operation of step S305, the image deformation unit 34 deforms the superimposed image according to the content and emotion of the participant's speech. The control information generation unit 37 performs the operation of step S307. Note that the estimation unit 33 performs the operation of step S303 before the start of the image deformation process by the image deformation unit 34.

[0111] As described above, the remote conferencing apparatus 3 of the present embodiment acquires an image of a participant's face and the participant's voice, and detects a shielded portion of a participant with a part of the face shielded from the image. The remote conferencing apparatus 3 estimates the content and emotion of the participant's speech from the image or voice. The remote conferencing apparatus 3 deforms a partial image, which is an image of a shielded part of the participant's face, according to the content and emotion of the participant's speech. The remote conferencing apparatus 3 generates a superimposed image in which the deformed partial image is superimposed on a range corresponding to the shielded portion in the participant's image. Since it is possible to show the other participants the superimposed image in which the partial image deformed according to the content and emotion of the participant's speech is superimposed, it becomes possible for the other participants to grasp the expression of the participant who is participating in the remote conference with a part of the face shielded.

[0112] When the remote conferencing apparatus 3 of this embodiment cannot detect a masked portion, it generates output control information for causing the acquired image to be displayed on the display unit at a timing corresponding to the acquired audio. Since a masked portion can be detected while the participant is wearing a mask, the remote conferencing apparatus 3 transmits output control information for causing the superimposed image to be displayed on the display unit to the receiving terminal 4. After the participant removes the mask, since the remote conferencing apparatus 3 cannot detect the masked portion, it transmits output control information for causing the acquired image in which the image is not deformed to be displayed on the display unit. Thereby, when a participant using the transmitting terminal 2 removes the mask during the remote conference, it is possible to stop the display on the display unit 44 of the receiving terminal 4 of the superimposed image in which the deformed partial image is superimposed on the range corresponding to the masked portion.

[0113] [Modification Example 1 of the Second Embodiment] The remote conferencing apparatus according to Modification Example 1 of the second embodiment uses machine learning for the deformation process of the partial image. The image deformation unit of this modification example has a model generation function. The learning data includes, for example, face images of a plurality of persons in which at least one of the content and emotion of speech is different, information indicating the emotion represented by each face image, and information indicating the content of speech of each person photographed in the plurality of face images. The image deformation unit of this modification example generates a deformation model for deforming the partial image according to the content and emotion of speech based on the learning data. The image deformation unit uses the generated deformation model to deform the partial image according to the estimated content and emotion of speech. The input to the deformation model is partial image data, estimated emotion information, and estimated speech information. The output from the deformation model is partial image data indicating the deformed partial image.

[0114] [Modification Example 2 of the Second Embodiment] In addition to the deformation process of the partial image, the image deformation unit of the remote conferencing device according to the second modification example 2 also deforms the unobscured portions of the face images of the participants acquired by the acquisition unit according to the content and emotion of the speech. For example, as shown in FIG. 4, when the mouth of participant TP1 is obscured by a mask, the image deformation unit of this modification example deforms the unobscured portion (the eyes in the example of FIG. 4) of the face image of participant TP1 in the image data acquired by the acquisition unit according to the content and emotion of the speech.

[0115] The superimposing unit of this modification example generates a superimposed image in which the partial image deformed by the image deformation unit is superimposed on a range corresponding to the obscured portion in the image deformed by the image deformation unit.

[0116] The image deformation unit of the remote conferencing device according to this modification example also deforms the unobscured portions of the face images of the participants according to the content and emotion of the speech. The superimposing unit of the remote conferencing device according to this modification example generates a superimposed image in which the unobscured portions of the face images of the participants are also deformed. Thereby, the remote conferencing device according to this modification example can match the expression represented by the deformed partial image with the expression represented by the unobscured portion of the face image of the participant. Thereby, the remote conferencing device according to this modification example can match the expression of the deformed partial image with the expression of the unobscured portion of the face image. For this reason, the remote conferencing device according to this modification example can reduce the possibility that an unnatural expression is displayed on the display unit of the receiving terminal.

[0117] [Second Modification Example of the Second Embodiment] The remote conferencing device according to the second modification example 3 of the second embodiment is a transmitting terminal. For example, by installing a remote conferencing execution program that realizes the functions of the remote conferencing device 3 in the transmitting terminal 2, the transmitting terminal has the functions of the remote conferencing device 3. Regarding this modification example, the differences from the remote conferencing device 3 of the second embodiment will be described. Note that the configurations of the detection unit, estimation unit, image deformation unit, superimposing unit, and control information generation unit are the same as those of the remote conferencing device 3 in the second embodiment shown in FIG. 3, and thus the description thereof will be omitted.

[0118] In the acquisition unit of this modification example, voice data is input from the voice input unit, and image data is input from the imaging unit. In this way, the acquisition unit of this modification example acquires an image of the participant's face and the participant's voice.

[0119] The partial image information stored in the partial image storage unit of this modification example stores at least the partial image information of the participants who participate in the remote conference using the transmission terminal.

[0120] The conference information storage unit of this modification example stores communication destination information indicating the server that controls the remote conference.

[0121] The transmission unit of this modification example transmits output control information to the server that is the communication destination indicated by the communication destination information. Note that the output control information is transmitted to the reception terminal via the server.

[0122] Note that the remote conference execution device may be a reception terminal. When the reception terminal has the function of the remote conference execution device, it acquires the participant identification information, image data, and voice data of the participant using the transmission terminal, which are transmitted from the transmission terminal via the server that controls the remote conference. Further, the output control unit controls the voice output unit and the display unit using the output control information generated by the control information generation unit. When the reception terminal has the function of the remote conference execution device, the configurations of the detection unit, estimation unit, image deformation unit, superimposition unit, and control information generation unit are the same as those of the remote conference execution device 3 in the second embodiment shown in FIG. 3.

[0123] [Third Embodiment] Next, the remote conference execution device 5 in the third embodiment of the present invention will be specifically described.

[0124] FIG. 11 is a block diagram showing a configuration example of a remote conference execution system according to the third embodiment of the present invention. As shown in FIG. 11, the remote conference execution system includes a transmission terminal 6, a remote conference execution device 5, and a reception terminal 4. Further, the transmission terminal 6 is connected to an imaging device 7 and a voice input device 8.

[0125] In the third embodiment, the remote conferencing execution device 5 basically includes the configuration and functions of the remote conferencing execution device 3 of the second embodiment. The remote conferencing execution device 5 of the third embodiment is different from the remote conferencing execution device 3 of the second embodiment in the following points. The remote conferencing execution device 5 of the third embodiment is different in that it identifies the speaker among a plurality of participants based on the images and voices of each participant acquired by the acquisition unit 51. Further, the detection unit 52, the estimation unit 33, the image deformation unit 34, and the superimposition unit 35 are different in that they execute processing on the speaker.

[0126] With reference to FIGS. 11 to 13, each configuration of the remote conferencing execution system of the present embodiment will be described in detail.

[0127] FIG. 12 is a diagram schematically showing the state of a remote conference in the third embodiment. FIG. 12 shows the states of participants (in the example of FIG. 12, participant TP1, participant TP2, participant TP3, and participant TP4) photographed by the photographing device 7 connected to the transmission terminal 6. As shown in FIG. 12, it is assumed that participants participating in a remote conference wear masks to participate in the remote conference in order to avoid infection such as COVID-19.

[0128] The photographing device 7 is installed at a position where a plurality of participants (in the example of FIG. 12, participants TP1 to TP4) participating in the remote conference can be photographed. In the example shown in FIG. 12, the photographing device 7 is installed above the external display that displays the conference materials. In the example shown in FIG. 12, the photographing device 7 communicates with the transmission terminal 6 by wire, but the photographing device 7 may communicate with the transmission terminal 6 wirelessly. The photographing device 7 performs photographing and transmits the image data to the transmission terminal 6. The photographing device 7 corresponds to the photographing unit 21 of the transmission terminal 2 of the second embodiment.

[0129] FIG. 13 is a diagram schematically showing the image IM2 photographed by the photographing device 7 of the third embodiment. FIG. 13 is an example of an image when the participants TP1 to TP4 shown in FIG. 12 face the photographing device 7 and are photographed. The participants photographed by the photographing device 7 do not have to face the photographing device 7.

[0130] The voice input device 8 is installed at a position where it can receive the voice inputs of a plurality of participants (in the example of FIG. 12, participants TP1 to TP4) participating in the remote conference. The voice input device 8 receives voice inputs. The voice input device 8 transmits voice data, which is data corresponding to the voice, to the transmission terminal 6. The voice input device 8 corresponds to the voice input unit 22 of the transmission terminal 2 in the second embodiment.

[0131] Referring to FIG. 11, the configuration of the transmission terminal 6 in this embodiment will be described.

[0132] The transmission terminal 6 includes a transmission / reception unit 61. In the example of FIG. 12, the transmission terminal 6 is a notebook personal computer.

[0133] The transmission / reception unit 61 receives image data from the imaging device 7. The transmission / reception unit 61 receives voice data from the voice input device 8. The transmission / reception unit 61 associates the image data and the voice data and transmits them to the remote conference execution device 5.

[0134] Referring to FIG. 11, the configuration of the remote conference execution device 5 in this embodiment will be described.

[0135] The remote conference execution device 5 includes an acquisition unit 51, a detection unit 52, an estimation unit 33, an image deformation unit 34, and a superimposition unit 35. The speaker identification unit 53 receives an input from at least the acquisition unit 51. The conference information storage unit 39 and the feature amount storage unit 54 will be described later.

[0136] Also, since the configurations of the partial image storage unit 36, the transmission unit 38, and the conference information storage unit 39 of the remote conference execution device 5 in this embodiment are the same as those in the second embodiment shown in FIG. 3, the same reference numerals as those in FIG. 3 are assigned to the corresponding elements and the common description is omitted.

[0137] The acquisition unit 51 acquires images of the faces of the participants and the voices of the participants. The acquisition unit 51 associates the voice data and the image data and outputs them to the speaker identification unit 53 and the control information generation unit 37.

[0138] The speaker identification unit 53 is an example of a speaker identification means. The speaker identification unit 53 identifies the speaker who is speaking among a plurality of participants based on the images and voices of each participant acquired by the acquisition unit 51. The speaker identification unit 53 identifies the range of the image in which the speaker is photographed and generates speaker range information. The speaker range information indicates the range of the image in which the speaker is photographed. The speaker identification unit 53 associates the speaker range information, the participant identification information of the speaker, the image data, and the voice data and outputs them to the detection unit 52, the estimation unit 33, and the superimposition unit 35.

[0139] When the speaker cannot be identified, the speaker identification unit 53 notifies the control information generation unit 37 that the speaker cannot be identified.

[0140] The image recognition performed by the speaker identification unit 53 will be described in detail.

[0141] Voice data and image data are input to the speaker identification unit 53 from the acquisition unit 51. The speaker identification unit 53 extracts the feature amounts of the face images from the image data. When a plurality of participants are included in the image of the image data, the speaker identification unit 53 extracts the feature amounts of the face images of each of the participants included in the image. The speaker identification unit 53 extracts the feature amounts from the face images using an arbitrarily set method. The speaker identification unit 53 collates whether or not the feature amounts of the face images whose similarity with the feature amounts of the face images indicating the extraction results is equal to or more than a predetermined value are stored in the feature amount storage unit 54.

[0142] The feature amount storage unit 54 stores in advance, in association with each other, the participant identification information of the participants participating in the remote conference, the face image feature amounts that are the feature amounts of the face images of the participants participating in the remote conference, and the voice feature amounts that are the feature amounts of the voices of the participants participating in the remote conference.

[0143] FIG. 14 is a diagram showing an example of face image feature amounts and voice feature amounts stored in the feature amount storage unit 54 of the remote conference execution apparatus 5. In the example of FIG. 14, for each of the participants whose participant identification information is "ID1" and "ID2", a face image feature amount that is a feature amount of the face image of that participant and a voice feature amount that is a feature amount of the voice are associated and stored in the feature amount storage unit 54.

[0144] The speaker identification unit 53 performs the following operation by collating whether or not a face image feature amount having a similarity with the face image feature amount indicating the extraction result equal to or greater than a predetermined value is stored in the feature amount storage unit 54. The speaker identification unit 53 identifies the participant identification information associated with and stored in the feature amount storage unit 54 for a face image feature amount having a similarity with the face image feature amount indicating the extraction result equal to or greater than a predetermined value. The speaker identification unit 53 identifies the participant identification information for each of the face image feature amounts extracted from one image. By the speaker identification unit 53 identifying the participant identification information using the face image feature amount, it is possible to identify the participant identification information of each of the plurality of participants photographed by the photographing apparatus 7. Also, hereinafter, the process in which the speaker identification unit 53 identifies the participant identification information using the face image feature amount is referred to as an image recognition process.

[0145] The voice recognition performed by the speaker identification unit 53 will be described in detail.

[0146] The speaker identification unit 53 extracts a feature amount from the voice corresponding to the voice data using an arbitrary method set in advance. The speaker identification unit 53 collates whether or not a voice feature amount having a similarity with the voice feature amount indicating the extraction result equal to or greater than a predetermined value is stored in the feature amount storage unit 54. The speaker identification unit 53 identifies the participant identification information associated with and stored in the feature amount storage unit 54 for a voice feature amount having a similarity with the voice feature amount indicating the extraction result equal to or greater than a predetermined value. By the speaker identification unit 53 identifying the participant identification information using the voice feature amount, it is possible to identify the participant identification information of the speaker who uttered the voice input to the voice input device 8. Also, hereinafter, the process in which the speaker identification unit 53 identifies the participant identification information using the voice feature amount is referred to as a voice recognition process.

[0147] The speaker identification unit 53 determines whether the participant identification information identified in the image recognition process includes the participant identification information identified in the voice recognition process. If the participant identification information identified in the image recognition process includes the participant identification information identified in the voice recognition process, the speaker identification unit 53 determines that the speaker has been identified. When it is determined that the speaker has been identified, the speaker identification unit 53 identifies the participant of the participant identification information identified in the voice recognition process as the speaker.

[0148] If the participant identification information identified in the image recognition process does not include the participant identification information identified in the voice recognition process, or if the participant identification information cannot be identified in the voice recognition process, the speaker identification unit 53 determines that the speaker cannot be identified. For example, when it is assumed that the speaker cannot be identified, it is assumed that a participant outside the shooting area of the imaging device 7 is speaking. Alternatively, it is assumed that the participant identification information of the speaker cannot be identified in the image recognition process because the speaker photographed by the imaging device 7 is looking down or facing the other side.

[0149] Note that the speaker identification unit 53 may identify the speaker from the image. For example, when the shielded part is the eye, among the images of each participant acquired by the acquisition unit 51, the mouth part of the image corresponding to the speaker moves. The speaker identification unit 53 may identify the participant from the image and identify the participant in which movement is detected in the image of the mouth part among the participants as the speaker.

[0150] The detection unit 52 receives the speaker range information, the participant identification information of the speaker, the image data, and the voice data from the speaker identification unit 53. The detection unit 52 detects the shielded part of the speaker whose face part is shielded from the acquired image in the range indicated by the speaker range information. The detection unit 52 associates the partial information indicating the shielded part, the information indicating the range of the shielded part, and the participant identification information of the speaker and outputs them to the image deformation unit 34. When the shielded part cannot be detected, the detection unit 52 notifies the control information generation unit 37 that the shielded part cannot be detected.

[0151] The speaker range information, the participant identification information of the speaker, the image data, and the voice data are input from the speaker identification unit 53 to the estimation unit 33. The estimation unit 33 estimates the content and emotion of the speaker's speech from the image of the part corresponding to the range indicated by the speaker range information or the voice acquired by the acquisition unit 51.

[0152] Each of the emotion estimation unit 331, the speech estimation unit 332, and the output unit 333 of the estimation unit 33 is the same as that of the second embodiment except that, instead of the image data, an image of a part corresponding to the range indicated by the speaker range information is used. Therefore, the same reference numerals as those in FIG. 3 are assigned to the corresponding elements, and the description of the configuration of the estimation unit 33 is omitted.

[0153] Since the configuration of the image transformation unit 34 is the same as the configuration in the second embodiment shown in FIG. 3, the same reference numerals as those in FIG. 3 are assigned to the corresponding elements, and the description is omitted.

[0154] Since the configuration of the superimposing unit 35 of the remote conference execution device 5 in the present embodiment is the same as each of the configurations in the second embodiment shown in FIG. 3, the same reference numerals as those in FIG. 3 are assigned to the corresponding elements, and the common description is omitted.

[0155] The speaker range information, the participant identification information of the speaker, the image data, and the voice data are input to the superimposing unit 35. Timing information, partial image data indicating the deformed partial image, information indicating the range of the shielded part, and the participant identification information of the speaker with a part of the face shielded are input from the image transformation unit 34 to the superimposing unit 35.

[0156] The superimposing unit 35 generates a superimposed image in which the partial image deformed by the image transformation unit 34 is superimposed on a range corresponding to the shielded part in the speaker's image. The superimposing unit 35 associates the timing information, the participant identification information of the speaker, and the superimposed image data of the superimposed image and outputs them to the control information generation unit 37.

[0157] FIG. 15 is a schematic diagram for explaining a process of superimposing a partial image PIM1' by the remote conferencing apparatus 5 of the third embodiment on an image IM2 captured in a remote conference. FIG. 15 is an example in the case where timing information, partial image data of the deformed partial image PIM1', and participant identification information indicating the speaker TP1 are input from the image deformation unit 34 to the superimposing unit 35. The superimposing unit 35 generates a superimposed image IM2' in which the partial image PIM1' is superimposed on the image IM2 based on information indicating the range of the shielding portion of the speaker TP1 in the image IM2.

[0158] When notified that the speaker cannot be identified or the shielding portion cannot be detected, the control information generation unit 37 performs the following process. The control information generation unit 37 generates output control information for causing the display unit (in this example, the display unit 44 of the receiving terminal 4) to display the image acquired by the acquisition unit 51 at a timing corresponding to the voice acquired by the acquisition unit 51. Since the process of the control information generation unit 37 in this embodiment when the speaker can be identified and the shielding portion can be detected is the same as the process performed by the control information generation unit 37 in the second embodiment when the shielding portion can be detected, the description thereof is omitted.

[0159] As described above, the detection unit 52, the estimation unit 33, the image deformation unit 34, and the superimposing unit 35 execute processing on the speaker.

[0160] Since each configuration of the receiving terminal 4 in this embodiment is the same as each of the configurations in the second embodiment shown in FIG. 3, the corresponding elements are denoted by the same reference numerals as in FIG. 3 and the description thereof is omitted.

[0161] In this way, the remote conference execution device 5 of the present embodiment acquires an image of the face of a participant and the voice of the participant, and detects a shielded portion of a participant whose part of the face is shielded from the image. The remote conference execution device 5 estimates the content and emotion of the participant's speech from the image or voice. The remote conference execution device 5 deforms a partial image, which is an image of a shielded part of the participant's face, according to the content and emotion of the participant's speech. The remote conference execution device 5 generates a superimposed image in which the deformed partial image is superimposed on a range corresponding to the shielded part in the participant's image. Since it is possible to show other participants a superimposed image in which a partial image deformed according to the content and emotion of the participant's speech is superimposed, it becomes possible for other participants to grasp the expression of a participant who is participating in the remote conference with a part of the face shielded.

[0162] Next, with reference to FIG. 16, an operation example of the remote conference execution device 5 of the present embodiment will be described. FIG. 16 is a flowchart showing an operation example of the remote conference execution device 5.

[0163] The acquisition unit 51 acquires an image of the face of a participant and the voice of the participant by receiving them from the transmission terminal 6 (step S401).

[0164] The speaker identification unit 53 identifies a speaker who is speaking among a plurality of participants based on the images and voices of the respective participants acquired by the acquisition unit 51 (step S402).

[0165] When the speaker cannot be identified (step S402, NO), the speaker identification unit 53 notifies the control information generation unit 37 that the speaker cannot be identified. Further, the detection unit 52 does not perform the operation of step S403. The estimation unit 33 does not perform the operation of step S404. The image deformation unit 34 does not perform the operations of steps S405 and S406. The superimposing unit 35 does not perform the operation of step S407.

[0166] When the speaker can be identified (step S402, YES), the speaker identification unit 53 associates the speaker range information, the participant identification information of the speaker, the image data, and the voice data and outputs them to the detection unit 52, the estimation unit 33, and the superimposition unit 35.

[0167] The detection unit 52 detects the shielded part of the speaker whose face part is shielded from the image in the range indicated by the speaker range information (step S403).

[0168] When the shielded part can be detected (step S403, YES), the detection unit 52 associates the partial information indicating the shielded part, the information indicating the range of the shielded part, and the participant identification information of the speaker and outputs them to the image deformation unit 34.

[0169] When the shielded part cannot be detected (step S403, NO), the detection unit 52 notifies the control information generation unit 37 that the shielded part cannot be detected. Also, the estimation unit 33 does not perform the operation in step S404. The image deformation unit 34 does not perform the operations in steps S405 and S406. The superimposition unit 35 does not perform the operation in step S407.

[0170] The estimation unit 33 estimates the content and emotion of the speaker's speech from the image of the part corresponding to the range indicated by the speaker range information or the voice acquired by the acquisition unit 51 (step S404). The estimation unit 33 associates the participant identification information of the speaker, the timing information, the estimated emotion information indicating the result of emotion estimation, and the estimated speech information indicating the result of speech estimation and outputs them to the image deformation unit 34.

[0171] The image deformation unit 34 reads out the partial image of the speaker from the partial image storage unit 36 based on the partial information and the participant identification information of the speaker (step S405).

[0172] The image deformation unit 34 deforms the partial image read from the partial image storage unit 36 according to the content and emotion of the speaker's speech (step S406). The image deformation unit 34 outputs to the superimposition unit 35 by associating the timing information, the partial image data indicating the deformed partial image, the information indicating the range of the shielding portion, and the participant identification information of the speaker.

[0173] The superimposition unit 35 generates a superimposed image by superimposing the partial image deformed by the image deformation unit 34 on a range corresponding to the shielding portion in the speaker's image (step S407). The superimposition unit 35 outputs the timing information, the participant identification information of the speaker with a part of the face shielded, and the superimposed image data of the superimposed image to the control information generation unit 37.

[0174] The control information generation unit 37 generates output control information (step S408). When the superimposed image data is input from the superimposition unit 35, the control information generation unit 37 performs the following operation in step S408. When the superimposed image data is input, the control information generation unit 37 generates output control information for displaying the superimposed image on the display unit (in this example, the display unit 44 of the receiving terminal 4) at the timing of the voice for which the content and emotion of the speech have been estimated.

[0175] When notified that the speaker cannot be identified or the shielding portion cannot be detected, the control information generation unit 37 performs the following process in step S408. The control information generation unit 37 generates output control information based on the voice and image acquired by the acquisition unit 51. The control information generation unit 37 outputs the output control information to the transmission unit 38.

[0176] The transmission unit 38 transmits the output control information to the communication destination indicated by the communication destination information (step S409).

[0177] As described above, the remote conference execution device 5 of the present embodiment acquires an image of the face of a participant and the voice of the participant, and detects a shielded portion of a participant whose part of the face is shielded from the image. The remote conference execution device 5 estimates the content and emotion of the participant's speech from the image or voice. The remote conference execution device 5 deforms a partial image, which is an image of a shielded part of the participant's face, according to the content and emotion of the participant's speech. The remote conference execution device 5 generates a partial image in which the deformed partial image is superimposed on a range corresponding to the shielded part in the participant's image. Since it is possible to show other participants a superimposed image in which a partial image deformed according to the content and emotion of the participant's speech is superimposed, it becomes possible for other participants to grasp the emotion of a participant who is participating in the remote conference with part of their face shielded.

[0178] The remote conference execution device 5 of the present embodiment identifies a speaker who is a participant speaking among a plurality of participants based on the images and voices of each participant acquired by the acquisition unit 51. The remote conference execution device 5 acquires a partial face image corresponding to the detected shielded part of the identified speaker from the partial image storage unit 36 in which the partial images of the participants are stored. The remote conference execution device 5 deforms the partial image of the speaker according to the content and emotion of the speaker's speech, and superimposes the partial image on a range corresponding to the shielded part of the identified speaker. The remote conference execution device 5 of the present embodiment can identify the speaker even when the mouths of a plurality of participants are shielded. As a result, when a participant using the receiving terminal 4 of the remote conference views images of a plurality of participants with part of their faces shielded, the participant can easily identify the speaker. In addition, a participant using the receiving terminal 4 of the remote conference can easily grasp the expression of the speaker.

[0179] [Hardware Configuration Example] The procedures shown in each of the above embodiments can be realized by a remote conference execution program that causes an information processing device (computer) functioning as a remote conference execution device to realize these functions as the devices.

[0180] A configuration example of hardware resources for realizing each of the remote conferencing execution devices (1, 3, 5) in each of the above-described embodiments of the present invention using a single information processing device (computer) will be described. Note that the remote conferencing execution device may be realized using at least two information processing devices physically or functionally. Also, the remote conferencing execution device may be realized as a dedicated device. Further, only some functions of the remote conferencing execution device may be realized using an information processing device.

[0181] FIG. 17 is a diagram schematically showing a hardware configuration example of an information processing device capable of realizing the remote conferencing execution device according to each embodiment of the present invention. The information processing device 9 includes a communication interface 91, an input / output interface 92, an arithmetic unit 93, a storage device 94, a non-volatile storage device 95, and a drive device 96.

[0182] For example, the acquisition unit 11 of the remote conferencing execution device 1 in FIG. 1 can be realized by the communication interface 91 and the arithmetic unit 93. The detection unit 12, the estimation unit 13, the image deformation unit 14, and the superimposition unit 15 of the remote conferencing execution device 1 in FIG. 1 can be realized by the arithmetic unit 93.

[0183] The communication interface 91 is a communication means for the remote conferencing execution device of each embodiment to communicate with an external device by wire or / and wirelessly. Note that when the remote conferencing execution device is realized using at least two information processing devices, they may be connected so as to be able to communicate with each other via the communication interface 91.

[0184] The input / output interface 92 is a man-machine interface such as a keyboard which is an example of an input device and a display as an output device.

[0185] The arithmetic unit 93 is realized by an arithmetic processing unit such as a general-purpose CPU (Central Processing Unit) or a microprocessor, or a plurality of electric circuits. The arithmetic unit 93 can, for example, read various programs stored in the non-volatile storage device 95 into the storage device 94 and execute processing according to the read programs.

[0186] The storage device 94 is a memory device such as a RAM (Random Access Memory) that can be referenced by the arithmetic unit 93, and stores programs, various data, and the like. The storage device 94 may be a volatile memory device.

[0187] The non-volatile storage device 95 is a non-volatile storage device such as a ROM (Read Only Memory) or a flash memory, and can store various programs, data, and the like.

[0188] The drive device 96 is a device that processes, for example, reading and writing of data to the recording medium 97 described later.

[0189] The recording medium 97 is an arbitrary recording medium capable of recording data, such as an optical disk, a magneto-optical disk, or a semiconductor flash memory.

[0190] Each embodiment of the present invention may configure a remote conferencing execution device by the information processing device 9 illustrated in FIG. 17, for example. And each embodiment of the present invention may be realized by supplying a program capable of realizing the functions described in the above embodiments to this remote conferencing execution device.

[0191] In this case, it is possible to realize the embodiment by the arithmetic unit 93 executing the program supplied to the remote conferencing execution device. Also, it is possible to configure not all but some of the functions of the remote conferencing execution device by the information processing device 9.

[0192] Furthermore, the above program may be recorded on a recording medium 97, and the remote conference execution device may be configured such that the above program is appropriately stored in the non-volatile memory device 95 at the time of shipment or operation of the remote conference execution device. In this case, as the method for supplying the above program, a method of installing it inside the remote conference execution device using an appropriate jig may be adopted at the manufacturing stage before shipment or at the operation stage. Also, as the method for supplying the above program, a general procedure such as a method of downloading from the outside via a communication line such as the Internet may be adopted.

[0193] Note that each of the above-described embodiments is a preferred embodiment of the present invention, and various modifications can be made without departing from the gist of the present invention.

[0194] As described above, the present invention has been described with reference to the embodiments, but the present invention is not limited to the above embodiments. Various changes can be made to the configuration and details of the present invention within the scope of the present invention that can be understood by those skilled in the art.

[0195] Some or all of the above embodiments may be described as follows in the appended claims, but are not limited thereto.

[0196] (Appended Claim 1) A computer, an acquisition function for acquiring an image of a participant's face and the participant's voice, a detection function for detecting a shielded portion of the participant with a part of the face shielded from the image, an estimation function for estimating the content and emotion of the participant's speech from the image or the voice, an image deformation function for deforming a partial image, which is an image of the part of the participant's face, according to the content and emotion of the participant's speech, a superimposing function for generating a superimposed image by superimposing the deformed partial image on a range corresponding to the shielded portion in the participant's image, A remote conference execution program for realizing the above.

[0197] (Appended Claim 2) The estimation function estimates the emotion based on the analysis result of the voice or the analysis result of the change in the unobscured part of the face image of the participant, The estimation function estimates the content of the utterance based on the analysis result of the voice, The image deformation function deforms the partial image based on the result of the estimation by the estimation function The remote conferencing execution program according to Supplementary Note 1.

[0198] (Supplementary Note 3) The image deformation function also deforms the unobscured part of the face image of the participant according to the content of the utterance and the emotion. The remote conferencing execution program according to Supplementary Note 1 or Supplementary Note 2.

[0199] (Supplementary Note 4) The detection function detects the obscured part of the participant whose mouth or eyes are obscured. The remote conferencing execution program according to any one of Supplementary Notes 1 to 3.

[0200] (Supplementary Note 5) Further provided is a speaker identification function for identifying a speaker who is a participant speaking among a plurality of the participants based on the images and voices of the respective participants acquired by the acquisition function, The detection function, the estimation function, the image deformation function, and the superimposition function execute processing on the speaker. The remote conferencing execution program according to any one of Supplementary Notes 1 to 4.

[0201] (Supplementary Note 6) Acquire an image of the face of the participant and the voice of the participant, Detect the obscured part of the participant whose part of the face is obscured from the image, Estimate the content and emotion of the utterance of the participant from the image or the voice, Deform a partial image, which is an image of the part of the face of the participant, according to the content of the utterance and the emotion of the participant, Generate a superimposed image by superimposing the deformed partial image on a range corresponding to the occluded portion in the image of the participant. Remote conference execution method.

[0202] (Appendix 7) Estimate the emotion based on the analysis result of the voice or the analysis result of the change in the unoccluded portion of the face image of the participant. Estimate the content of the speech based on the analysis result of the voice. Deform the partial image based on the result of the estimation. The remote conference execution method according to Appendix 6.

[0203] (Appendix 8) Also deform the unoccluded portion of the face image of the participant according to the content of the speech and the emotion. The remote conference execution method according to Appendix 6 or Appendix 7.

[0204] (Appendix 9) Detect the occluded portion of the participant whose mouth or eyes are occluded. The remote conference execution method according to any one of Appendices 6 to 8.

[0205] (Appendix 10) Based on the acquired images and voices of each participant, identify the speaker who is the participant speaking among the plurality of participants. The detection process, the estimation process, the image deformation process, and the superimposition process are executed for the speaker. The remote conference execution method according to any one of Appendices 6 to 9.

[0206] (Appendix 11) An acquisition means for acquiring an image of a participant's face and the voice of the participant, A detection means for detecting the occluded portion of the participant whose part of the face is occluded from the image, An estimation means for estimating the content and emotion of the speech of the participant from the image or the voice. Image deformation means for deforming a partial image, which is an image of the part of the face of the participant, according to the content and emotion of the speech of the participant; Superposition means for generating a superimposed image by superimposing the deformed partial image on a range corresponding to the shielded part in the image of the participant; A remote conference execution device comprising:

[0207] (Appendix 12) The estimation means estimates the emotion based on the analysis result of the voice or the analysis result of the change in the unshielded part of the face image of the participant. The estimation means estimates the content of the speech based on the analysis result of the voice. The image deformation means deforms the partial image based on the result of the estimation by the estimation means. The remote conference execution device according to Appendix 11.

[0208] (Appendix 13) The image deformation means also deforms the unshielded part of the face image of the participant according to the content and emotion of the speech. The remote conference execution device according to Appendix 11 or Appendix 12.

[0209] (Appendix 14) The detection means detects the shielded part of the participant whose mouth or eyes are shielded. The remote conference execution device according to any one of Appendices 11 to 13.

[0210] (Appendix 15) The remote conference execution device further comprises speaker identification means for identifying a speaker, who is a participant speaking among a plurality of participants, based on the images and voices of each participant acquired by the acquisition means. The detection means, the estimation means, the image deformation means, and the superposition means execute processing on the speaker. The remote conference execution device according to any one of Appendices 11 to 14.

[0211] This application claims priority based on Japanese Patent Application No. 2021-200592 filed on December 10, 2021, and incorporates the entire disclosure thereof herein.

Explanation of Reference Numerals

[0212] 1, 3, 5 Remote Conference Execution Device 11, 31, 51 Acquisition Unit 12, 32, 52 Detection Unit 13, 33 Estimation Unit 331 Emotion Estimation Unit 332 Speech Estimation Unit 333 Output Unit 14, 34 Image Transformation Unit 15, 35 Superimposition Unit 36 Partial Image Storage Unit 37 Control Information Generation Unit 38 Transmission Unit 39 Conference Information Storage Unit 53 Speaker Identification Unit 54 Feature Quantity Storage Unit 4 Receiving Terminal 41 Receiving Unit 42 Output Control Unit 43 Voice Output Unit 44 Display Unit 2, 6 Transmitting Terminal 21 Photographing Unit 22 Voice Input Unit 23 Transmission Unit 61 Transceiving Unit 7 Photographing Device 8 Voice Input Device 9 Information Processing Device 91 Communication Interface 92 Input / Output Interface 93 Arithmetic Unit 94 Storage Device 95 Non-Volatile Storage Device 96 Drive Device 97 Recording Medium

Claims

1. A computer has an acquisition function for acquiring an image of a participant's face and the participant's voice, a detection function for detecting a shielded part of the participant from the image where a part of the face is shielded, an estimation function for estimating the content and emotion of the participant's speech from the image or the voice, an image deformation function for deforming a partial image, which is an image of the part of the participant's face, according to the content and emotion of the participant's speech, a superimposition function for generating a superimposed image by superimposing the deformed partial image on a range corresponding to the shielded part in the participant's image, A remote conference execution program for realizing the above.

2. The estimation function estimates the emotion based on the analysis result of the voice or the analysis result of the change in the unshielded part of the participant's face image, The estimation function estimates the content of the speech based on the analysis result of the voice, The image deformation function deforms the partial image based on the result of the estimation by the estimation function The remote conference execution program according to claim 1.

3. The image deformation function also deforms the unshielded part of the participant's face image according to the content and emotion of the speech The remote conference execution program according to claim 1 or claim 2.

4. The detection function detects the shielded part of the participant whose mouth or eyes are shielded The remote conference execution program according to any one of claims 1 to 3.

5. Further provided is a speaker identification function for identifying a speaker, who is a participant speaking among a plurality of participants, based on the images and voices of each participant acquired by the acquisition function, The detection function, the estimation function, the image deformation function, and the superimposition function execute processing on the speaker. The remote conference execution program according to any one of claims 1 to 4.

6. A computer acquires an image of a participant's face and the participant's voice, detects a shielded portion of the participant with a part of the face shielded from the image, estimates the content and emotion of the participant's speech from the image or the voice, deforms a partial image that is an image of the part of the participant's face according to the content and emotion of the participant's speech, generates a superimposed image by superimposing the deformed partial image on a range corresponding to the shielded portion in the participant's image. A remote conference execution method.

7. estimates the emotion based on the analysis result of the voice or the analysis result of the change in the unshielded part of the participant's face image, estimates the content of the speech based on the analysis result of the voice, deforms the partial image based on the result of the estimation The remote conference execution method according to claim 6.

8. also deforms the unshielded part of the participant's face image according to the content and emotion of the speech, The remote conference execution method according to claim 6 or claim 7.

9. detects the shielded portion of the participant with the mouth or eyes shielded The remote conference execution method according to any one of claims 6 to 8.

10. an acquisition means for acquiring an image of a participant's face and the participant's voice, a detection means for detecting a shielded portion of the participant with a part of the face shielded from the image, an estimation means for estimating the content and emotion of the participant's speech from the image or the voice, Image deformation means for deforming a partial image, which is an image of the part of the face of the participant, according to the content and emotion of the speech of the participant; Superposition means for generating a superimposed image by superimposing the deformed partial image on a range corresponding to the occluded part in the image of the participant; A remote conferencing execution device comprising the above.

Citation Information

Patent Citations

  • Communication conference system

    JP2000020683A

  • Substitute image display and tv phone apparatus

    JP2003037826A

  • Conference system, conference method and program

    JP2014225801A

  • Information processing device, information processing method, and program

    JP2018109924A

  • Speaker detection system, speaker detection method, and program

    JP2020155944A