Remote conferencing support program, remote conferencing support device, and remote conferencing support method
The remote conferencing support program addresses communication disruptions by enabling participants to use predefined text or audio signals when privacy mode is on, ensuring smooth and timely participation while maintaining privacy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KK TOSHIBA
- Filing Date
- 2023-04-03
- Publication Date
- 2026-04-13
AI Technical Summary
In remote conferencing systems, participants who disable their microphones and cameras for privacy concerns cannot communicate smoothly with others, and there is uncertainty about their consent for responding, leading to disrupted communication.
A remote conferencing support program that includes an acquisition function to capture media signals, a detection function to identify predetermined keyword or gesture inputs, and a transmission function to send corresponding text or audio signals instead of live voice or video when privacy mode is enabled.
Enables smooth communication by allowing participants to respond with pre-defined text or audio signals without live transmission, maintaining privacy and ensuring timely participation without disrupting others.
Smart Images

Figure 0007844384000001 
Figure 0007844384000002 
Figure 0007844384000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a remote conference support program, a remote conference support device, and a remote conference support method.
Background Art
[0002] In a remote conference system, each participant who is geographically separated from each other communicates audio and video using their own communication terminal. Each participant may set the microphone and camera mounted on their communication terminal to off during the conference from the perspective of privacy protection and the like.
[0003] However, while the participant has disabled the microphone and camera, the participant cannot communicate smoothly with other participants. For example, when the participant is asked for an answer by another participant, the participant cannot immediately reply. On the other hand, other participants cannot determine whether the consent of the participant has been obtained. Therefore, a remote conference system that ensures the privacy of each participant and promotes smooth communication between each participant is desired.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The problem to be solved by the present invention is to promote smooth communication in a remote conference.
Means for Solving the Problems
[0006] The remote conferencing support program according to the embodiment provides a computer with an acquisition function, a detection function, and a transmission function. The acquisition function acquires media signals related to the user's voice or video, and control information for the media signals. The detection function detects a detection signal corresponding to the signal model from the media signals by applying a signal model to the media signals. The transmission function transmits the detection signal or a media file corresponding to the detection signal to the external device if the media signals are not transmitted to the external device due to the control information. [Brief explanation of the drawing]
[0007] [Figure 1] A block diagram showing an example configuration of a remote conferencing system according to the first embodiment. [Figure 2] A block diagram showing an example of the functional configuration of a communication terminal according to the first embodiment. [Figure 3] A block diagram showing an example of the functional configuration of the transmission control unit according to the first embodiment. [Figure 4] A figure showing a first example of the display screen of a communication terminal according to the first embodiment. [Figure 5] A diagram showing an example of keyword information according to the first embodiment. [Figure 6] This figure shows a second example of the display screen of a communication terminal according to the first embodiment. [Figure 7] A block diagram showing an example of the functional configuration of the transmission control unit according to the second embodiment. [Figure 8] A diagram showing an example of a keyword list according to the second embodiment. [Figure 9] A diagram showing an example of an audio input signal according to the second embodiment. [Figure 10] A block diagram showing an example of the functional configuration of the transmission control unit according to the third embodiment. [Figure 11] A diagram showing an example of gesture information according to the third embodiment. [Figure 12] A diagram showing an example of the display screen of a communication terminal according to the third embodiment. [Figure 13] A block diagram showing an example of the functional configuration of the transmission control unit according to the fourth embodiment. [Figure 14] A diagram showing an example of operation information according to the fourth embodiment. [Figure 15] A diagram showing an example of the display screen of a communication terminal according to the fourth embodiment. [Figure 16] A block diagram showing an example configuration of a signal processing device according to the fifth embodiment. [Figure 17] A flowchart showing an example of the operation of the signal processing device according to the fifth embodiment. [Modes for carrying out the invention]
[0008] The following description of the remote conferencing support program, remote conferencing support device, and remote conferencing support method according to the embodiments will be given with reference to the drawings. In the following embodiments, parts with the same reference numerals perform similar operations, and redundant explanations will be omitted as appropriate.
[0009] (First Embodiment) Figure 1 is a block diagram showing an example configuration of a remote conferencing system 100 according to the first embodiment. The remote conferencing system 100 is a system for conducting remote conferencing. The remote conferencing system 100 includes a remote conferencing device 101, an internet 102, and a plurality of communication terminals 103. The remote conferencing device 101 and the communication terminals 103 are connected to each other via the internet 102 so that they can communicate with one another.
[0010] The remote conferencing device 101 is a device for conducting remote conferencing. The remote conferencing device 101 functions as a server in the remote conferencing system 100. The remote conferencing device 101 may also be a workstation capable of performing high-speed information processing. The remote conferencing device 101 is connected to the internet 102 by wire or wireless connection. The remote conferencing device 101 receives transmission data T (e.g., video, audio, text) sent from the communication terminal 103 via the internet 102. The remote conferencing device 101 processes the received transmission data T as needed and then transmits the processed transmission data T to the communication terminal 103.
[0011] The communication terminal 103 is a terminal that communicates various types of data or information with the remote conferencing device 101. The communication terminal 103 functions as a client in the remote conferencing system 100. The communication terminal 103 may be a personal computer (PC), a tablet terminal, or a smartphone. The communication terminal 103 is connected to the Internet 102 by wire or wirelessly. The communication terminal 103 transmits the transmission data T regarding the user to the remote conferencing device 101 via the Internet 102. The communication terminal 103 receives the processed transmission data T transmitted from the remote conferencing device 101 as the reception data R. The communication terminal 103 presents the received reception data R to the user in a predetermined manner. The communication terminal 103 is an example of a "remote conferencing support device".
[0012] FIG. 2 is a block diagram showing a functional configuration example of the communication terminal 103 according to the first embodiment. The communication terminal 103 includes a communication unit 201, a transmission control unit 202, a video input unit 203A, an operation input unit 203B, an audio input unit 203C, a reception control unit 204, a video output unit 205A, a text output unit 205B, and an audio output unit 205C.
[0013] The communication unit 201 establishes communication with the remote conferencing device 101 via the Internet 102. After the establishment of communication, the communication unit 201 transmits the transmission data T output from the transmission control unit 202 to the remote conferencing device 101. The communication unit 201 receives the processed transmission data T transmitted from the remote conferencing device 101 as the reception data R, and outputs the received reception data R to the reception control unit 204. The communication unit 201 is an example of a transmission unit or a reception unit.
[0014] The transmission control unit 202 generates the transmission data T by selectively or processing various input signals (e.g., video input signal VI, operation input signal OI, audio input signal AI) as necessary. The transmission control unit 202 outputs the generated transmission data T to the communication unit 201. The transmission control unit 202 is an example of a detection unit (see FIG. 3).
[0015] The video input unit 203A generates a video input signal VI by acquiring the user and the user's background video input from the camera. The video input unit 203A outputs the generated video input signal VI to the transmission control unit 202. The camera may be a built-in camera mounted on the communication terminal 103 or an external camera connected to the communication terminal 103. The video input unit 203A is an example of an acquisition unit.
[0016] The operation input unit 203B generates an operation input signal OI by acquiring user operation input from an input device. The operation input unit 203B outputs the generated operation input signal OI to the transmission control unit 202. The operation input signal OI may be a signal related to mouse movement or click, key input from a keyboard, tap or flick from a touchscreen, or pen input from a pen tablet. The input device may be a built-in input device mounted on the communication terminal 103, or an external input device connected to the communication terminal 103. The operation input unit 203B is an example of an acquisition unit.
[0017] The voice input unit 203C generates a voice input signal AI by acquiring the user's voice input from the microphone and ambient sounds around the user. The voice input unit 203C outputs the generated voice input signal AI to the transmission control unit 202. The microphone may be a built-in microphone mounted on the communication terminal 103 or an external microphone connected to the communication terminal 103. The voice input unit 203C is an example of an acquisition unit.
[0018] The receiving control unit 204 generates various output signals (e.g., video output signal VO, text output signal TO, audio output signal AO) by decomposing the received data R output from the communication unit 201. The receiving control unit 204 outputs the generated output signals to the video output unit 205A, text output unit 205B, or audio output unit 205C.
[0019] The video output unit 205A generates output video by processing or reconstructing the video output signal VO output from the receiving control unit 204 as needed. The video output unit 205A outputs the generated output video to a display device. This display device may be a built-in display device mounted on the communication terminal 103, or an external display device connected to the communication terminal 103.
[0020] The text output unit 205B outputs predetermined text to a display device based on the text output signal TO output from the reception control unit 204. Similarly, the display device may be a built-in display device mounted on the communication terminal 103, or an external display device connected to the communication terminal 103.
[0021] The audio output unit 205C generates output audio by processing or reconstructing the audio output signal AO output from the receiving control unit 204 as needed. The audio output unit 205C outputs the generated output audio to an audio device. This audio device may be a built-in audio device mounted on the communication terminal 103, or an external audio device connected to the communication terminal 103.
[0022] Figure 3 is a block diagram showing an example of the functional configuration of the transmission control unit 202 according to the first embodiment. The transmission control unit 202 includes an operation detection unit 301, a video control unit 302, an audio control unit 303, a keyword utterance detection unit 304, a keyword model storage unit 305, a keyword output control unit 306, a keyword information storage unit 307, and an integration unit 308.
[0023] The operation detection unit 301 generates video control information VC, transmission text information TX, and voice control information AC by decomposing the operation input signal OI output from the operation input unit 203B. The operation detection unit 301 outputs the generated video control information VC to the video control unit 302. The operation detection unit 301 outputs the generated transmission text information TX to the integration unit 308. The operation detection unit 301 outputs the generated voice control information AC to the voice control unit 303 and the keyword output control unit 306.
[0024] The video control unit 302 controls the output of the video input signal VI output from the video input unit 203A according to the video control information VC output from the operation detection unit 301. When the video control information VC is "ON", the video control unit 302 outputs the video input signal VI to the integration unit 308. Conversely, when the video control information VC is "OFF", the video control unit 302 stops outputting the video input signal VI. In other words, the video control unit 302 functions as a "gate" that determines whether or not to output the video input signal VI.
[0025] The voice control unit 303 controls the output of the voice input signal AI output from the voice input unit 203C according to the voice control information AC output from the operation detection unit 301. When the voice control information AC is "ON", the voice control unit 303 outputs the voice input signal AI to the integration unit 308. Conversely, when the voice control information AC is "OFF", the voice control unit 303 stops outputting the voice input signal AI. In other words, the voice control unit 303 functions as a "gate" that determines whether or not to output the voice input signal AI.
[0026] The keyword utterance detection unit 304 detects an audio signal corresponding to a predetermined keyword utterance from the audio input signal AI output from the audio input unit 203C by applying a keyword model stored in the keyword model storage unit 305 to the audio input signal AI output from the audio input unit 203C. When the keyword utterance detection unit 304 detects such an audio signal, it outputs an ID corresponding to the detected audio signal to the keyword output control unit 306.
[0027] The keyword model storage unit 305 stores the keyword model applied by the keyword utterance detection unit 304. The keyword model may be a machine learning model (e.g., linear regression, logistic regression, random forest, decision tree, k-nearest neighbors, support vector machine, naive Bayes, regularization, neural network). For example, the keyword model detects each keyword included in a given keyword utterance. In particular, the keyword model may detect the phoneme sequence or syllable sequence that constitutes the pronunciation of each keyword and determine the presence or absence of the given keyword utterance based on the detected phoneme sequence or syllable sequence. The structure of the neural network may be a known structure (e.g., fully connected, convolutional, recurrent).
[0028] The keyword model may be pre-trained with training data. This training data may be a speech corpus containing a large vocabulary, or it may be speech data collecting keyword utterances related to typical keywords. This speech data may include keyword utterances made by users of the communication terminal 103. Of course, a pre-trained keyword model may be re-trained with new training data. The pre-trained keyword model can detect user keyword utterances with high accuracy.
[0029] First, the keyword output control unit 306 reads the transmission text information TX corresponding to the ID output from the keyword utterance detection unit 304 from the keyword information storage unit 307. Second, the keyword output control unit 306 controls the output of the read transmission text information TX according to the voice control information AC output from the operation detection unit 301. If the voice control information AC is "ON", the keyword output control unit 306 stops the output of the transmission text information TX. Conversely, if the voice control information AC is "OFF", the keyword output control unit 306 outputs the transmission text information TX to the integration unit 308.
[0030] In other words, when the voice control information AC is "ON", the voice control unit 303 outputs the voice input signal AI to the integration unit 308, and the keyword output control unit 306 does not output the transmission text information TX. Conversely, when the voice control information AC is "OFF", the voice control unit 303 does not output the voice input signal AI, and the keyword output control unit 306 outputs the transmission text information TX to the integration unit 308. As a result, depending on whether the voice control information AC is "ON" or "OFF", either the voice input signal AI or the transmission text information TX is output.
[0031] The keyword information storage unit 307 stores keyword information that associates the pronunciation and transmitted text information TX corresponding to the ID output from the keyword utterance detection unit 304 (see Figure 5).
[0032] The integration unit 308 generates transmission data T by integrating the transmission text information TX output from the operation detection unit 301 or the keyword output control unit 306, the video input signal VI output from the video control unit 302, and the audio input signal AI output from the audio control unit 303. The integration unit 308 outputs the generated transmission data T to the communication unit 201.
[0033] Figure 4 shows a first example of the display screen of the communication terminal 103 according to the first embodiment. In the following, we assume that four users (S, Y, K, and T) are participating in a remote conference using their own communication terminals 103. Display screen 400A shows the display screen of user S's communication terminal 103. Display screen 400B shows the display screen of user T's communication terminal 103.
[0034] Display screens 400A and 400B show the window 401 of the remote conferencing application. Window 401 includes video control buttons 402, audio control buttons 403, display name 404, audio stop mark 405, video stop mark 406, participant video 407, conference chat display field 408, and conference chat input field 409.
[0035] The video control button 402 is a button used by the user to switch whether or not to transmit video from their communication terminal 103. The user can switch between "video transmission state" and "video stopped state" by toggling the video control button 402 with a click or other operation. This switches the video control information VC to "ON" or "OFF".
[0036] The voice control button 403 is a button that allows the user to switch whether or not to transmit voice from their communication terminal 103. The user can switch between "voice transmission state" and "voice stop state" by toggling the voice control button 403 by clicking or other means. This switches the voice control information AC to "ON" or "OFF".
[0037] Display screen 400A shows the "video transmission status" and "audio transmission status". In this state, the video control information VC and audio control information AC are "ON", and user S's video and audio are being transmitted to other users. Display screen 400B shows the "video stopped status" and "audio stopped status". In this state, the video control information VC and audio control information AC are "OFF", and user T's video or audio is not being transmitted to other users.
[0038] Display name 404 indicates the name registered by the user in advance. Audio stop mark 405 indicates that the user is in "audio stop state". Video stop mark 406 indicates that the user is in "video stop state". Participant video 407, when the user is in "video transmission state", displays the video transmitted by that user instead of video stop mark 406.
[0039] The conference chat display area 408 displays the text entered by the user in the conference chat input area 409, along with the user's display name 404. This allows the text entered by the user to be shared with other users.
[0040] The conference chat input field 409 is a field for users to enter text. Users enter the desired text into the conference chat input field 409 using keyboard input or other methods. The entered text is output as the transmitted text information TX.
[0041] Figure 5 shows an example of keyword information according to the first embodiment. In this example, table 200A registers six records as keyword information. For example, the record related to ID "1" includes "OK" as the transmission text information TX corresponding to the pronunciation "ok desu". Similarly, each of the records related to IDs "2" to "6" includes a unique pronunciation and a unique transmission text information TX.
[0042] Keyword information may be selected, edited, or registered by the user of the communication terminal 103. The transmitted text information TX may include attributes related to text formatting (e.g., size, font, color) in a format such as HTML (HyperText Markup Language). Furthermore, instead of the transmitted text information TX, an ID or URL (Uniform Resource Locator) of an image or video that can be commonly referenced by each communication terminal 103 may be registered.
[0043] Figure 6 shows a second example of the display screen of the communication terminal 103 according to the first embodiment. In the following, we assume that user S says to other users (Y, K, T), "Is this okay, everyone?" In response to this question, we assume that user Y says, "It's fine," user K says, "Okay," and user T says, "Good." Display screens 400C and 400D show the display screen of user S's communication terminal 103 in the above case.
[0044] According to display screens 400C and 400D, user S's voice control information AC is "ON," so user S's speech is transmitted to and played back by other users. Similarly, user K's voice control information AC is "ON," so user K's speech is transmitted to and played back by other users. On the other hand, user (Y, T)'s voice control information AC is "OFF," so user (Y, T)'s speech is not transmitted to other users.
[0045] At this time, user Y's communication terminal 103 operates as follows: The keyword utterance detection unit 304 detects "2" as the ID corresponding to user Y's utterance "It's alright". The keyword output control unit 306 refers to table 200A and outputs "It's alright" as the transmission text information TX corresponding to ID "2". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0046] Meanwhile, user T's communication terminal 103 operates as follows: The keyword utterance detection unit 304 detects "4" as the ID corresponding to user T's utterance "Like". The keyword output control unit 306 refers to table 200A and outputs "Like!" as the transmission text information TX corresponding to ID "4". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0047] As a result of the above operation, as shown on display screen 400C, the utterances of users (Y, T) are displayed in the conference chat display field 408 of each user's communication terminal 103. The conference chat display field 408 displays the user's (Y, T) display name 404 and the transmitted text information TX.
[0048] Alternatively, as shown on display screen 400D, the user's (Y, T) utterance is superimposed as a box 450 on the user's (Y, T) video stop mark 406 (or participant video 407). The box 450 contains transmitted text information TX and is displayed for a predetermined time. This allows each user to intuitively understand which user made the utterance. If the transmitted text information TX is an image ID or URL, the image may be superimposed on the video stop mark 406 (or participant video 407). By displaying an image, the user can express emotions or nuances in addition to linguistic information.
[0049] According to the first embodiment described above, when the voice control information AC is "OFF", the user's communication terminal 103 detects a predetermined keyword utterance from the user's voice input signal AI. Instead of the voice input signal AI, the communication terminal 103 transmits the text information TX corresponding to the detected keyword utterance as transmission data T to the remote conferencing device 101.
[0050] Therefore, users (Y, T) whose voice control information AC is "OFF" can immediately respond to user S who asked a question by uttering a predetermined keyword. Meanwhile, user S can quickly confirm agreement with other users (Y, K, T) and smoothly conduct the meeting. Furthermore, users (Y, T) do not feel any privacy concerns or worries about the meeting being disrupted due to the transmission of ambient sounds around them. Users (Y, T) do not need to take the time to switch the voice control button 403 to "ON" before speaking, so they can communicate their intentions to other users in a timely manner.
[0051] In addition, users can check Table 200A in advance to understand the pronunciation of detected keywords and the text information TX to be sent. This allows users to take precautions to prevent unintended utterances from being detected and unintended text information TX from being sent. In other words, users can participate in meetings with peace of mind.
[0052] Furthermore, the communication terminal 103 may display a confirmation window before sending the detected keyword. For example, the confirmation window may include the text, "Send a 'Like!'. Is that OK?", and GUI buttons "Yes" and "No". If the user selects the GUI button "Yes", the communication terminal 103 sends the "Like!". Conversely, if the user selects the GUI button "No", the communication terminal 103 does not send the "Like!". This reduces the risk of the communication terminal 103 sending a falsely detected keyword. In other words, users can participate in the meeting with greater peace of mind.
[0053] Furthermore, the communication terminal 103 may detect the user's emotions from the user's voice input signal AI. For example, the keyword model storage unit 305 stores an emotion model for detecting laughter or anger. The keyword utterance detection unit 304 applies the emotion model to the voice input signal AI and outputs ID "1" when laughter is detected, and ID "2" when anger is detected.
[0054] For example, the keyword information storage unit 307 registers "(laugh)" as the transmission text information TX corresponding to ID "1" and "(angry)" as the transmission text information TX corresponding to ID "2". The keyword output control unit 306 outputs the transmission text information TX corresponding to the ID output from the keyword utterance detection unit 304.
[0055] For example, consider a scenario where a user tells a joke, and another user laughs with the voice control information AC "OFF". In this case, the other user's laughter is detected, and the text "(laugh)" is displayed in the conference chat display area 408 along with the other user's display name 404. This allows the user who told the joke to be informed of the other user's reaction, facilitating smoother communication.
[0056] (Second Embodiment) Figure 7 is a block diagram showing an example of the functional configuration of the transmission control unit 202 according to the second embodiment. According to the second embodiment, the transmission control unit 202 stores a voice input signal AI for a certain period of time. When a predetermined keyword utterance is detected, the transmission control unit 202 reads the voice signal for the section containing this keyword utterance from the stored voice input signal AI and transmits it.
[0057] The transmission control unit 202 includes an operation detection unit 301, a video control unit 302, an audio control unit 303, a keyword utterance detection unit 501, a keyword model storage unit 502, a keyword output control unit 503, an input audio storage unit 504, and an integration unit 308.
[0058] The keyword utterance detection unit 501 detects speech segment information corresponding to a predetermined keyword utterance from the speech input signal AI output from the speech input unit 203C by applying a keyword model stored in the keyword model storage unit 502 to the speech input signal AI output from the speech input unit 203C. For example, if the start time of the utterance is "1.7 seconds ago" and the end time of the utterance is "0.3 seconds ago", the keyword utterance detection unit 501 outputs the speech segment information [1.7,0.3] to the keyword output control unit 503.
[0059] The keyword model storage unit 502 first stores the keyword model applied by the keyword utterance detection unit 501. The keyword model stored in the keyword model storage unit 502 is the same as the keyword model stored in the keyword model storage unit 305. Secondly, the keyword model storage unit 502 stores a list of keywords that the keyword model should detect (see Figure 8).
[0060] First, the keyword output control unit 503 reads the voice signal (detected voice signal DA) corresponding to the speech interval information output from the keyword speech detection unit 501 from the voice input signal AI stored in the input voice storage unit 504. Second, the keyword output control unit 503 controls the output of the detected voice signal DA according to the voice control information AC output from the operation detection unit 301. If the voice control information AC is "ON", the keyword output control unit 503 stops the output of the detected voice signal DA. Conversely, if the voice control information AC is "OFF", the keyword output control unit 503 outputs the detected voice signal DA to the integration unit 308.
[0061] In other words, when the voice control information AC is "ON", the voice control unit 303 outputs the voice input signal AI to the integration unit 308, and the keyword output control unit 503 does not output the detected voice signal DA. Conversely, when the voice control information AC is "OFF", the voice control unit 303 does not output the voice input signal AI, and the keyword output control unit 503 outputs the detected voice signal DA to the integration unit 308. As a result, depending on whether the voice control information AC is "ON" or "OFF", either the voice input signal AI or the detected voice signal DA is output.
[0062] The input voice storage unit 504 stores voice input signals AI from the current time up to a predetermined time prior, and sequentially updates its stored contents. The input voice storage unit 504 sets the storage time for the voice input signals AI based on the number of characters or syllables of the keywords stored in the keyword model storage unit 502. In particular, the input voice storage unit 504 sets the storage time so that it can store the entire voice of the keyword. Typically, the storage time is "4.0 seconds" (see Figure 9).
[0063] The integration unit 308 generates transmission data T by integrating the transmission text information TX output from the operation detection unit 301, the video input signal VI output from the video control unit 302, and the audio input signal AI output from the audio control unit 303 or the detected audio signal DA output from the keyword output control unit 503. The integration unit 308 outputs the generated transmission data T to the communication unit 201.
[0064] Figure 8 shows an example of a keyword list according to the second embodiment. In this example, table 200B registers six records as a keyword list. For example, the record related to ID "1" includes the pronunciation "ok desu". Similarly, each of the records related to IDs "2" to "6" includes a unique pronunciation.
[0065] Figure 9 shows an example of a voice input signal AI according to the second embodiment. In this example, the waveform data 500 represents the user's voice input signal AI over a certain period of time. The waveform data 500 is waveform data that covers an interval 510 from "4.0 seconds ago" to "0 seconds ago" (current time). For the waveform data 500, the horizontal axis represents time, and the vertical axis represents amplitude. The waveform data 500 is stored in the input voice storage unit 504.
[0066] For example, consider a case where the speech segment information corresponding to the utterance "It's alright" is [1.7,0.3]. In this case, the keyword output control unit 503 identifies the segment 520 within segment 510 that corresponds to the speech segment information [1.7,0.3]. The keyword output control unit 503 reads out the detected speech signal DA in the identified segment 520. Note that the keyword output control unit 503 may extend the segment in which it reads out the detected speech signal DA to account for errors in the speech segment information. In the above example, if the keyword output control unit 503 considers an error of "0.2 seconds", it reads out the detected speech signal DA in the segment from "1.9 seconds ago" to "0.1 seconds ago". This ensures that the entire detected speech signal DA is reliably detected.
[0067] Referring again to Figure 6, an example of the display screen of the communication terminal 103 according to the second embodiment will be described. In the following, as in the first embodiment, we will assume that user S says to other users (Y, K, T), "Is this okay, everyone?" In response to this question, we will assume that user Y says, "It's fine," user K says, "Okay," and user T says, "Good." Display screens 400C and 400D show the display screen of user S's communication terminal 103 in the above case.
[0068] According to display screens 400C and 400D, user S's voice control information AC is "ON," so user S's speech is transmitted to and played back by other users. Similarly, user K's voice control information AC is "ON," so user K's speech is transmitted to and played back by other users. On the other hand, user (Y, T)'s voice control information AC is "OFF."
[0069] At this time, user Y's communication terminal 103 operates as follows: The keyword utterance detection unit 501 detects utterance section information [1.7,0.3] corresponding to user Y's utterance "It's alright". The keyword output control unit 503 reads out the detected voice signal DA of section 520 corresponding to this utterance section information by referring to the waveform data 500. The integration unit 308 generates transmission data T including the detected voice signal DA and outputs the generated transmission data T to the communication unit 201.
[0070] Meanwhile, the user T's communication terminal 103 operates as follows: The keyword utterance detection unit 501 detects utterance section information corresponding to the user T's utterance "Like". The keyword output control unit 503 reads the detected voice signal DA for the section corresponding to this utterance section information from the input voice storage unit 504. The integration unit 308 generates transmission data T including the detected voice signal DA and outputs the generated transmission data T to the communication unit 201.
[0071] As a result of the above operation, on display screens 400C and 400D, the user's (Y, T) speech is transmitted to and played back to the other user.
[0072] According to the second embodiment described above, when the voice control information AC is "OFF", the user's communication terminal 103 detects a detection voice signal DA corresponding to a predetermined keyword utterance from the user's voice input signal AI. Instead of the voice input signal AI, the communication terminal 103 transmits the detection voice signal DA as transmission data T to the remote conferencing device 101.
[0073] Therefore, users (Y, T) whose voice control information AC is "OFF" can immediately respond to user S who asked a question by uttering a predetermined keyword, as if they had switched voice control information AC to "ON" and spoken. On the other hand, user S can quickly confirm agreement with other users (Y, K, T) and smoothly conduct the meeting. Because the intonation and tone of the voice uttered by users (Y, T) are transmitted, users (Y, T) can convey nuances to other users that cannot be conveyed through text or other linguistic information.
[0074] In addition, while the user (Y, T) is not speaking the specified keyword, ambient sounds around the user are not transmitted to other users. That is, during the above period, the user (Y, T) does not feel any privacy concerns or worries about the meeting being disrupted by the transmission of ambient sounds around them. The user (Y, T) does not need to take the trouble of switching the voice control button 403 to "ON" before speaking, so they can communicate their intentions to other users in a timely manner.
[0075] (Third embodiment) Figure 10 is a block diagram showing an example of the functional configuration of the transmission control unit 202 according to the third embodiment. According to the third embodiment, the transmission control unit 202 detects a predetermined gesture from the video input signal VI and transmits transmission text information TX corresponding to the detected gesture.
[0076] The transmission control unit 202 includes an operation detection unit 301, a video control unit 302, an audio control unit 303, a gesture detection unit 601, a gesture model storage unit 602, a gesture output control unit 603, a gesture information storage unit 604, and an integration unit 308.
[0077] The gesture detection unit 601 detects a gesture signal corresponding to a predetermined gesture from the video input signal VI output from the video input unit 203A by applying a gesture model stored in the gesture model storage unit 602 to the video input signal VI output from the video input unit 203A. When the gesture detection unit 601 detects the gesture signal, it outputs the ID corresponding to the detected gesture signal to the gesture output control unit 603.
[0078] The gesture model storage unit 602 stores the gesture model applied by the gesture detection unit 601. The gesture model may be a machine learning model, similar to the keyword model. For example, the gesture model may detect a sequence of poses that constitute a predetermined gesture and determine the presence or absence of the predetermined gesture based on the detected sequence of poses.
[0079] The gesture model may be pre-trained with training data. The training data may be a video corpus containing a large number of gestures, or video data that collects gesture videos related to typical gestures. This video data may include gesture videos made by the user of the communication terminal 103. Of course, the trained gesture model may be retrained with new training data. The trained gesture model can detect the user's gestures with high accuracy.
[0080] First, the gesture output control unit 603 reads the transmission text information TX corresponding to the ID output from the gesture detection unit 601 from the gesture information storage unit 604. Second, the gesture output control unit 603 controls the output of the read transmission text information TX according to the video control information VC output from the operation detection unit 301. If the video control information VC is "ON", the gesture output control unit 603 stops the output of the transmission text information TX. Conversely, if the video control information VC is "OFF", the gesture output control unit 603 outputs the transmission text information TX to the integration unit 308.
[0081] In other words, when the video control information VC is "ON", the video control unit 302 outputs the video input signal VI to the integration unit 308, and the gesture output control unit 603 does not output the transmission text information TX. Conversely, when the video control information VC is "OFF", the video control unit 302 does not output the video input signal VI, and the gesture output control unit 603 outputs the transmission text information TX to the integration unit 308. As a result, depending on whether the video control information VC is "ON" or "OFF", either the video input signal VI or the transmission text information TX is output.
[0082] The gesture information storage unit 604 stores gesture information that associates the content of the gesture corresponding to the ID output from the gesture detection unit 601 with the transmitted text information TX (see Figure 11).
[0083] The integration unit 308 generates transmission data T by integrating the transmission text information TX output from the operation detection unit 301 or the gesture output control unit 603, the video input signal VI output from the video control unit 302, and the voice input signal AI output from the voice control unit 303. The integration unit 308 outputs the generated transmission data T to the communication unit 201.
[0084] Figure 11 shows an example of gesture information according to the third embodiment. In this example, table 200C registers three records as gesture information. For example, the record for ID "1" includes "yes, yes" as the transmitted text information TX corresponding to the gesture "shake head up and down twice". Similarly, the records for IDs "2" and "3" each include a unique gesture and a unique transmitted text information TX.
[0085] Gesture information may be selected, edited, or registered by the user of the communication terminal 103. The transmitted text information TX may include attributes related to text formatting in a format such as HTML. Furthermore, instead of the transmitted text information TX, an ID or URL of an image or video that can be commonly referenced by each communication terminal 103 may be registered.
[0086] Figure 12 shows an example of the display screen of the communication terminal 103 according to the third embodiment. In the following, we assume that user S says to other users (Y, K, T), "Is this alright, everyone?" In response to this question, we assume that users (K, T) perform the action of "shaking their heads up and down twice," and user Y performs the action of "giving a thumbs-up and thrusting out a fist." Display screens 400E and 400F show the display screen of user S's communication terminal 103 in the above case.
[0087] According to display screens 400E and 400F, user S's voice control information AC is "ON," so user S's speech is transmitted to and played back by other users. On the other hand, user K's video control information VC is "ON," so user K's video is transmitted to and played back by other users. On the other hand, user (Y, T)'s video control information VC is "OFF," so user (Y, T)'s video is not transmitted to other users.
[0088] At this time, user Y's communication terminal 103 operates as follows: The gesture detection unit 601 detects "3" as the ID corresponding to user Y's action "thumbs up and fist raised". The gesture output control unit 603 refers to table 200C and outputs "Like!" as the transmission text information TX corresponding to ID "3". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0089] Meanwhile, user T's communication terminal 103 operates as follows: The gesture detection unit 601 detects "1" as the ID corresponding to user T's action "shakes head up and down twice". The gesture output control unit 603 refers to table 200C and outputs "yes, yes" as the transmission text information TX corresponding to ID "1". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0090] As a result of the above operation, the actions of users (Y, T) are displayed in the conference chat display field 408 of each user's communication terminal 103, as shown in the display screen 400E. The conference chat display field 408 displays the user's (Y, T) display name 404 and the sent text information TX.
[0091] Alternatively, as shown on display screen 400F, the actions of users (Y, T) are superimposed as gesture videos 460 onto the user's (Y, T) video stop marker 406 (or participant video 407). For example, user Y's video stop marker 406 will have a gesture video 460 of a thumbs-up and fist thrust superimposed. On the other hand, user T's video stop marker 406 will have a gesture video 460 of a head bobbing up and down twice superimposed. The gesture videos 460 are played for a predetermined amount of time. This allows each user to intuitively understand which user performed which gesture. By playing the gesture video 460, the user can communicate emotions or nuances to other users.
[0092] According to the third embodiment described above, when the video control information VC is "OFF", the user's communication terminal 103 detects a predetermined gesture from the user's video input signal VI. Instead of the video input signal VI, the communication terminal 103 transmits to the remote conferencing device 101 the transmission data T, which is transmission text information TX or the like corresponding to the detected gesture.
[0093] Therefore, users (Y, T) whose video control information VC is "OFF" can immediately respond to user S who asked a question by performing a predetermined gesture. Meanwhile, user S can quickly confirm agreement with other users (Y, K, T) and smoothly conduct the meeting. Furthermore, users (Y, T) do not feel any privacy concerns or worries that the transmission of video of their surroundings will disrupt the meeting. Users (Y, T) do not need to take the time to switch the video control button 402 to "ON" before performing a gesture, so they can communicate their intentions to other users in a timely manner.
[0094] In addition, users can check Table 200C in advance to understand the gestures that will be detected and the text information (TX) that will be sent. This allows users to take precautions to prevent unintended gestures from being detected and unintended text information (TX) from being sent. In other words, users can participate in meetings with peace of mind.
[0095] The communication terminal 103 may also detect the user's facial expression from the user's video input signal VI. For example, the gesture model storage unit 602 stores facial expression models for detecting a smiling face or an angry face. The gesture detection unit 601 applies the facial expression model to the video input signal VI and outputs ID "1" when a smiling face is detected, and ID "2" when an angry face is detected.
[0096] For example, the gesture information storage unit 604 registers "(laugh)" as the transmission text information TX corresponding to ID "1" and "(angry)" as the transmission text information TX corresponding to ID "2". The gesture output control unit 603 outputs the transmission text information TX corresponding to the ID output from the gesture detection unit 601.
[0097] For example, consider a scenario where a user tells a joke, and another user laughs while their video control information (VC) is set to "OFF". In this case, the other user's laughter is detected, and the text "(laugh)" is displayed in the conference chat display area 408 along with the other user's display name 404. This allows the user who told the joke to be informed of the other user's reaction, facilitating smoother communication.
[0098] (Fourth Embodiment) Figure 13 is a block diagram showing an example of the functional configuration of the transmission control unit 202 according to the fourth embodiment. According to the fourth embodiment, the transmission control unit 202 detects a predetermined operation pattern from the operation input signal OI and transmits transmission text information TX corresponding to the detected operation pattern.
[0099] The transmission control unit 202 includes an operation detection unit 701, an operation information storage unit 702, a video control unit 302, an audio control unit 303, and an integration unit 308.
[0100] First, the operation detection unit 701 generates video control signal VC, transmission text information TX, and voice control information AC by decomposing the operation input signal OI output from the operation input unit 203B. Second, the operation detection unit 701 detects an ID corresponding to a predetermined operation pattern from the operation input signal OI and outputs the transmission text information TX corresponding to the detected ID to the integration unit 308. That is, the transmission text information TX is output when (i) text is entered into the conference chat input field 409, and when (ii) a predetermined operation pattern is detected from the operation input signal OI.
[0101] The operation information storage unit 702 stores operation information that associates the operation pattern to be detected by the operation detection unit 701 from the operation input signal OI with the ID and transmission text information TX corresponding to that operation pattern (see Figure 14).
[0102] The integration unit 308 generates transmission data T by integrating the transmission text information TX output from the operation detection unit 701, the video input signal VI output from the video control unit 302, and the audio input signal AI output from the audio control unit 303. The integration unit 308 outputs the generated transmission data T to the communication unit 201.
[0103] Figure 14 shows an example of operation information according to the fourth embodiment. In this example, table 200D registers four records as operation information. For example, the record for ID "1" includes "OK" as the transmission text information TX corresponding to the operation pattern "Ctrl+O, Ctrl+K". Similarly, each of the records for IDs "2" to "4" includes a unique operation pattern and a unique transmission text information TX.
[0104] The operation pattern "Ctrl" refers to the control key on the keyboard, and the letters "O, K, D, J, A, R" refer to the respective keys on the keyboard. The operation pattern "ML" refers to the left mouse button, and "MR" refers to the right mouse button. The plus sign "+" means to perform the operations to the left and right of the sign simultaneously, and the comma sign "," means to perform the operation to the left of the sign, followed by the operation to the right of the sign. Furthermore, operation information may be selected, edited, or registered by the user of the communication terminal 103.
[0105] Table 200D may also register operation patterns using input devices other than the keyboard and mouse. For example, Table 200D may register operation patterns using a mouse pointer, or operation patterns using taps or flicks on a touchscreen.
[0106] Figure 15 shows an example of the display screen of the communication terminal 103 according to the fourth embodiment. In the following, we assume that user S says to other users (Y, K, T), "Is this alright, everyone?" In response to this question, we assume that user Y says, "It's fine," user K performs the operation "ML+MR," and user T performs the operation "Ctrl+O, Ctrl+K." Display screen 400G shows the display screen of user S's communication terminal 103 in the above case.
[0107] According to display screen 400G, user S's voice control information AC is "ON," so user S's speech is transmitted to and played back by other users. Similarly, user Y's voice control information AC is "ON," so user Y's speech is transmitted to and played back by other users. On the other hand, users (K, T)'s voice control information AC is "OFF."
[0108] At this time, user K's communication terminal 103 operates as follows: The operation detection unit 701 detects "4" as the ID corresponding to user K's operation "ML+MR". By referring to table 200D, the operation detection unit 701 outputs "Like!" as the transmission text information TX corresponding to ID "4". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0109] Meanwhile, user T's communication terminal 103 operates as follows: The operation detection unit 701 detects "1" as the ID corresponding to user T's operation "Ctrl+O, Ctrl+K". By referring to table 200D, the operation detection unit 701 outputs "OK" as the transmission text information TX corresponding to ID "1". The integration unit 308 generates transmission data T including the transmission text information TX and outputs the generated transmission data T to the communication unit 201.
[0110] As a result of the above operation, as shown on display screen 400G, the operations of the users (K, T) are displayed in the conference chat display field 408 of each user's communication terminal 103. The conference chat display field 408 displays the user's (K, T) display name 404 and the sent text information TX.
[0111] According to the fourth embodiment described above, when the voice control information AC is "OFF", the user's communication terminal 103 detects a predetermined operation pattern from the user's operation input signal OI. Instead of the voice input signal AI, the communication terminal 103 transmits the transmission text information TX, etc., corresponding to the detected operation pattern as transmission data T to the remote conferencing device 101.
[0112] Therefore, users (K, T) whose voice control information AC is "OFF" can immediately respond to user S who made a request by inputting a predetermined operation pattern. Meanwhile, user S can quickly confirm agreement with other users (Y, K, T) and smoothly conduct the meeting. Furthermore, users (K, T) do not feel any privacy concerns or worries about the meeting being disrupted due to the transmission of ambient sounds around them. Users (K, T) do not need to take the trouble to type text into the meeting chat input field 409, so they can communicate their intentions to other users in a timely manner. In particular, since the communication terminal 103 uses user operation input instead of a signal model, the risk of false detection can be reduced.
[0113] According to the first to fourth embodiments described above, the communication terminal 103 detects a detection signal corresponding to a signal model or operation pattern from one of the three types of input signals (video input signal VI, operation input signal OI, and voice input signal AI). Not limited to this example, the communication terminal 103 may detect multiple detection signals in any combination of these three types of input signals. This allows the user to choose a method that is easy for them to use and to communicate their intentions to other users in the chosen method, thereby promoting smoother communication.
[0114] (Fifth embodiment) Figure 16 is a block diagram showing an example configuration of the signal processing device 800 according to the fifth embodiment. The signal processing device 800 is a device that processes various signals. The signal processing device 800 may be a personal computer (PC), a tablet terminal, or a smartphone. The signal processing device 800 may be mounted on the communication terminal 103, or it may be the communication terminal 103 itself. The signal processing device 800 is an example of a "remote conferencing support device".
[0115] The signal processing device 800 includes, as its components, a processing circuit 81, a storage device 82, an input device 83, an output device 84, and a communication device 85. Each component is connected to the others in a communicative manner via a common signal communication path, which is a bus (BUS).
[0116] The processing circuit 81 is a circuit that controls the overall operation of the signal processing device 800. The processing circuit 81 includes at least one processor. A processor refers to a circuit such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or a programmable logic device (e.g., a Simple Programmable Logic Device (SPLD), a Complex Programmable Logic Device (CPLD), or a Field Programmable Gate Array (FPGA)). If the processor is a CPU, the CPU realizes each function by reading and executing each program stored in the memory device 82. If the processor is an ASIC, each function is directly incorporated into the ASIC as a logic circuit. The processor may be configured as a single circuit or as a combination of multiple independent circuits. The processing circuit 81 realizes each part (acquisition unit 811, detection unit 812, transmission unit 813, and system control unit 814). Processing circuit 81 is an example of a processing unit.
[0117] The acquisition unit 811 acquires various types of data or information. Firstly, the acquisition unit 811 acquires media signals related to the user's voice or video (e.g., video input signal VI, audio input signal AI). Secondly, the acquisition unit 811 acquires control information for the media signals (e.g., video control information VC, audio control information AC).
[0118] The detection unit 812 detects various types of data or information. For example, the detection unit 812 applies a signal model (e.g., keyword model, emotion model, gesture model, facial expression model) to the media signal acquired by the acquisition unit 811, thereby detecting a detection signal (e.g., voice signal, emotion signal, gesture signal, facial expression signal) corresponding to the signal model from the media signal.
[0119] The transmitting unit 813 transmits various types of data or information. For example, depending on the control information acquired by the acquisition unit 811, the transmitting unit 813 transmits the media signal acquired by the acquisition unit 811 or the detection signal detected by the detection unit 812 to an external device. If the control information is "ON", the transmitting unit 813 transmits the media signal to the external device. Conversely, if the control information is "OFF", the transmitting unit 813 transmits the detection signal or a media file (e.g., text, image, music, audio, video) corresponding to the detection signal to the external device.
[0120] The system control unit 814 has the function of controlling various operations performed by the processing circuit 81. For example, the system control unit 814 provides an operating system (OS) for the processing circuit 81 to implement each part (acquisition unit 811, detection unit 812, transmission unit 813).
[0121] The storage device 82 stores various types of data or information. The storage device 82 may be a storage medium readable by the processor (e.g., magnetic storage medium, electromagnetic storage medium, optical storage medium, semiconductor memory), or it may be a drive device that reads or writes data or information to and from the storage medium. The storage device 82 stores programs that enable the processing circuit 81 to implement each part (acquisition unit 811, detection unit 812, transmission unit 813, system control unit 814). The storage device 82 may also store various types of signals (media signals, detection signals) or media files. The storage device 82 is an example of a storage unit.
[0122] The input device 83 is a device that inputs various types of data or information to the signal processing device 800. The input device 83 may be a mouse, keyboard, buttons, panel switches, slider switches, trackballs, control panels, touchscreens, pen tablets, cameras, or microphones. The input device 83 is an example of an input unit.
[0123] The output device 84 is a device that outputs various types of data or information. The output device 84 may be a display, a speaker, or an earphone. If the output device 84 is a display, the display may accept various operations on the data or information displayed via GUI buttons or the like. The output device 84 is an example of an output unit, display unit, or sound unit.
[0124] The communication device 85 is a device that communicates various types of data or information with an external device. The external device may be a remote conferencing device 101. The communication device 85 is an example of a communication unit.
[0125] The processing circuit 81, storage device 82, input device 83, output device 84, or communication device 85 may each represent a part of the communication terminal 103 according to the first to fourth embodiments.
[0126] Figure 17 is a flowchart showing an example of operation of the signal processing device 800 according to the fifth embodiment. This example of operation may be started in response to a start command from the user.
[0127] (Step S1) First, the signal processing device 800 acquires media signals and control information using the acquisition unit 811. Specifically, the acquisition unit 811 acquires media signals and control information from the input device 83.
[0128] (Step S2) Next, the signal processing device 800 uses the detection unit 812 to detect a detection signal from the media signal acquired in step S1. Specifically, the detection unit 812 applies a signal model to the media signal to detect a detection signal corresponding to the signal model from the media signal.
[0129] (Step S3) Here, the signal processing device 800 determines the signal state of the control information acquired in step S1 by the transmitting unit 813. If the signal state is "ON" (step S3-ON), the process proceeds to step S4A. If the signal state is "OFF" (step S3-OFF), the process proceeds to step S4B.
[0130] (Step S4A) In this case, the signal processing device 800 transmits the media signal acquired in step S1 to an external device via the transmission unit 813. After step S4A, the signal processing device 800 completes its series of operations.
[0131] (Step S4B) In this case, the signal processing device 800 transmits the detection signal detected in step S2, or the media file corresponding to the detection signal, to an external device via the transmission unit 813. After step S4B, the signal processing device 800 completes the series of operations.
[0132] According to the fifth embodiment described above, the signal processing device 800 can perform the same operations as the communication terminal 103 according to the first to fourth embodiments. In other words, the signal processing device 800 can achieve the same effects as those achieved by the operation of the communication terminal 103.
[0133] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of symbols]
[0134] 81…Processing circuit, 82…Storage device, 83…Input device, 84…Output device, 85…Communication device, 100…Remote conferencing system, 101…Remote conferencing device, 102…Internet, 103…Communication terminal, 200A, 200B, 200C, 200D…Table, 201…Communication unit, 202…Transmission control unit, 203A…Video input unit, 203B…Operation input unit, 203C…Audio input unit, 204…Reception control unit, 205A…Video output unit, 205B…Text output unit, 205C…Audio output unit, 301, 701…Operation detection unit, 302…Video control unit, 303…Audio control unit, 304, 501…Keyword utterance detection unit, 305, 502…Keyword model storage unit, 306, 503…Keyword output control unit, 307…Keyword information storage unit, 308…Integration unit 400A, 400B, 400C, 400D, 400E, 400F, 400G…Display screen, 401…Window, 402…Video control button, 403…Audio control button, 404…Display name, 405…Audio stop mark, 406…Video stop mark, 407…Participant video, 408…Conference chat display field, 409…Conference chat input field, 450…Box, 460…Gesture video, 500…Waveform data, 504…Input audio storage unit, 510, 520…Section, 601…Gesture detection unit, 602…Gesture model storage unit, 603…Gesture output control unit, 604…Gesture information storage unit, 702…Operation information storage unit, 800…Signal processing unit, 811…Acquisition unit, 812…Detection unit, 813…Transmission unit, 814…System control unit
Claims
1. On the computer, An acquisition function that acquires an audio input signal related to the user's voice and control information of the said audio input signal, A detection function that applies a keyword model to the aforementioned voice input signal to detect a voice signal corresponding to a predetermined keyword utterance from the voice input signal, If the control information prevents the audio input signal from being transmitted to the external device, a transmission function is provided to transmit the audio signal or a media file corresponding to the audio signal to the external device. A remote conferencing support program that makes this possible.
2. On the computer, An acquisition function that acquires an audio input signal related to the user's voice and control information of the said audio input signal, A detection function that detects an emotion signal corresponding to a predetermined emotion from the voice input signal by applying an emotion model to the voice input signal, If the control information prevents the audio input signal from being transmitted to the external device, a transmission function is provided to transmit a media file corresponding to the emotion signal to the external device. A remote conferencing support program that makes this possible.
3. The acquisition function acquires the video input signal related to the user's video, The detection function detects a gesture signal corresponding to a predetermined gesture from the video input signal by applying a gesture model to the video input signal. A remote conferencing support program according to claim 1 or claim 2.
4. The transmission function, if the video input signal is not transmitted to the external device due to control information of the video input signal, transmits the media file corresponding to the detected gesture signal to the external device. The remote conferencing support program according to claim 3.
5. The acquisition function acquires the video input signal related to the user's video, The detection function detects an expression signal corresponding to a predetermined expression from the video input signal by applying an expression model to the video input signal. A remote conferencing support program according to claim 1 or claim 2.
6. The transmission function, if the video input signal is not transmitted to the external device due to control information of the video input signal, transmits the media file corresponding to the detected facial expression signal to the external device. The remote conferencing support program according to claim 5.
7. The acquisition function acquires operation input signals related to the operation pattern entered by the user, The detection function detects a predetermined operation pattern from the operation input signal, The transmission function, if the audio input signal is not transmitted to the external device due to the control information, transmits the media file corresponding to the detected operation pattern to the external device. A remote conferencing support program according to claim 1 or claim 2.
8. The aforementioned media file may be text, images, music, audio, or video. A remote conferencing support program according to claim 1 or claim 2.
9. The detection function is The ID corresponding to the aforementioned audio signal is detected, From keyword information that associates pronunciation and transmitted text information for each ID, the transmitted text information corresponding to the ID is detected. The transmission function transmits the transmission text information to the external device. The remote conferencing support program according to claim 1.
10. The keyword information is selected, edited, or registered by the user. The remote conferencing support program according to claim 9.
11. An acquisition unit that acquires an audio input signal related to the user's voice and control information of the audio input signal, A detection unit that applies a keyword model to the aforementioned voice input signal to detect a voice signal corresponding to a predetermined keyword utterance from the voice input signal, If the control information prevents the audio input signal from being transmitted to the external device, a transmission unit transmits the audio signal or a media file corresponding to the audio signal to the external device. A remote conferencing support device equipped with the following features.
12. An acquisition unit that acquires an audio input signal related to the user's voice and control information of the audio input signal, A detection unit that applies an emotion model to the voice input signal to detect an emotion signal corresponding to a predetermined emotion from the voice input signal, If the control information prevents the audio input signal from being transmitted to the external device, the transmission unit transmits a media file corresponding to the emotion signal to the external device. A remote conferencing support device equipped with the following features.
13. To acquire an audio input signal related to the user's voice and control information for the said audio input signal, By applying a keyword model to the aforementioned audio input signal, an audio signal corresponding to a predetermined keyword utterance is detected from the audio input signal. If the control information prevents the audio input signal from being transmitted to the external device, the audio signal or a media file corresponding to the audio signal will be transmitted to the external device. A remote conferencing support method that includes the following features.
14. To acquire an audio input signal related to the user's voice and control information for the said audio input signal, By applying an emotion model to the aforementioned audio input signal, an emotion signal corresponding to a predetermined emotion is detected from the audio input signal. If the control information prevents the audio input signal from being transmitted to the external device, the media file corresponding to the emotion signal is transmitted to the external device. A remote conferencing support method that includes the following features.
Citation Information
Patent Citations
Intelligent detection and automatic correction of erroneous audio settings in a video conference
CN114079746A
Method, device, and program for voice communication record generation
JP2004279897A
Meeting management device, and meeting management method
JP2011066794A
Speaker detection system, speaker detection method, and program
JP2020155944A
Video communication device and video display method
JP2022083548A