Communication terminal, determination method, and program

The communication terminal's advanced mute management system, which considers audio intensity and predetermined keywords, addresses the issue of unintended mute releases in conference voice systems, ensuring that mute operations align with the user's intention.

JP7694217B2Active Publication Date: 2025-06-18OKI ELECTRIC INDUSTRY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021116986
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-15
Publication Date
2025-06-18
Estimated Expiration
2041-07-15

AI Technical Summary

Technical Problem

Existing conference voice systems release mute simply by capturing voice, leading to unintended mute releases due to background noise or unintended voice capture, not aligning with the user's intention.

Method used

A communication terminal with an audio data generation unit, a mute state determination unit, and a mute release determination unit that assesses the intensity of input audio and audio from other terminals to determine if the mute state should be released, specifically considering the presence of predetermined keywords in the input audio.

Benefits of technology

The solution ensures that the mute state is released only when the user intends to speak, reducing unnecessary mute releases and improving user experience by aligning the mute operation with the user's intention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694217000001
    Figure 0007694217000001
  • Figure 0007694217000002
    Figure 0007694217000002
  • Figure 0007694217000003
    Figure 0007694217000003
Patent Text Reader

Abstract

To provide a communication terminal configured to perform operation for unmuting in accordance with the intention of a user.SOLUTION: A communication terminal which can construct a voice transmitting / receiving system for transmitting / receiving voice data to / from another communication terminal includes: a voice data generation unit which generates first voice data from input voice input to the communication terminal; a mute-state determination unit which determines whether the communication terminal is in mute state in which the first voice data is not transmitted; and an unmuting determination unit which determines, when the terminal is in the mute state, whether to release the mute state on the basis of a first voice level indicating intensity of the input voice and a second voice level indicating intensity of voice indicated by second voice data transmitted from the other communication terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a communication terminal, a determination method, and a program.

Background Art

[0002] A conference system that conducts a conference using a communication terminal is known. For example, Patent Document 1 discloses a conference voice system including a plurality of microphones, a microcomputer including voice level detection means and voice data storage means, and a speaker.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The conference voice system described in Patent Document 1 is provided with an auto-mute release device that releases the mute when voice is captured by the microphone. However, in the conference voice system described in Patent Document 1, since the mute is released simply by capturing voice, voice that does not include the intention of the user to speak, for example, coughing or noise, may be captured by the microphone, and the mute may be released when the user does not intend to speak.

[0005] The present invention has been made in view of the above points, and an object thereof is to provide a communication terminal capable of performing an operation related to mute release in a manner consistent with the intention of the user.

Means for Solving the Problems

[0006] The communication terminal according to the present invention is a communication terminal capable of constructing an audio transmission / reception system for transmitting and receiving audio data with other communication terminals, and includes an audio data generation unit that generates first audio data from input audio input to the communication terminal, a mute state determination unit that determines whether or not the communication terminal is in a mute state in which the first audio data is not transmitted in the audio transmission / reception system, and a mute release determination unit that determines whether or not the mute state should be released based on a first audio level indicating the intensity of the input audio and a second audio level indicating the intensity of audio indicated by second audio data transmitted from the other communication terminal when the communication terminal is in the mute state. and, when the voice indicated by the first voice data includes a predetermined keyword, the mute release determination unit determines that the mute state should be released It is characterized by this.

[0007] Further, the determination method according to the present invention is a determination method by a communication terminal capable of constructing an audio transmission / reception system for transmitting and receiving audio data with other communication terminals, and includes an audio data generation step in which an audio data generation unit generates first audio data from input audio input to the communication terminal, a mute state determination step in which a mute state determination unit determines whether or not the communication terminal is in a mute state in which the first audio data is not transmitted in the audio transmission / reception system, and a mute release determination step in which when the mute state determination unit determines that the communication terminal is in the mute state, a mute release determination unit determines whether or not the mute state should be released based on a first audio level indicating the intensity of the input audio and a second audio level indicating the intensity of audio indicated by second audio data transmitted from the other communication terminal. and, when the voice indicated by the first voice data includes a predetermined keyword, the mute release determination unit determines that the mute state should be released It is characterized by this.

[0008] Further, the program according to the present invention is for a communication terminal capable of constructing an audio transmission / reception system for transmitting and receiving audio data with other communication terminals A program to be executed, a voice data generation step in which a voice data generation unit generates first voice data from input voice input to the communication terminal, a mute state determination step in which a mute state determination unit determines whether or not the communication terminal is in a mute state in which the first voice data is not transmitted in the voice transmission / reception system, and when the mute state determination unit determines that the communication terminal is in the mute state, a mute release determination step in which a mute release determination unit determines whether or not the mute state should be released based on a first voice level indicating the intensity of the input voice and a second voice level indicating the intensity of the voice indicated by second voice data transmitted from the other communication terminal. having, and when the voice indicated by the first voice data includes a predetermined keyword, the mute release determination unit determines that the mute state should be released It is a program.

Brief Description of Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be specifically described with reference to the drawings. In the drawings, the same components are denoted by the same reference numerals, and the description of overlapping components will be omitted.

Embodiment

[0011] FIG. 1 is a diagram showing a conference system 100 as a voice transmission / reception system according to Embodiment 1. In the following description, a case where the conference system 100 is a system in which three communication terminals 10, 11, and 12 and a conference server 14 are communicably connected via a network NW will be described. Of course, the number of communication terminals constituting the conference system 100 is not limited to the three shown in FIG. 1, and any number may be used as long as the system capabilities permit.

[0012] The network NW is, for example, a wired or wireless communication network capable of two-way data communication such as a WAN (Wide Area Network), a LAN (Local Area Network), or a public communication line (public line).

[0013] Each of the communication terminals 10, 11, and 12 is a communication terminal connectable to the conference server 14 via the network NW. By being communicably connected to each other by the conference server 14, the communication terminals 10, 11, and 12 can transmit and receive voice data to and from each other via the conference server 14. In the present embodiment, each of the communication terminals 10, 11, and 12 is a PC (Personal Computer) capable of transmitting and receiving voice data.

[0014] The conference server 14 is a communication device that individually establishes connections with each of the communication terminals 10, 11, and 12 via the network NW, and enables each of the communication terminals 10, 11, and 12 to transmit and receive voice data to and from each other.

[0015] In this embodiment, a conference application for constructing a conference system 100 is installed in each of the communication terminals 10, 11, and 12. The conference server 14 can make each of the communication terminals 10, 11, and 12 capable of transmitting and receiving voice data to and from each other by responding to connection requests from each of the communication terminals 10, 11, and 12 via the application.

[0016] Note that the conference application may be acquired by each of the communication terminals 10, 11, and 12 through communication via the network NW, for example, or may be acquired via a storage medium such as an optical disk like a DVD or a USB.

[0017] FIG. 2 is a block diagram showing the configuration of the communication terminal 10. Hereinafter, the communication terminals 11 and 12 also have the same configuration as the communication terminal 10.

[0018] The control unit 15 is a processing device including a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory). The CPU realizes various functions by reading and executing various programs stored in the ROM. The control unit 15 is a part that gives instructions and controls to each part in response to operations by a user of the communication terminal 10 (hereinafter also referred to as a user). In this embodiment, the control unit 15 executes the processing of the above-described conference application.

[0019] The input device 16 is an input device that receives input operations from the user of the communication terminal 10. The input device 16 is an input device for inputting information such as characters and numbers, such as a keyboard and a mouse, for example.

[0020] The microphone 17 is a voice input device that picks up the voice of the user of the communication terminal 10, for example, the voice uttered by the user, and converts it into an electrical signal. In other words, the microphone 17 is a voice data generation unit that generates voice data as first voice data from the input voice input to the communication terminal 10.

[0021] Speaker 18 is an audio output device that outputs the audio indicated by the audio data as the second audio data transmitted from communication terminals 11 and 12 based on the control of control unit 15. In this embodiment, the user of communication terminal 10 can have a voice call with each user of communication terminals 11 and 12 through microphone 17 and speaker 18.

[0022] Camera 19 is an imaging device that performs imaging based on the control of control unit 15. Camera 19 is, for example, a camera that captures the user of communication terminal 10.

[0023] Display 21 is a display device that performs screen display based on the control of control unit 15. When communicably connected to communication terminals 11 and 12, for example, a meeting user interface such as a window in which the video of camera 19, the ON / OFF status of the mute of the audio in communication terminal 10, and the user names of the users of communication terminals 11 and 12 participating in the meeting are displayed is displayed on display 21.

[0024] Note that display 21 may be a touch panel display in which a touch panel that receives an input operation from the user of communication terminal 10 as input device 16 and a display that performs screen display based on the control of control unit 15 are combined. When display 21 is a touch panel, display 21 functions as an input device in addition to or instead of input device 16.

[0025] Next, the functional blocks of control unit 15 will be described.

[0026] Communication unit 23 is a functional unit that transmits and receives data to and from communication terminals 11 and 12 according to the instructions of control unit 15. Communication unit 23 forms, for example, a communication interface for exchanging data via network NW together with a communication interface device such as a NIC (Network Interface Card), and is a part that transmits and receives data via network NW.

[0027] The communication unit 23 can be a transmission unit that transmits the voice data converted by the control unit 15 after being voice - input by the microphone 17 to the conference server 14. Also, the communication unit 23 can be a reception unit that receives voice data transmitted from other communication terminals via the conference server 14.

[0028] The mute state determination unit 24 is a determination unit that determines whether or not the communication terminal 10 is in a mute state where it does not transmit voice data in the conference system 100. For example, when voice muting is selected by the user's operation of the input device 16, the mute state determination unit 24 determines that the communication terminal 10 is in a mute state.

[0029] The mute release determination unit 25 is a determination unit that determines whether or not to release the mute state based on the voice input to the microphone 17 of the own terminal, that is, the communication terminal 10. Specifically, the mute release determination unit 25 determines whether or not to release the mute state by determining whether or not the user of the communication terminal 10 has made a voice intended for speaking.

[0030] For example, the mute release determination unit 25 determines whether or not the user of the communication terminal 10 has made a voice intended for speaking based on whether or not the voice level indicating the intensity of the voice input to the microphone 17 of the own terminal (hereinafter also referred to as the first voice level) has become equal to or higher than a predetermined threshold value (hereinafter also referred to as the first threshold value).

[0031] The first threshold value can be set, for example, from the history of the voice levels when the user of the own terminal has spoken. Also, the first threshold value is set to be larger than the voice levels of small noises such as the user's coughing sound or the mouse click sound, or environmental sounds.

[0032] The mute release determination unit 25 is also a determination unit that determines whether or not the users of the communication terminals 11 and 12 have made voices intended for speaking based on the voices indicated by the voice data transmitted from the other terminals, that is, the communication terminals 11 and 12.

[0033] For example, the mute release determination unit 25 determines whether the users of the communication terminals 11 and 12 are making an intended speech based on whether the voice level (hereinafter also referred to as the second voice level) indicating the strength of the voice indicated by the voice data transmitted from other terminals, i.e., the communication terminals 11 and 12, has become equal to or lower than a predetermined threshold (hereinafter also referred to as the second threshold).

[0034] The second threshold value is set based on a history of the voice levels when the users of the communication terminals 11 and 12 speak, so as to be smaller than the voice levels.

[0035] In this embodiment, when communication terminal 10 is in a muted state, if mute release determination unit 25 determines based on the first voice level that the user of communication terminal 10 has made an utterance intending to speak, and determines based on the second voice level that the users of communication terminals 11 and 12 have not made an utterance intending to speak, then it determines that the mute state should be released.

[0036] When the mute release determination unit 25 determines that the mute state should be released, the control unit 15 executes control to output a notification sound from the speaker 18 to notify the user of the communication terminal 10 that the communication terminal 10 is in a muted state.

[0037] The notification sound may be, for example, a simple alarm sound such as "beep beep" or a voice such as "Muted." Furthermore, the control unit 15 may cause the speaker 18 to output the notification sound and may also display "Muted" on the display 21.

[0038] A user of the communication terminal 10, for example, may become aware that the communication terminal 10 is in a muted state due to a notification sound output from the speaker 18 when speaking, and may then operate the input device 16 to unmute the communication terminal 10 and speak again.

[0039] In addition, the control unit 15 may output a notification sound from the speaker 18 and release the mute state of the communication terminal 10. As a result, it is possible to save the trouble of the user who has noticed that the communication terminal 10 is in the mute state by the notification sound from performing an operation to release the mute state of the communication terminal 10.

[0040] In other words, the user of the communication terminal 10 can speak as it is while recognizing that the communication terminal 10 was in the mute state and that the mute state has been released. Note that the control unit 15 does not necessarily have to output a notification sound and release the mute state of the communication terminal 10, and may simply release the mute state of the communication terminal 10 without a notification sound.

[0041] Further, when the mute state of the communication terminal 10 is manually or automatically released by the user of the communication terminal 10, the control unit 15 may notify the user of the communication terminal 10 that the mute state has been released by displaying a message such as "The mute state has been released" on the display 21.

[0042] FIG. 3 is a block diagram showing the configuration of the conference server 14. The control unit 27 includes a CPU, a ROM, and a RAM, and is a processing device that gives instructions and controls to each part of the conference server 14.

[0043] As described above, the control unit 27 makes the communication terminals 10, 11, and 12 capable of transmitting and receiving audio data to and from each other by responding to connection requests transmitted from each of the communication terminals 10, 11, and 12 via the conference application.

[0044] The mixing unit 28 of the control unit 27 has a mixer function that performs a synthesis process on the audio data transmitted from each of the communication terminals 10, 11, and 12 to generate one piece of audio data when each of the communication terminals 10, 11, and 12 is in a state where it can transmit and receive audio data to and from each other. The audio data generated by the mixing unit 28 is transmitted to each of the communication terminals 10, 11, and 12.

[0045] The communication unit 29 is a communication interface that transmits and receives data with external devices according to the instructions of the control unit 27. The communication unit 29 is, for example, a NIC for connecting to the network NW. The communication unit 29 can be a receiving unit that receives voice data transmitted from each of the communication terminals 10, 11, and 12. Further, the communication unit 29 can be a transmitting unit that transmits the voice data subjected to the synthesis process by the mixing unit 28 to each of the communication terminals 10, 11, and 12.

[0046] An example of the specific operation of the communication terminal 10 in this embodiment will be described below using a flowchart.

[0047] FIG. 4 is a flowchart showing a notification sound output routine RT1 executed in the control unit 15 of the communication terminal 10. The control unit 15 starts the notification sound output routine RT1, for example, using as a start trigger the establishment of a connection between the own terminal, that is, the communication terminal 10, and the communication terminals 11 and 12 via the conference server 14.

[0048] The control unit 15 first determines whether the communication terminal 10 is in the mute state via the mute state determination unit 24 (step S101). When the mute state determination unit 24 determines that the communication terminal 10 is not in the mute state (step S101: NO), the control unit 15 ends the notification sound output routine RT1.

[0049] When the mute state determination unit 24 determines that the communication terminal 10 is in the mute state (step S101: YES), the control unit 15 determines whether a first voice level indicating the intensity of the voice input to the microphone 17 via the mute release determination unit 25 is equal to or greater than a first threshold value (step S102).

[0050] When the mute release determination unit 25 determines that the first voice level is not equal to or greater than the first threshold value (step S102: NO), that is, when it is determined that the user of the communication terminal 10 is not making a voice intended for speaking, the control unit 15 ends the notification sound output routine RT1.

[0051] When the mute release determination unit 25 determines that the first voice level has become equal to or higher than the first threshold value (step S102: YES), that is, when it is determined that the user of the communication terminal 10 is making a voice with the intention of speaking, the control unit 15 determines whether the second voice level has become equal to or lower than the second threshold value via the mute release determination unit 25 (step S103).

[0052] When the control unit 15 determines that the mute release determination unit 25 has determined that the second voice level is not equal to or lower than the second threshold value (step S103: NO), that is, when it is determined that the users of the communication terminals 11 and 12 are making voices with the intention of speaking, the control unit 15 ends the notification sound output routine RT1.

[0053] When the control unit 15 determines that the mute release determination unit 25 has determined that the second voice level has become equal to or lower than the second threshold value (step S103: YES), that is, when it is determined that the users of the communication terminals 11 and 12 are not making voices with the intention of speaking, the control unit 15 causes a notification sound for notifying that the communication terminal 10 is in the mute state to be output from the speaker 18 (step S104).

[0054] As described above, in step S104, the control unit 15 causes a notification sound such as an alarm or voice for notifying the user of the communication terminal 10 that the communication terminal 10 is in the mute state to be output from the speaker 18. After step S104, the control unit 15 ends the notification sound output routine RT1.

[0055] As described above, according to the present embodiment, when the communication terminal 10 is in the mute state, if the mute release determination unit 25 determines that the user of the communication terminal 10 has made a voice with the intention of speaking based on the first voice level, and determines that the users of the communication terminals 11 and 12 have not made voices with the intention of speaking based on the second voice level, the control unit 15 causes a notification sound for notifying that the communication terminal 10 is in the mute state to be output from the speaker 18.

[0056] As a result, when the user of communication terminal 10 makes a voice intending to speak, the communication terminal 10 can be known to be in a mute state in a situation where the users of communication terminals 11 and 12 are not making voices intending to speak.

[0057] Also, when the control unit 15 controls to output a notification sound from the speaker 18 and release the mute state, the user of the communication terminal 10 can smoothly speak without performing an operation related to releasing the mute state or the like.

[0058] Therefore, according to the present embodiment, since the mute state is not released simply because the user's own voice is captured, or the mute state is not released when other conference participants are speaking during a conference, the operation related to releasing the mute can be performed in a manner consistent with the user's intention.

[0059] In the present embodiment, each of the communication terminals 10, 11, and 12 has been described as a PC, but it may be any terminal capable of transmitting and receiving voice data to and from each other via the conference server 14, and is not limited thereto. For example, each of the communication terminals 10, 11, and 12 may be a tablet terminal or a smartphone. Also, each of the communication terminals 10, 11, and 12 may be, for example, an IP (Internet Protocol) phone capable of switching the ON / OFF of the mute state, or a landline phone (analog phone).

[0060] Each of the communication terminals 10, 11, and 12 only needs to be able to transmit and receive voice data to and from each other via the conference server 14, and they may be different terminals from each other. For example, in the conference system 100, the communication terminal 10 may be a PC, the communication terminal 11 may be a smartphone, and the communication terminal 12 may be an IP phone.

[0061] In this embodiment, although the above-described conferencing application is installed in each of the communication terminals 10, 11, and 12, and the determination of the mute state and the determination of the user's speech are performed in each of the control units, the present invention is not limited thereto. For example, the determination of the mute state and the determination of the user's speech in each of the communication terminals 10, 11, and 12 described above may be performed by the conferencing server 14 on a Web application on a Web browser.

Embodiment

[0062] Hereinafter, a conferencing system 200 as a voice transmission / reception system according to Embodiment 2 will be described with reference to FIGS. 5 to 10. The conferencing system 200 is different from Embodiment 1 in that it has a voice recognition server 33, and the configurations of the communication terminals 30, 31, and 32 are different from those in Embodiment 1. The conferencing system has the same configuration as Embodiment 1 except for these points.

[0063] FIG. 5 is a diagram showing the configuration of the conferencing system 200. In the following description, a case will be described in which the conferencing system 200 is a system in which three communication terminals 30, 31, and 32, a conferencing server 14, and a voice recognition server 33 are communicably connected via a network NW. Of course, the number of communication terminals constituting the conferencing system 200 is not limited to the three shown in FIG. 5, and any number may be used as long as the system capabilities permit.

[0064] The voice recognition server 33 is a voice recognition server that converts voice data transmitted from the communication terminal 30 into text data and transmits the text data to the communication terminal 30. In this embodiment, the voice recognition server 33 is provided separately from the conferencing server 14.

[0065] FIG. 6 is a block diagram showing the configuration of the communication terminal 30. The control unit 34 is different from Embodiment 1 in the configuration of the mute release determination unit 35, and has the same configuration as Embodiment 1 in other respects. Hereinafter, the communication terminals 31 and 32 also have the same configuration as the communication terminal 30.

[0066] In this embodiment, the mute release determination unit 35 is composed of an audio level determination unit 35A and a keyword determination unit 35B.

[0067] When the communication terminal 30 is in the mute state, the audio level determination unit 35A determines whether the above-described first audio level is equal to or higher than the first threshold value, and also determines whether the above-described second audio level is equal to or lower than the second threshold value.

[0068] When the audio level determination unit 35A determines that the first audio level is equal to or higher than the first threshold value, the mute release determination unit 35 determines that the user of the communication terminal 30 has made a voice intended for speaking. Also, when the audio level determination unit 35A determines that the second audio level is equal to or lower than the second threshold value, the mute release determination unit 35 determines that the users of the communication terminals 31 and 32 have not made a voice intended for speaking.

[0069] The keyword determination unit 35B is a determination unit that compares the character string indicated by the text data transmitted from the speech recognition server 33 with the keywords stored in the keyword DB 36, and determines whether a predetermined keyword is included in the character string. Specifically, the keyword determination unit 35B determines whether a word having an intention to speak is included in the character string indicated by the above-described text data.

[0070] When the keyword determination unit 35B determines that a word having an intention to speak is included in the character string indicated by the above-described text data, the mute release determination unit 35 determines that the user of the communication terminal 30 has made a voice intended for speaking.

[0071] The keyword DB 36 is a database that holds a plurality of words having the above-described intention to speak. Note that the keyword DB 36 may be stored in an external storage device such as an external hard disk, and the control unit 34 may acquire the above-described keywords via the external storage device.

[0072] Here, the keywords held by the keyword DB36 described above will be explained with reference to FIG. 7.

[0073] FIG. 7 shows a keyword TB1 indicating an example of the keywords held by the keyword DB36. In the keyword TB1, the “type of keyword” indicates the situation in which the word stored in the keyword TB1 is used. Also, in the keyword TB1, the “example of keyword” shows an example of the word corresponding to each of the above-described types of keywords.

[0074] In the keyword TB1, the “words indicating greetings” are words such as “Good morning” and “Please take care of me,” which are mainly used at the start of a meeting.

[0075] Also, in the keyword TB1, the “words used when speaking to oneself” are words such as “Excuse me” and “Is that okay?” which are mainly used when one cuts into a conversation or starts a conversation oneself.

[0076] Also, in the keyword TB1, the “words used when being addressed by others” are words such as “That is” and “I understand,” which are mainly used when being asked for an explanation by others or when agreeing with the opinions of others.

[0077] Referring to FIG. 6 again, when the mute release determination unit 35 determines that the users of the communication terminals 31 and 32 are not making a voice intended for speaking based on the second voice level, the control unit 34 extracts a voice whose voice level is equal to or higher than the first threshold value for a certain period of time (for example, about the first 2 to 3 seconds), converts the voice into voice data, and transmits it to the voice recognition server 33.

[0078] In this embodiment, as described above, when the keyword determination unit 35B determines that the string indicated by the above-described text data includes a word having an intention to speak, the mute release determination unit 35 determines that the user of the communication terminal 30 is making a voice intended to speak.

[0079] When the mute release determination unit 35 determines that the user of the communication terminal 30 is making a voice intended to speak, the control unit 34 causes the speaker 18 to output a notification sound notifying that the communication terminal 30 is in the mute state, in the same manner as in the first embodiment. Note that the control unit 34 may cause the speaker 18 to output a notification sound and release the mute state of the communication terminal 30, in the same manner as in the first embodiment.

[0080] FIG. 8 is a block diagram showing the configuration of the voice recognition server 33. The control unit 37 is a processing device including a CPU, a ROM, and a RAM. The control unit 37 is a part that gives instructions and controls each part of the voice recognition server 33.

[0081] The voice recognition unit 38 in the control unit 37 is a part that performs voice recognition on the voice data transmitted from the communication terminal 30. Specifically, as described above, the voice recognition unit 38 converts the voice data transmitted from the communication terminal 30 into text data consisting of a character string by voice conversion.

[0082] For example, the voice recognition unit 38 extracts feature amounts such as the frequency and intensity of sound from the voice data transmitted from the communication terminal 30 (acoustic analysis), compares the feature amounts extracted by the acoustic analysis with information on sounds and words that have been learned in advance, extracts phonemes, which are the minimum units of voice (acoustic model), extracts combinations of sounds from the information database and recognizes them as words (pronunciation dictionary), and combines the phonemes extracted by the acoustic model and the words recognized by the pronunciation dictionary to recognize them as a meaningful sentence (language model), thereby enabling the recognition of voice as characters.

[0083] The communication unit 39 is a communication interface that transmits and receives data to and from the communication terminals 31 and 32 according to the instructions of the control unit 37. The communication unit 39 is, for example, a NIC for connecting to the network NW. The communication unit 39 can be a receiving unit that receives voice data transmitted from the communication terminal 30. Further, the communication unit 39 can be a transmitting unit that transmits text data generated by voice recognition to the communication terminal 30.

[0084] The mass storage device 41 is composed of, for example, a hard disk device, an SSD (solid state drive), a flash memory, etc., and stores various programs such as an operating system and software. In the present embodiment, the mass storage device 41 holds the acoustic model for the above-described voice recognition and information on sounds and words in the pronunciation dictionary.

[0085] Hereinafter, an example of each specific operation of the communication terminal 30 and the voice recognition server 33 in the present embodiment will be described using a flowchart.

[0086] FIG. 9 is a flowchart showing a notification sound output routine RT2 executed in the control unit 34 of the communication terminal 30. In FIG. 9, only the differences from the notification sound output routine RT1 executed in the control unit 15 of the communication terminal 10 according to the first embodiment will be described.

[0087] In step S103, when the mute release determination unit 25 determines that the second voice level has become equal to or lower than the second threshold (step S103: YES), the control unit 34 extracts approximately the first 2 to 3 seconds of the voice for which the first voice level has become equal to or higher than the first threshold, converts it into voice data, and transmits it to the voice recognition server 33 (step S201).

[0088] After step S201, the control unit 34 determines whether text data has been received from the voice recognition server 33 (step S202). When the control unit 34 determines that text data has not been received from the voice recognition server 33 (step S202: NO), it repeatedly executes step S202.

[0089] When the control unit 34 determines that it has received text data from the speech recognition server 33 (step S202: YES), it determines whether the text data contains keywords stored in the keyword DB 36 via the keyword determination unit 35B (step S203). That is, the keyword determination unit 35B determines whether the voice input to the microphone 17 of the own terminal is a word having an intention to speak.

[0090] When the control unit 34 determines that the keyword determination unit 35B has determined that the text data contains a word having an intention to speak (step S203: YES), that is, when the mute release determination unit 35 has determined that the user of the communication terminal 30 is making a voice intending to speak, it causes the speaker 18 to output a notification sound notifying that the communication terminal 10 is in a mute state (step S204).

[0091] When the control unit 34 determines that the keyword determination unit 35B has determined that the text data does not contain a keyword (step S203: NO), it ends the notification sound output routine RT2. The control unit 34 ends the notification sound output routine RT2 after step S204.

[0092] FIG. 10 is a flowchart showing a speech recognition routine RT3 executed in the control unit 37 of the speech recognition server 33. The control unit 37 starts the speech recognition routine RT3 using, for example, the establishment of a connection between the speech recognition server 33 and the communication terminal 30 via the network NW as a start trigger.

[0093] The control unit 37 determines whether it has received voice data from the communication terminal 30 (step S301). When the control unit 37 determines that it has received voice data from the communication terminal 30 (step S301: YES), it converts the voice indicated by the voice data into text data via the voice recognition unit 38 (step S302).

[0094] When the control unit 37 determines that voice data has not been received from the communication terminal 30 (step S301: NO), it ends the voice recognition routine RT3.

[0095] After step S302, the control unit 37 transmits the text data converted via the voice recognition unit 38 to the communication terminal 30 (step S303). After step S303, it ends the voice recognition routine RT3.

[0096] As described above, according to this embodiment, when the communication terminal 30 is in the mute state, if the mute release determination unit 35 determines that the first voice level is equal to or higher than the first threshold value and the second voice level is equal to or lower than the second threshold value, the control unit 34 transmits the voice data indicated by the voice equal to or higher than the first threshold value to the voice recognition server 33.

[0097] Then, the control unit 34 refers to the text data transmitted from the voice recognition server 33, and when the keyword determination unit 35B determines that the character string indicated by the text data includes a word having an intention to speak, it causes the speaker 18 to output a notification sound notifying that the communication terminal 30 is in the mute state.

[0098] Thereby, when the user of the communication terminal 30 emits a voice having a certain voice level, in a situation where the users of the communication terminals 11 and 12 are not speaking, the user can know that the communication terminal 30 is in the mute state when the voice input to the communication terminal 30 is a word having an intention to speak.

[0099] Also, when the control unit 34 outputs a notification sound from the speaker 18 and releases the mute state of the communication terminal 30, the user of the communication terminal 30 can speak without performing an operation to release the mute state of the communication terminal 30.

[0100] Therefore, according to this embodiment, similar to Embodiment 1, since it does not happen that the mute state is released just because one's own voice is captured, or the mute state is released when other conference participants are speaking during the conference, the operation regarding unmute can be performed in a manner that conforms to the user's intention.

[0101] In this embodiment, the voice recognition server 33, which plays a part in the function (unmute function) for releasing the mute state of the communication terminal 30, exists separately from the conference server 14. In other words, even when the conference server 14 changes, it is not necessary to change the voice recognition server 33 each time.

[0102] Therefore, for example, even when using a conference system constructed with different protocols for each conference, there is no need to perform different processes, such as generating voice data according to different protocols for each conference, in order to exhibit the above-described unmute function. Therefore, by providing the voice recognition server 33 separately from the conference server 14, it is possible to enhance the versatility of the above-described unmute function and the application equipped with the function.

[0103] For example, the above-described unmute function can be added as an add-on to various conference applications such as ZOOM (registered trademark), Skype (registered trademark), Teams (registered trademark), BlueJeans (registered trademark), Webex (registered trademark), etc., and the voice data of the conferences held in each conference application is transmitted to the voice recognition server 33, thereby realizing the above-described unmute function.

[0104] Note that in the notification sound output routine RT2, when the voice level determination unit 35A determines that the second voice level is equal to or lower than the second threshold value (step S103: YES), the control unit 34 transmits the voice having a voice level equal to or higher than the first threshold value to the voice recognition server 33 as voice data (step S201), but step S103 may not be executed.

[0105] That is, when the first voice level becomes equal to or higher than the first threshold value and the voice having the voice level equal to or higher than the first threshold value includes a word having an intention to speak, the control unit 34 may cause the speaker 18 to output a notification sound. Thereby, the control unit 34 can notify or cancel the mute state of the communication terminal 30 based only on the mode of the voice input to its own terminal.

[0106] In this embodiment, the voice recognition server 33 may have a keyword determination unit 35B instead of the communication terminal 30, and the mass storage device 41 may have a keyword DB 36. For example, the control unit 37 of the voice recognition server 33 may convert the voice data transmitted from the communication terminal 30 into text data by the voice recognition unit 38, and the keyword determination unit 35B may determine whether the character string indicated by the text data includes a keyword having an intention to speak, and may transmit the result of the determination to the communication terminal 30.

[0107] Thereby, when the determination result is that the above character string includes a word having an intention to speak based on the result of the keyword determination transmitted from the voice recognition server 33, the control unit 34 of the communication terminal 30 may output a notification sound of the mute state from the speaker 18.

[0108] In this embodiment, the voice recognition server 33 may be incorporated into each of the communication terminals 30, 31, and 32. For example, when the communication terminal 30 is an IP phone, the voice recognition server 33 may be incorporated into a private branch exchange (PBX) that connects a plurality of telephones. Also, the voice recognition server 33 may be incorporated into the conference server 14.

[0109] A series of processes in the control unit of each of the communication terminal, the conference server 14, and the voice recognition server 33 described in the first and second embodiments may be a program to be executed by a computer. Also, the program may be recorded on a computer-readable recording medium.

[0110] The type of the above-described recording medium is not particularly limited, and for example, it may be an optical disk, a hard disk, or a semiconductor memory such as a flash memory or an SSD. Further, the above program may be downloaded and installed in the communication terminal via communication.

[0111] The control routines shown in the above-described Example 1 and Example 2 are merely examples and can be appropriately selected and changed according to the application, usage conditions, etc.

Explanation of Signs

[0112] 10, 11, 12, 30, 31, 32 Communication terminal 14 Conference server 15, 27, 34, 37 Control unit 16 Input device 17 Microphone 18 Speaker 19 Camera 21 Display 23, 29, 39 Communication section 24 Mute state determination section 25, 35 Mute release determination section 26 Mixing section 33 Speech recognition server 35A Audio level determination section 35B Keyword determination section 36 Keyword DB 38 Speech conversion section 41 Mass storage device

Claims

1. A communication terminal capable of constructing a voice transmission / reception system for transmitting and receiving voice data with other communication terminals, a voice data generation unit that generates first voice data from input voice input to the communication terminal, a mute state determination unit that determines whether the communication terminal is in a mute state where the communication terminal does not transmit the first voice data in the voice transmission / reception system, and a mute release determination unit that determines whether to release the mute state based on a first voice level indicating the intensity of the input voice and a second voice level indicating the intensity of the voice indicated by second voice data transmitted from the other communication terminal when in the mute state. The communication terminal is characterized in that the mute release determination unit determines that the mute state should be released when a predetermined keyword is included in the voice indicated by the first voice data.

2. The communication terminal according to claim 1, wherein the mute release determination unit determines that the mute state should be released when the first voice level is equal to or higher than a first threshold value and the second voice level is equal to or lower than a second threshold value.

3. The communication terminal according to claim 1 or 2, wherein the predetermined keyword is a word having an intention to speak.

4. The communication terminal according to any one of claims 1 to 3, wherein a notification sound is output when the mute release determination unit determines that the mute state should be released.

5. The communication terminal according to any one of claims 1 to 4, wherein when the mute release determination unit determines that the mute state should be released, the mute state in the voice transmission / reception system is released.

6. A determination method by a communication terminal capable of constructing a voice transmission / reception system for transmitting and receiving voice data with other communication terminals, comprising: A voice data generation step in which a voice data generation unit generates first voice data from input voice input to the communication terminal; A mute state determination step in which a mute state determination unit determines whether or not the communication terminal is in a mute state in which the first voice data is not transmitted in the voice transmission / reception system; When the mute state determination unit determines that the communication terminal is in the mute state, a mute release determination step in which a mute release determination unit determines whether or not to release the mute state based on a first voice level indicating the intensity of the input voice and a second voice level indicating the intensity of the voice indicated by second voice data transmitted from the other communication terminal; The determination method is characterized in that the mute release determination unit determines that the mute state should be released when a predetermined keyword is included in the voice indicated by the first voice data.

7. A program for causing a communication terminal capable of constructing a voice transmission / reception system for transmitting and receiving voice data with other communication terminals to execute, comprising: A voice data generation step in which a voice data generation unit generates first voice data from input voice input to the communication terminal; A mute state determination step in which a mute state determination unit determines whether or not the communication terminal is in a mute state in which the first voice data is not transmitted in the voice transmission / reception system; When the mute state determination unit determines that the communication terminal is in the mute state, a mute release determination step in which a mute release determination unit determines whether or not to release the mute state based on a first voice level indicating the intensity of the input voice and a second voice level indicating the intensity of the voice indicated by second voice data transmitted from the other communication terminal; A program for the mute release determination unit to determine that the mute state should be released when a predetermined keyword is included in the voice indicated by the first voice data.

Citation Information

Patent Citations

  • Voice switch for speaking equipment

    JP1998308816A

  • Information processor, program, and information processing system

    JP2019184800A

  • Audio processing system, conferencing system, audio processing method, and audio processing program

    JP2020198588A

  • Information processing method, information processing device, and information processing program

    JP2022016997A

  • Conference audio system

    WO2007013180A1