Speech processing system, speech processing method, and program

Through the processor 12 in the voice processing system, the speech signal output is optimized by using clarity and speech identity determination, which solves the problem of impaired conversation comfort in a multi-user environment, and realizes clear voice communication and a comfortable conversation experience.

CN120303954APending Publication Date: 2025-07-11PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380083798.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-11-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In an environment where multiple users exist in the same stronghold, users are prone to hear direct voice from other users and voice via the conference system, resulting in impairment of conversation comfort.

Method used

Through the processor 12 in the voice processing system, the speech component contained in the output voice signal is controlled by using clarity calculation and speech identity determination, ensuring that the volume of the direct voice is reduced or the external sound intake function is turned on when specific conditions are met, the noise cancellation function is reduced, and the voice signal output is optimized.

Benefits of technology

In a multi-user environment, users can clearly hear direct voice, reduce noise interference, maintain conversational comfort, and improve clarity and comfort of voice communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303954A_ABST
    Figure CN120303954A_ABST
Patent Text Reader

Abstract

A speech processing system (1) is provided with a first input I / F (10), a second input I / F (11), and a processor (12). A first input I / F (10) acquires a first voice signal (Sig1) via a communication line. A second input I / F (11) acquires a second voice signal (Sig2) based on the voice collected by the microphone (2). The processor (12) outputs an output voice signal (Sig3) based on the first voice signal (Sig1) and the second voice signal (Sig2) to the speaker (3). The processor (12) includes, in the output voice signal (Sig3), a signal in which a component corresponding to the first voice signal (Sig1) has been reduced, when both a first condition that the first voice signal (Sig1) and a second voice signal (Sig2) include voice signals based on voice emitted by the same person and a second condition that the second voice signal (Sig2) is clear are satisfied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a voice processing system and the like for processing voices emitted from speakers. Background Art

[0002] For example, a voice communication terminal is disclosed in Patent Document 1. The voice communication terminal is a device that controls the voice output of at least one of a plurality of terminals participating in a multipoint voice communication system, and includes a voice configuration unit and an interlocutor management unit. The voice configuration unit sets a sound source configuration when outputting voices from other terminals. The interlocutor management unit detects a speaker and an interlocutor who is the other party of the speaker from among the plurality of terminals, and detects a conversation group based on the combination of the detected speaker and interlocutor. The voice configuration unit changes the setting of the sound source configuration according to a change in the detected conversation group.

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2012-108587 Summary of the Invention

[0006] Problems to be Solved by the Invention

[0007] The present disclosure provides a voice processing system and the like that are less likely to impair the comfort of conversation even in an environment where multiple users exist at the same location.

[0008] Means for Solving the Problems

[0009] A voice processing system according to one aspect of the present disclosure includes a first input interface, a second input interface, and a signal processing circuit. The first input interface acquires a first voice signal via a communication line. The second input interface acquires a second voice signal based on voices collected by a microphone. The signal processing circuit outputs an output voice signal based on the first voice signal and the second voice signal to a speaker. The signal processing circuit includes, in the output voice signal, a signal in which a component corresponding to the first voice signal is reduced when both a first condition and a second condition are satisfied, the first condition being that both the first voice signal and the second voice signal include voice signals based on voices emitted by the same person, and the second condition being that the second voice signal is clear.

[0010] In the speech processing method according to one aspect of the present disclosure, a first speech signal is acquired via a communication line. In the speech processing method, a second speech signal based on speech collected by a microphone is acquired. In the speech processing method, when both a first condition and a second condition are satisfied, a signal with a component corresponding to the first speech signal reduced is included in an output speech signal and output to a speaker. The first condition is that both the first speech signal and the second speech signal include speech signals based on speech uttered by the same person, and the second condition is that the second speech signal is clear.

[0011] A program according to one aspect of the present disclosure causes one or more processors to execute the speech processing method.

[0012] Advantageous Effects of the Invention

[0013] According to the speech processing system and the like of the present disclosure, there is an advantage that the comfort of a conversation is not easily impaired even in an environment where both offline conversations and online conversations coexist. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is an explanatory diagram of problems in communication using a conference system.

[0015] Figure 2 It is a block diagram showing an example of the overall configuration of a speech processing system according to an embodiment.

[0016] Figure 3 It is an explanatory diagram of a first determination operation for determining the clarity of a second speech signal.

[0017] Figure 4 It is an explanatory diagram of a second determination operation for determining the clarity of a second speech signal.

[0018] Figure 5 It is a flowchart showing an example of the operation of a speech processing system according to an embodiment.

[0019] Figure 6 It is a flowchart showing an example of the calculation of parameters required for determining the clarity of a second speech signal.

[0020] Figure 7 It is a flowchart showing an example of the calculation of parameters required for determining speech identity.

[0021] Figure 8 It is an explanatory diagram of an outline of the operation of a speech processing system according to an embodiment.

[0022] Figure 9 It is an explanatory diagram of the advantages of a speech processing system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0023] [1. Insights underlying the present disclosure]

[0024] First, the inventors' perspectives will be described below.

[0025] Conventionally, technologies for simultaneously conducting meetings and other forms of communication among multiple locations using, for example, a conferencing system via an MCU (Multipoint Control Unit) or a web conferencing service such as Zoom (registered trademark) have been known. In such communication, each participant wears a device equipped with a microphone and a speaker (e.g., a headset, etc.) to have conversations with other participants. Additionally, in recent years, each participant can also have conversations with other participants in the same virtual space or while observing the same virtual space by wearing a device that uses XR (x Reality) technology (e.g., a head-mounted display or smart glasses, etc.). In such communication, there are sometimes multiple participants at the same location, and the following problems exist.

[0026] Figure 1 It is an explanatory diagram of the problems in communication using a conferencing system. In Figure 1 , the conferencing system 100 is a server provided by the above-mentioned MCU-based conferencing system or web conferencing service. In Figure 1 The example shown represents two users U1 and U2 at the first location A1, user U3 at the second location A2, and user U4 at the third location A3 using the conferencing system 100 to conduct a meeting online. Users U1 to U4 can all hear the voices of other users by outputting voices based on the voice signals transmitted via the conferencing system 100 from the speakers. As an example, when user U1 emits a voice such as "Hello", other users U2, U3, and U4 can hear the voice "Hello" emitted by user U1 by outputting this voice based on the voice signal transmitted via the conferencing system 100 from the speakers.

[0027] Here, in Figure 1 The example shown, there are two users U1 and U2 at the first location A1. Therefore, at the first location A1, the voice emitted by one of the two users U1 and U2 can be directly heard by the other user without passing through the conferencing system 100. In this case, for example, when user U2 at the first location A1 emits a voice V2 such as "Hello", user U1 directly hears the voice V2 emitted by user U2 and also hears the voice V1 emitted by user U2 transmitted via the conferencing system 100.

[0028] As described above, when there are multiple users at the same base point, the users at that base point hear both the direct voice from other users at that base point and the voice via the conference system 100. Therefore, there is a problem that it is difficult to hear the voice emitted by other users. In addition, the voice of other users at that base point via the conference system 100 reaches the user's ear later than the direct voice from other users. Therefore, there is a problem that at the timing when a user hears the direct voice from other users and wants to speak, the voice of other users via the conference system 100 reaches the user's ear, so that the user's speech is obstructed and it is difficult to speak. Thus, there is a problem that in an environment where there are multiple users at the same base point, the comfort of the conversation is easily impaired.

[0029] In view of the above, the inventors have completed the present disclosure.

[0030] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below are all general or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions of the constituent elements, connection methods, steps, order of steps, etc. shown in the following embodiments are examples and are not intended to limit the present disclosure. In addition, the constituent elements in the following embodiments that are not described in the independent claims are described as optional constituent elements.

[0031] In addition, the drawings are schematic diagrams and are not necessarily strictly illustrated. In addition, in each drawing, the same reference numerals are given to substantially the same structures, and redundant explanations may be omitted or simplified.

[0032] (Embodiment)

[0033] [2. Structure]

[0034] [2-1. Overall Structure]

[0035] First, use Figure 2 to describe the overall structure of the voice processing system involved in the embodiment. Figure 2 is a block diagram showing the overall structure of the voice processing system involved in the embodiment. The voice processing system 1 is a system for outputting a voice based on a voice signal from a speaker 3 when a voice signal is obtained from the outside. In the embodiment, the voice processing system 1 is implemented by a voice call device 4.

[0036] The voice call device 4 can communicate with the conference system 100 via a network such as the Internet. In addition, the voice call device 4 can also communicate with the conference system 100 via a LAN (Local Area Network).

[0037] The voice call device 4 is a device worn on the user's head or neck, and is divided into a closed-type voice call device, an open-type voice call device, and a voice call device capable of switching between closed-type and open-type. The closed-type voice call device is a device that covers the user's ear canal (eardrum), and includes, for example, an earphone (Headset) of the in-ear headphone type or an earphone (Headphone) of the over-ear headphone type. The open-type voice call device is a device that does not cover the user's ear canal, and includes, for example, a neck-worn speaker or a wearable device of the goggle type for XR. The voice call device capable of switching between closed-type and open-type is a device that can switch between the function of covering the user's ear canal and the function of not covering the user's ear canal, and includes, for example, an in-ear headphone type earphone that can be switched by opening and closing a plate of a housing portion, or an over-ear headphone type earphone. In addition, in the voice call device 4, the main body portion that performs voice processing, etc., and the earphone portion including a microphone and a speaker may be integrally formed or separately formed.

[0038] The voice processing system 1 can be applied to any one of a closed-type voice call device, an open-type voice call device, and a voice call device capable of switching between closed-type and open-type. Hereinafter, as an example, the case where the voice call device 4 is a closed-type voice call device will be described.

[0039] As described above, the conference system 100 is, for example, a conference system via an MCU or a server provided by a Web conference service. If the conference system 100 receives a voice signal output from a voice call device 4 worn by an arbitrary user, it performs appropriate correction processing on the received voice signal, and then sends the corrected voice signal to one or more other users respectively wearing one or more voice call devices 4. The correction processing may include, for example, noise suppression processing for reducing noise included in the received voice signal. In addition, the correction processing may include, for example, frequency correction processing for emphasizing the audible frequency range of a person in the received voice signal. In addition, the conference system 100 may not perform correction processing on the received voice signal.

[0040] [2-2. Structure of Voice Call Device (Voice Processing System)]

[0041] Next, the structure of the voice call device 4 (voice processing system 1) will be specifically described. As Figure 2 shown, the voice call device 4 includes a microphone 2, a first input interface (hereinafter, referred to as "first input I / F (Interface)") 10, a second input interface (hereinafter, referred to as "second input I / F") 11, a processor 12, a memory 13, and a speaker 3.

[0042] The microphone 2 is a sound collection device that acquires the voices around the voice call device 4 and outputs a second voice signal Sig2 based on the acquired voices. Specifically, the microphone 2 is a capacitive microphone, a dynamic microphone, an MEMS (Micro ElectroMechanical Systems) microphone, etc., but there is no particular limitation. In addition, the microphone can be omnidirectional or directional.

[0043] The speaker 3 outputs a voice based on the output voice signal Sig3 output from the processor 12. The speaker 3 is a speaker that emits sound waves toward the ear canal of the user wearing the voice call device 4 and can be, for example, a bone conduction speaker.

[0044] The first input I / F 10 is, for example, a wireless communication interface. Based on wireless communication standards such as Wi-Fi (registered trademark), it communicates with the conference system 100 via a network and thereby receives the first voice signal Sig1 transmitted from the conference system 100. In other words, the first input I / F 10 acquires the first voice signal Sig1 via a communication line. The first voice signal Sig1 is a voice signal mainly based on the voices emitted by other users. The first input I / F 10 outputs the acquired first voice signal Sig1 to the processor 12.

[0045] The second input I / F 11 is an interface that receives the second voice signal Sig2 output from the microphone 2. In other words, the second input I / F 11 acquires the second voice signal Sig2 based on the voices collected by the microphone 2. The second input I / F 11 outputs the acquired second voice signal Sig2 to the processor 12.

[0046] The processor 12 is, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor), etc. The processor 12 performs information processing to output an output voice signal Sig3 based on the first voice signal Sig1 acquired by the first input I / F 10 and the second voice signal Sig2 acquired by the second input I / F 11 to the speaker 3. The above information processing is realized by the processor 12 executing a computer program stored in the memory 13. The processor 12 is an example of the signal processing circuit of the voice processing system 1.

[0047] In the processor 12, as functional components, there are included a clarity calculation unit 121, a clarity determination unit 122, a first feature quantity calculation unit 123, a second feature quantity calculation unit 124, a speech identity determination unit 125, an output speech determination unit 126, an output speech control unit 127, an external sound input switching unit 128, and an ANC (Active Noise Cancelling) control unit 129. Each of the above functions is realized, for example, by the processor 12 executing a computer program stored in the memory 13.

[0048] The clarity calculation unit 121 calculates the feature quantity of the second voice signal Sig2 used when the clarity determination unit 122 determines whether the second voice signal Sig2 is clear. Here, the second voice signal Sig2 being clear means that the SNR (Signal to Noise Ratio) of the frequency band corresponding to human speech (hereinafter referred to as the "speech frequency band") in the second voice signal Sig2 is higher than a threshold value and the features of human speech are clear. In other words, the second voice signal Sig2 being clear means that when a person hears the speech based on the second voice signal Sig2 output from the speaker 3, the degree to which the person can understand its content.

[0049] The clarity calculation unit 121 calculates the above SNR and the spectral envelope of the second voice signal Sig2 as the feature quantity of the second voice signal Sig2. Specifically, the clarity calculation unit 121 calculates the spectral contrast of the second voice signal Sig2 by performing appropriate signal processing on the second voice signal Sig2, and calculates the SNR of the speech frequency band in the second voice signal Sig2 based on the calculated spectral contrast. In addition, the clarity calculation unit 121 calculates the MFCC (Mel-Frequency Cepstral Coefficient) of the second voice signal Sig2. MFCC is the coefficient of the cepstrum used as a feature quantity in speech recognition, etc., and is obtained by transforming the power spectrum compressed using a Mel filter bank into a logarithmic power spectrum and applying an inverse discrete cosine transform to the logarithmic power spectrum. MFCC corresponds to the spectral envelope.

[0050] The clarity determination unit 122 uses the feature quantity of the second voice signal Sig2 calculated by the clarity calculation unit 121 to determine whether the second condition that the second voice signal Sig2 is clear is satisfied. The determination operation of the clarity determination unit 122 will be described in detail in [2-3. Determination of clarity] described later.

[0051] The first feature quantity calculation unit 123 calculates the feature quantity of the first speech signal Sig1 used when the speech identity determination unit 125 determines whether the first speech signal Sig1 and the second speech signal Sig2 are both speech signals based on the speech of the same person. The first feature quantity calculation unit 123 calculates the fundamental frequency of the first speech signal Sig1 and the spectral envelope of the first speech signal Sig1 as the feature quantity of the first speech signal Sig1. Specifically, the first feature quantity calculation unit 123 calculates the cepstrum of the first speech signal Sig1, and calculates the fundamental frequency of the first speech signal Sig1 based on the calculated cepstrum. The cepstrum is obtained by applying the Fourier transform to calculate the power spectrum of the first speech signal Sig1, transforming the calculated power spectrum into a logarithmic power spectrum, and further applying the Fourier transform to the logarithmic power spectrum. In addition, the first feature quantity calculation unit 123 calculates the spectral envelope by calculating the MFCC of the first speech signal Sig1. In addition, the first feature quantity calculation unit 123 calculates the timing of the appearance of vowels in the first speech signal Sig1 based on the calculated spectral envelope.

[0052] The second feature quantity calculation unit 124 calculates the feature quantity of the second speech signal Sig2 used when the speech identity determination unit 125 determines whether the first speech signal Sig1 and the second speech signal Sig2 are both speech signals based on the speech of the same person. The second feature quantity calculation unit 124 calculates the fundamental frequency of the second speech signal Sig2 and the spectral envelope of the second speech signal Sig2 as the feature quantity of the second speech signal Sig2. Specifically, the second feature quantity calculation unit 124 calculates the cepstrum of the second speech signal Sig2, and calculates the fundamental frequency of the second speech signal Sig2 based on the calculated cepstrum. In addition, the second feature quantity calculation unit 124 calculates the spectral envelope by calculating the MFCC of the second speech signal Sig2. In addition, the second feature quantity calculation unit 124 calculates the timing of the appearance of vowels in the second speech signal Sig2 based on the calculated spectral envelope.

[0053] In addition, the spectral envelope of the second speech signal Sig2 may be calculated by either the clarity calculation unit 121 or the second feature quantity calculation unit 124. In the embodiment, it is assumed that the second feature quantity calculation unit 124 calculates the spectral envelope of the second speech signal Sig2 for description. Therefore, the clarity calculation unit 121 may not calculate the spectral envelope of the second speech signal Sig2. In addition, when the spectral envelope of the second speech signal Sig2 is calculated by only one of the clarity calculation unit 121 and the second feature quantity calculation unit 124, the calculated spectral envelope is shared by the other party.

[0054] The speech identity determination unit 125 determines whether the first condition is satisfied, using the feature amounts of the first speech signal Sig1 calculated by the first feature amount calculation unit 123 and the feature amounts of the second speech signal Sig2 calculated by the second feature amount calculation unit 124. The first condition is that both the first speech signal Sig1 and the second speech signal Sig2 include speech signals based on speech uttered by the same person. In the embodiment, the speech identity determination unit 125 determines that the first condition is satisfied when (i) the fundamental frequency of the first speech signal Sig1 is the same as the fundamental frequency of the second speech signal Sig2, and (ii) the timing and types of vowels appearing in the first speech signal Sig1 are the same as the timing and types of vowels appearing in the second speech signal Sig2. On the other hand, the speech identity determination unit 125 determines that the first condition is not satisfied when at least one of the above (i) and (ii) is not satisfied.

[0055] As described above, the timing of vowel appearance in each speech signal can be detected based on the spectral envelope of each speech signal. In addition, since the first speech signal Sig1 passes through the communication line, the first input I / F 10 acquires it with a delay compared to the second speech signal Sig2. Therefore, the speech identity determination unit 125 takes the above delay into account when making the determination for (ii).

[0056] In this way, the processor 12 (speech identity determination unit 125) determines whether the first condition is satisfied based on the correlation between the components corresponding to vowels in the first speech signal Sig1 and the components corresponding to vowels in the second speech signal Sig2. Specifically, the speech identity determination unit 125 determines that the first condition is satisfied when (i) the difference between the fundamental frequency of the first speech signal Sig1 and the fundamental frequency of the second speech signal Sig2 is calculated and the calculated difference is below the threshold, and (ii) the difference between the time when a vowel appears in the first speech signal Sig1 and the time when a vowel appears in the second speech signal Sig2 is calculated and the calculated difference is below the threshold, and the types of vowels appearing in the first speech signal Sig1 are the same as the types of vowels appearing in the second speech signal Sig2. On the other hand, the speech identity determination unit 125 determines that the first condition is not satisfied when the difference calculated in (i) or (ii) exceeds the threshold. In addition, the speech identity determination unit 125 may also calculate the correlation coefficient between the spectral envelope calculated in the first speech signal Sig1 and the spectral envelope calculated in the second speech signal Sig2 in the determination of (ii), and determine whether the calculated correlation coefficient is below the threshold. Further, the speech identity determination unit 125 may also set it as satisfied for (ii) when a certain condition is satisfied in the determination of (ii).

[0057] In addition, the speech identity determination unit 125 may determine whether the first condition is satisfied only based on whether (i) is satisfied, or may determine whether the first condition is satisfied only based on whether (ii) is satisfied. Additionally, the speech identity determination unit 125 may also determine that the first condition is satisfied when at least one of (i) and (ii) is satisfied, and determine that the first condition is not satisfied when both (i) and (ii) are not satisfied.

[0058] Alternatively, the speech identity determination unit 125 may, instead of (ii), determine whether the pattern of consecutive occurrences of vowels in the first speech signal Sig1 is the same as the pattern of consecutive occurrences of vowels in the second speech signal Sig2. In this case, the speech identity determination unit 125 may also not consider the above-mentioned delay.

[0059] Here, the following method is also considered: based on the similarity between the waveform of the first speech signal Sig1 and the waveform of the second speech signal Sig2, if the waveform similarity is equal to or greater than a threshold, it is determined that the first condition is satisfied, and if the waveform similarity is less than the threshold, it is determined that the first condition is not satisfied. The "waveform" mentioned here is the waveform of the amplitude of the signal, that is, the waveform of the sound pressure level. However, as described above, the first speech signal Sig1 is a speech signal after correction processing is performed in the conference system 100, so the waveform of the first speech signal Sig1 will be different from the waveform of the second speech signal Sig2. Therefore, as described above, the speech identity determination unit 125 determines whether the first condition is satisfied by a method different from the method based on waveform similarity.

[0060] In addition, if the first speech signal Sig1 is not subjected to correction processing in the conference system 100, the speech identity determination unit 125 may also determine whether the first condition is satisfied based on waveform similarity. For example, the speech identity determination unit 125 may also determine whether the first condition is satisfied based on whether the change in the sound pressure level of the first speech signal Sig1 is substantially the same as the change in the sound pressure level of the second speech signal Sig2, or in other words, based on the correlation between the envelope of the amplitude of the first speech signal Sig1 and the envelope of the amplitude of the second speech signal Sig2.

[0061] The output voice determination unit 126 determines which of the first situation and the second situation it is in based on the determination result of whether the second condition is satisfied in the clarity determination unit 122 and the determination result of whether the first condition is satisfied in the speech identity determination unit 125. The first situation is a situation where the distance between the user and other users is relatively close, and the user can easily directly hear the voice emitted by other users. The second situation is a situation other than the first situation. The second situation includes, for example, a situation where the distance between the user and other users is relatively far, and the user has difficulty directly hearing the voice emitted by other users. The output voice determination unit 126 determines that it is in the first situation when both the first condition and the second condition are satisfied. On the other hand, the output voice determination unit 126 determines that it is in the second situation when at least one of the first condition and the second condition is not satisfied.

[0062] Based on the determination result of the output voice determination unit 126, the output voice control unit 127 controls the voice signal included in the output voice signal Sig3. Specifically, when the output voice determination unit 126 determines that it is in the first situation, the output voice control unit 127 controls in such a way as to reduce the volume of the voice based on the first voice signal Sig1 output from the speaker 3. The "reducing the volume of the voice based on the first voice signal Sig1" mentioned here means making the volume of the voice based on the first voice signal Sig1 lower than the volume of the voice based on the first voice signal Sig1 in the second situation (in other words, the default volume). In addition, the output voice control unit 127 turns on the external sound input function by controlling the external sound input switching unit 128, and turns off the noise cancellation function by controlling the ANC control unit 129.

[0063] That is, the processor 12 (output voice control unit 127) includes a signal with a reduced component corresponding to the first voice signal Sig1 in the output voice signal Sig3 when both the first condition and the second condition are satisfied. In addition, in the embodiment, the processor 12 reduces the component corresponding to the first voice signal Sig1 by reducing the volume of the voice based on the first voice signal Sig1, but it is not limited to this. For example, the processor 12 may also not include the first voice signal Sig1 in the output voice signal Sig3 when both the first condition and the second condition are satisfied. In addition, for example, it may be that the processor 12 performs suppression processing on the first voice signal Sig1 based on the second voice signal Sig2 when both the first condition and the second condition are satisfied, and includes the processed first voice signal Sig1 in the output voice signal Sig3.

[0064] In addition, when the processor 12 (output voice control unit 127) satisfies both the first condition and the second condition, it activates the external sound input function, that is, includes the second voice signal Sig2 in the output voice signal Sig3. In addition, the second voice signal Sig2 included in the output voice signal Sig3 may also be a signal processed through voice processing such as noise reduction processing or equalization processing.

[0065] On the other hand, when the output voice control unit 127 determines in the output voice determination unit 126 that it is in the second situation, it controls in such a way that the volume of the voice based on the first voice signal Sig1 output from the speaker 3 is set to the default volume. In addition, the output voice control unit 127 turns off the external sound input function by controlling the external sound input switching unit 128, and activates the noise cancellation function by controlling the ANC control unit 129.

[0066] That is, when the processor 12 (output voice control unit 127) does not satisfy at least one of the first condition and the second condition, it includes the first voice signal Sig1 in the output voice signal Sig3 and turns off the external sound input function, that is, does not include the second voice signal Sig2 in the output voice signal Sig3. In addition, the first voice signal Sig1 included in the output voice signal Sig3 may also be a signal processed through voice processing such as noise reduction processing or equalization processing.

[0067] In addition, when the processor 12 (output voice control unit 127) does not satisfy at least one of the first condition and the second condition, it activates the noise cancellation function, that is, further includes a voice signal with a phase opposite to that of the second voice signal Sig2 in the output voice signal Sig3.

[0068] The external sound input switching unit 128 switches the on / off of the external sound input function for taking in the voice around the user by being controlled by the output voice control unit 127. When the external sound input function is on, the speaker 3 outputs a voice based on the output voice signal Sig3 including the second voice signal Sig2. On the other hand, when the external sound input function is off, the speaker 3 outputs a voice based on the output voice signal Sig3 that does not include the second voice signal Sig2.

[0069] The ANC control unit 129 is controlled by the output voice control unit 127 to switch the on / off of the noise cancellation function. When the noise cancellation function is on, the ANC control unit 129 generates a voice signal with a phase opposite to that of the second voice signal Sig2, and includes the generated voice signal in the output voice signal Sig3. In this case, the speaker 3 outputs a voice based on the voice signal with a phase opposite to that of the second voice signal Sig2. Thus, near the user's ear, the voice that is the source of the second voice signal Sig2 and the voice based on the voice signal with a phase opposite to that of the second voice signal Sig2 cancel each other out, so the user can hardly hear these voices. On the other hand, when the noise cancellation function is off, the ANC control unit 129 does not generate a voice signal with a phase opposite to that of the second voice signal Sig2.

[0070] The memory 13 is a storage device that stores computer programs executed by the processor 12 and information required to implement various functions. The memory 13 is implemented by, for example, a semiconductor memory. The memory 13 can be implemented as an internal memory of the processor 12 instead of an external memory of the processor 12.

[0071] [2-3. Determination of clarity]

[0072] Hereinafter, the determination operation of whether the second voice signal Sig2 is clear by the clarity determination unit 122 will be described in detail. In the embodiment, the clarity determination unit 122 performs a first determination operation and a second determination operation. When it is determined to be clear in any of the determination operations, it is determined that the second voice signal Sig2 is clear, that is, the second condition is satisfied. On the other hand, when the clarity determination unit 122 is determined to be unclear in at least one of the first determination operation and the second determination operation, it is determined that the second voice signal Sig2 is unclear, that is, the second condition is not satisfied.

[0073] Figure 3 It is an explanatory diagram of the first determination operation for determining the clarity of the second voice signal Sig2. Figure 3 The spectral contrast of the second voice signal Sig2 is shown. In Figure 3 , the vertical axis represents the frequency band of the second voice signal Sig2, and the horizontal axis represents time (unit: "second"). In addition, in Figure 3 , light and dark indicate the level of SNR. The brighter it is, the higher the SNR, and the darker it is, the lower the SNR.

[0074] In the first determination operation, the clarity determination unit 122 is in the voice section ( Figure 3In the interval enclosed by the rectangular frame (for example, an interval of several tenths of a second), the SNR of the voice band in the second voice signal Sig2 is compared with a threshold value. Also, in the first determination operation, if the SNR is higher than the threshold value, the clarity determination unit 122 determines that the second voice signal Sig2 is clear, and if the SNR is lower than the threshold value, the clarity determination unit 122 determines that the second voice signal Sig2 is unclear.

[0075] Here, the SNR of the voice band in the voice signal can be calculated, for example, as a representative value of the SNRs of the respective bands included in the voice band in the voice signal. The representative value is, for example, an average value, a median value, a maximum value, or a mode value, etc. Additionally, the SNR of the voice band in the voice signal can also be calculated, for example, as the ratio of the representative value of the SNRs of the respective bands included in the voice band to the representative value of the SNRs of the respective bands outside the voice band. In the latter case, for example, even when the surroundings of the user are relatively noisy due to a large operating sound of an exhaust fan or the like and the SNR is relatively high in all bands, it is easier for the clarity determination unit 122 to determine whether the second voice signal Sig2 is clear.

[0076] Figure 3 (a) shows that in the voice interval enclosed by the rectangular frame, the SNR of the voice band ( Figure 3 the band indicated by the arrow in (a)) in the second voice signal Sig2 is lower than the threshold value. Therefore, in Figure 3 the example shown in (a), the clarity determination unit 122 determines that the second voice signal Sig2 is unclear in the first determination operation.

[0077] On the other hand, Figure 3 (b) shows that in the voice interval enclosed by the rectangular frame, the SNR of the voice band ( Figure 3 the band indicated by the arrow in (b)) in the second voice signal Sig2 is higher than the threshold value. Therefore, in Figure 3 the example shown in (b), the clarity determination unit 122 determines that the second voice signal Sig2 is clear in the first determination operation.

[0078] Figure 4 is an explanatory diagram of the second determination operation for determining the clarity of the second voice signal Sig2. Figure 4 represents the spectrum of the second voice signal Sig2 in the above voice interval. In Figure 4 , the vertical axis represents the amplitude value of the second voice signal Sig2, and the horizontal axis represents the frequency of the second voice signal Sig2. Additionally, in Figure 4 , the solid line L1 represents the spectral envelope, and the dash-dot line represents the tendency of the spectral envelope.

[0079] In the second determination operation, the clarity determination unit 122 calculates the kurtosis of the spectral envelope in each of the first frequency band B1, the second frequency band B2, and the third frequency band B3 in the voice section, and compares the calculated kurtosis with a threshold value. Further, in the second determination operation, if the kurtosis is higher than the threshold value in any of the frequency bands B1, B2, and B3, the clarity determination unit 122 determines that the second voice signal Sig2 is clear, and if the kurtosis is lower than the threshold value in at least one frequency band, the clarity determination unit 122 determines that the second voice signal Sig2 is unclear.

[0080] The first frequency band B1 is a frequency band corresponding to the first formant of a vowel in human speech. The second frequency band B2 is a frequency band corresponding to the second formant of a vowel in human speech. The third frequency band B3 is a frequency band corresponding to formants after the second formant of a vowel in human speech.

[0081] Here, each of the frequency bands B1 to B3 is a frequency band corresponding to the formant of a vowel in Japanese. Therefore, for example, when determining whether the second voice signal Sig2 is clear for a language other than Japanese such as English, the clarity determination unit 122 only needs to calculate the kurtosis of the spectral envelope in one or more frequency bands corresponding to the formant of a vowel in that language, and compare the calculated kurtosis with a threshold value.

[0082] Kurtosis is an index indicating the sharpness of the probability density function or frequency distribution of a probability variable. The higher the kurtosis, the more the distribution has a sharper peak and a longer and thicker tail compared to the normal distribution (in other words, the change around the peak of the spectral envelope is steeper), and the lower the kurtosis, the more the distribution has a rounder peak and a shorter and thinner tail compared to the normal distribution (in other words, the change of the spectral envelope is smoother).

[0083] In the second determination operation, when the kurtosis is higher than the threshold value in each of the frequency bands B1 to B3 as described above, the clarity determination unit 122 determines that the characteristics of the vowel in human speech are significantly manifested, that is, the human speech is clear enough to hear the vowel.

[0084] Figure 4 of (a) and Figure 4 of (b) both show the case where a person emits the vowel "o" in the voice section. Figure 4 In (a), as shown by the solid line L1 and the dash-dot line, the spectral envelope is smooth in each of the frequency bands B1 to B3, that is, the kurtosis is lower than the threshold value in each of the frequency bands B1 to B3. Therefore, in Figure 4 the example shown in (a) of, the clarity determination unit 122 determines that the second voice signal Sig2 is unclear in the second determination operation.

[0085] On the other hand, Figure 4(b), as indicated by the solid line and the single-dot chain line, shows that the peak values of the spectral envelope appear in each of the frequency bands B1 to B3, and the change around the peak values is steep, that is, the kurtosis is higher than the threshold value in each of the frequency bands B1 to B3. Therefore, in Figure 4 In the example shown in (b) of, the clarity determination unit 122 determines that the second voice signal Sig2 is clear in the second determination operation.

[0086] In this way, the processor 12 (clarity determination unit 122) determines whether the second condition is satisfied based at least on the component corresponding to the vowel in the second voice signal Sig2.

[0087] In addition, in the embodiment, the clarity determination unit 122 performs both the first determination operation and the second determination operation, but is not limited thereto. For example, it may also be determined whether the second condition is satisfied only by the second determination operation. However, considering the influence of background noise such as reverberation in space, the clarity determination unit 122 can determine the clarity of the second voice signal Sig2 with higher accuracy when performing both the first determination operation and the second determination operation.

[0088] [3. Operation]

[0089] Hereinafter, using Figure 5 An example of the operation of the voice call device 4 (voice processing system 1) according to the embodiment, that is, an example of the voice processing method, will be described. Figure 5 is a flowchart showing an example of the operation of the voice processing system 1 according to the embodiment.

[0090] First, if the second input I / F 11 acquires the second voice signal Sig2 (S101: Yes), the processor 12 holds the acquired second voice signal Sig2 in the buffer. Hereinafter, unless otherwise specified, the "second voice signal Sig2" corresponds to the second voice signal Sig2 held in the buffer. Then, if the first input I / F 10 acquires the first voice signal Sig1 (S102: Yes), the processor 12 calculates and updates the delay time (S103). Specifically, the processor 12 calculates the delay time by calculating the difference between the time point when the first input I / F 10 acquires the first voice signal Sig1 and the time point when the second input I / F 11 acquires the second voice signal Sig2, and updates the previous delay time to the calculated delay time. In addition, if the calculated delay time is the same as the previous delay time, the processor 12 does not perform the update.

[0091] Next, the processor 12 corrects the time deviation between the first voice signal Sig1 and the second voice signal Sig2 based on the delay time so that the start time point of the first voice signal Sig1 coincides with the start time point of the second voice signal Sig2 (S104).

[0092] Next, the processor 12 calculates the parameters required for determining the clarity of the second voice signal Sig2 based on the second voice signal Sig2 (S105). Hereinafter, Figure 6 Step S105 will be described in detail.

[0093] Figure 6 FIG. is a flowchart showing a calculation example of the parameters required for determining the clarity of the second voice signal Sig2. First, the processor 12 detects the voice section of the second voice signal Sig2 (S201). For example, the processor 12 detects the voice section starting from the time point after a predetermined time has elapsed from the start time point of the second voice signal Sig2. The voice section is, for example, an interval of a few tenths of a second.

[0094] Next, the processor 12 calculates the spectral contrast in the detected voice section (S202). Then, the processor 12 calculates the SNR of the voice band in the second voice signal Sig2 based on the calculated spectral contrast (S203).

[0095] In addition, the processor 12 calculates the feature amount of the second voice signal Sig2 in the detected voice section in parallel with steps S202 and S203 or either before or after steps S202 and 203 (S204). Here, the processor 12 calculates the fundamental frequency of the second voice signal Sig2 and the spectral envelope of the second voice signal Sig2 as the feature amount of the second voice signal Sig2. Next, the processor 12 stores the calculated feature amount of the second voice signal Sig2 in the memory 13 (S205).

[0096] Next, the processor 12 calculates the kurtosis of the spectral envelope of the second voice signal Sig2 in the detected voice section (S206). Specifically, the processor 12 calculates the kurtosis of the spectral envelope in each of the first frequency band B1, the second frequency band B2, and the third frequency band B3 in the detected voice section.

[0097] Return Figure 5, the processor 12 determines the clarity of the second voice signal Sig2 (S106). Specifically, the processor 12 performs a first determination operation of comparing the SNR of the voice band of the second voice signal Sig2 with a threshold value in the detected voice section. In addition, the processor 12 performs a second determination operation of comparing the kurtosis of the spectral envelope of each of the frequency bands B1 to B3 with a threshold value in the detected voice section. And, when the processor 12 determines that it is clear in either the first determination operation or the second determination operation, it determines that the second voice signal Sig2 is clear, that is, the second condition is satisfied. On the other hand, when the processor 12 determines that it is not clear in at least one of the first determination operation and the second determination operation, it determines that the second voice signal Sig2 is not clear, that is, the second condition is not satisfied.

[0098] When it is determined that the second voice signal Sig2 is clear, that is, the second condition is satisfied (S106: Yes), the processor 12 then calculates the parameters required for determining the speech identity based on the first voice signal Sig1 and the second voice signal Sig2 (S107). Hereinafter, Figure 7 Step S107 will be described in detail.

[0099] Figure 7 FIG. is a flowchart showing a calculation example of the parameters required for determining the speech identity. First, the processor 12 detects the voice section of the first voice signal Sig1 (S301). For example, the processor 12 uses the time point after a predetermined time has elapsed from the start time point of the first voice signal Sig1 as the starting point to detect the voice section. The detected voice section is the same section as the voice section of the second voice signal Sig2.

[0100] Next, the processor 12 reads the feature amount of the second voice signal Sig2 stored in the memory 13 (S302). In addition, the processor 12 calculates the feature amount of the first voice signal Sig1 in the detected voice section either in parallel with step S302 or before or after step S302 (S303). Here, the processor 12 calculates the fundamental frequency of the first voice signal Sig1 and the spectral envelope of the first voice signal Sig1 as the feature amount of the first voice signal Sig1.

[0101] Return Figure 5, the processor 12 determines the speech identity (S108). Specifically, the processor 12 determines that the speakers are the same, that is, the first condition is satisfied, when (i) the fundamental frequency of the first voice signal Sig1 is the same as the fundamental frequency of the second voice signal Sig2, and (ii) the timing of the vowel appearance in the first voice signal Sig1 is the same as the timing of the vowel appearance in the second voice signal Sig2. On the other hand, when at least one of the above (i) and (ii) is not satisfied, the processor 12 determines that the speakers are different, that is, the first condition is not satisfied. Additionally, here, if the difference between the two comparison objects is below the threshold, the processor 12 determines that the two comparison objects are the same.

[0102] When it is determined that the speakers are the same, that is, the first condition is satisfied (S108: Yes), both the first condition and the second condition are satisfied. Therefore, the processor 12 determines that it is in the first state and reduces the volume of the speech (i.e., the communication speech) based on the first voice signal Sig1 output from the speaker 3 (S109). Additionally, the processor 12 turns off the noise cancellation function (S110) and turns on the external sound input function (S111). Furthermore, the order of executing steps S109 to S111 is not limited to this order.

[0103] On the other hand, when at least one of the first condition and the second condition is not satisfied, that is, when it is determined that the second voice signal Sig2 is unclear (S106: No), or when it is determined that the speakers are different (S108: No), the processor 12 determines that it is in the second state. And the processor 12 sets the volume of the communication speech to the default volume (S112). Additionally, the processor 12 turns on the noise cancellation function (S113) and turns off the external sound input function (S114). Furthermore, the order of executing steps S112 to S114 is not limited to this order.

[0104] Additionally, steps S112 to S114 are also executed when the second input I / F 11 does not acquire the second voice signal Sig2 (S101: No), or when the first input I / F 10 does not acquire the first voice signal Sig1 (S102: No).

[0105] And the processor 12 repeats the above series of processes during the period before the call ends (S115: No). On the other hand, if the call ends (S115: Yes), the processor 12 ends the operation.

[0106] Figure 8 It is a schematic explanatory diagram of the operation example of the voice processing system 1 according to the embodiment. Figure 8 It shows a series of operations of the voice call device 4 (voice processing system 1) worn by the user U1 when there are two users U1 and U2 at the same base point.

[0107] As shown in Figure 8 (a) of FIG. [0000264], if another user U2 emits a voice V2, the voice is converted into a second voice signal Sig2 by the microphone 2, and thus the second input I / F 11 acquires the second voice signal Sig2. Further, the processor 12 detects a voice section, and in the detected voice section, calculates the fundamental frequency and MFCC (spectral envelope) of the second voice signal Sig2 as the feature amounts of the second voice signal Sig2. In addition, the processor 12 stores the calculated fundamental frequency and MFCC of the second voice signal Sig2 in the memory 13. Further, the voice V2 emitted by the other user U2 is transmitted to the conference system 100 as a first voice signal Sig1.

[0108] Next, as shown in Figure 8 (b) of FIG. [0000265], the processor 12 calculates the SNR of the voice band of the second voice signal Sig2 and the kurtosis of the spectral envelope in the detected voice section. Further, the processor 12 uses the calculated SNR of the voice band of the second voice signal Sig2 and the kurtosis of the spectral envelope to determine whether the second voice signal Sig2 is clear, that is, whether the second condition is satisfied.

[0109] In addition, as shown in Figure 8 (c) of FIG. [0000266], if the first input I / F 10 acquires the first voice signal Sig1 transmitted from the conference system 100, the processor 12 detects the voice section of the first voice signal Sig1, and in the detected voice section, calculates the fundamental frequency and MFCC (spectral envelope) of the first voice signal Sig1 as the feature amounts of the first voice signal Sig1. Further, the processor 12 reads the fundamental frequency and MFCC of the second voice signal Sig2 from the memory 13, and performs comparison with the fundamental frequency and MFCC of the first voice signal Sig1, thereby determining whether the speakers are the same, that is, whether the first condition is satisfied.

[0110] When both the first condition and the second condition are satisfied, that is, when it is determined that the first state is established, as shown in Figure 8 (d) of FIG. [0000267], the processor 12 reduces the volume of the communication voice (voice based on the first voice signal Sig1) output from the speaker 3, or does not reproduce the communication voice from the speaker 3. In addition, the processor 12 turns off the noise canceling function and turns on the external sound input function. Thereby, the user U1 can hardly hear the voice via the conference system 100 and mainly hears the direct and clear voice from the other user U2 with respect to the voice V2 emitted by the other user U2.

[0111] On the other hand, when at least one of the first condition and the second condition is not satisfied, that is, when it is determined that the second state is established, as shown in Figure 8As shown in (e), the processor 12 causes the communication voice (voice based on the first voice signal Sig1) to be output from the speaker 3. In addition, the processor 12 turns on the noise cancellation function and turns off the external sound input function. Thereby, the user U1 can hardly hear the direct and unclear voice from the other user U2 regarding the voice V2 emitted by the other user U2, and mainly hears the voice via the conference system 100.

[0112] [4. Effects, etc.]

[0113] Hereinafter, use Figure 9 to illustrate the advantages of the voice processing system 1 according to the embodiment. Figure 9 It is an explanatory diagram of the advantages of the voice processing system 1 according to the embodiment. Figure 9 It shows that two users U1 and U2 at the first base point A1, a user U3 at the second base point A2, and a user U4 at the third base point A3 use the conference system 100 to conduct a conference online.

[0114] Figure 9 (a) shows a situation where two users U1 and U2 at the first base point A1 are in relatively close positions to each other, and the user U1 can easily directly hear the voice V2 such as "Hello" emitted by the other user U2. In such a situation, the voice call device 4 (voice processing system 1) worn by the user U1 determines that both the first condition and the second condition are satisfied, that is, it is in the first situation, and includes a signal with a reduced component corresponding to the first voice signal Sig1 in the output voice signal Sig3. Here, the voice processing system 1 does not include the first voice signal Sig1 in the output voice signal Sig3, that is, does not reproduce the voice based on the first voice signal Sig1 from the speaker 3.

[0115] Therefore, the user U1 can directly hear the clear voice V2 such as "Hello" emitted by the other user U2, without hearing the voice V1 emitted by the other user U2 sent via the conference system 100. That is, regarding the voice emitted by the other user U2, the user U1 does not hear both the direct voice from the other user U2 and the voice via the conference system 100 (communication line), so it is easy to hear the voice emitted by the other user U2. Therefore, in the voice processing system 1, there is an advantage that the comfort of the conversation is not easily impaired even in an environment where there are multiple users at the same base point.

[0116] In addition, in the embodiment, when the voice processing system 1 determines that it is in the first situation, it turns on the external sound input function, that is, includes the second voice signal Sig2 in the output voice signal Sig3. Therefore, there is an advantage that it is easier to hear the direct voice from the other user U2 by taking in the voice around the user U1.

[0117] Figure 9 In the case of (b), two users U1 and U2 at the first base point A1 are located at positions relatively far from each other, and it is difficult for user U1 to directly hear the voice V2 such as "Hello" emitted by the other user U2. In such a situation, the voice call device 4 (voice processing system 1) worn by user U1 determines that at least the second condition is not satisfied, that is, it is in the second situation, includes the first voice signal Sig1 in the output voice signal Sig3, and does not include the second voice signal Sig2 in the output voice signal Sig3.

[0118] Therefore, user U1 can hear the voice V1 such as "Hello" emitted by the other user U2 transmitted via the conference system 100, and hardly hears the direct voice and unclear voice V2 from the other user U2. That is, regarding the voice emitted by the other user U2, user U1 does not hear both the direct voice from the other user U2 and the voice via the conference system 100 (communication line), so it is easy to hear the voice emitted by the other user U2. Therefore, in the voice processing system 1, there is an advantage that the comfort of the conversation is not easily impaired even in an environment where multiple users exist at the same base point.

[0119] In addition, in the embodiment, when the voice processing system 1 determines that it is in the second situation, it turns on the noise cancellation function, that is, includes a voice signal with a phase opposite to that of the second voice signal Sig2 in the output voice signal Sig3. Therefore, by removing the noise around user U1 including the direct voice from the other user U2, there is an advantage that user U1 can more easily hear the voice of the other user U2 transmitted via the conference system 100 (communication line).

[0120] In addition, when multiple users (here users U1 to U4) emit voices simultaneously, the voice processing system 1 determines that both the first condition and the second condition are not satisfied, that is, it is in the second situation. In this case, the voices of the other users U2 to U4 from the conference system 100 are output from the speaker 3. In such a situation, the conversations of each of the users U1 to U4 are temporarily stopped, so the advantages of the voice processing system 1 are not hindered. User U1 adopting the voice processing system 1 only needs to be able to enjoy the above advantages at least in the situation where the users U1 to U4 emit voices alternately.

[0121] [5. Other Embodiments]

[0122] The above describes the embodiments, but the present disclosure is not limited to the above embodiments.

[0123] In the above-described embodiment, the processor 12 may also not include the external sound input switching unit 128 and the ANC control unit 129. In this case, the voice processing system 1 may also not execute Figure 5 steps S110, S111, S113, and S114 in the flowchart shown. More specifically, in the case where the voice call device 4 is an open-type voice call device, the processor 12 may include the external sound acquisition switching unit 128 and the ANC control unit 129, but may also not include them. In addition, in the case where the voice call device 4 is an open-type voice call device, the processor 12 may also include the ANC control unit 129. In this case, the voice processing system 1 may also not execute Figure 5 steps S111 and S114 in the flowchart shown. Additionally, in the case where the voice call device 4 is a closed-type voice call device, the processor 12 preferably includes the external sound input switching unit 128, but may also not include the external sound input switching unit 128, for example, if a certain amount of external sound can be heard. Also, in the case where the voice call device 4 is a closed-type voice call device, the processor 12 preferably includes the ANC control unit 129, but may also not include the ANC control unit 129 if the external sound is reduced to a certain extent due to the user's ear canal being blocked.

[0124] Furthermore, in the above-described embodiment, for example, in a VR conference implemented while moving indoors, or in a situation where the characteristic amount of clarity such as the indoor ambient noise changes, the output voice control unit 127, the external sound input switching unit 128, and the ANC control unit 129 are each always controlled, but they may also not be controlled for a certain period of time. More specifically, the voice processing system 1 may also not execute Figure 5 steps S112 to S114 and S109 to S111 in the flowchart shown for a certain period of time (e.g., several milliseconds). In this case, the output voice control unit 127, the external sound input switching unit 128, and the ANC control unit 129 may each be controlled at regular intervals to prevent high-frequency control. In addition, the time for controlling the output voice control unit 127, the time for controlling the external sound acquisition switching unit 128, and the time for controlling the ANC control unit 129 may also be different from each other.

[0125] In addition, in the above-described embodiment, the voice processing system 1 is implemented by a single device (the voice call device 4), but it may also be implemented by multiple devices. In the case where the voice processing system 1 is implemented by multiple devices, the functional components included in the voice processing system 1 can be arbitrarily allocated to the multiple devices. Further, for example, the voice processing system 1 may also be implemented by a server having a first input I / F 10, a second input I / F 11, and a processor 12. In this case, the voice processing system 1 can obtain the second voice signal Sig2 from the microphone 2 or output the voice based on the output voice signal Sig3 from the speaker 3 by communicating with a device having the microphone 2 and the speaker 3.

[0126] In addition, the communication method between the devices in the above-described embodiment is not particularly limited. In the above-described embodiment, when two devices communicate with each other, a relay device (not shown) may exist between the two devices.

[0127] In addition, the order of the processes described in the above-described embodiment is an example. The order of multiple processes may be changed, or multiple processes may be executed in parallel. In addition, the process executed by a specific processing unit may also be executed by another processing unit. Further, a part of the digital signal processing described in the above-described embodiment may also be implemented by analog signal processing.

[0128] In addition, in the above-described embodiment, each component may also be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0129] In addition, each component may also be implemented by hardware. For example, each component may be a circuit (or an integrated circuit). These circuits may be integrated into one circuit as a whole, or may be separate circuits. In addition, these circuits may be general-purpose circuits or dedicated circuits, respectively.

[0130] In addition, the whole or specific technical solutions of the present disclosure may also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM. In addition, it may also be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium. For example, the present disclosure may be executed as a voice processing method executed by a computer, or may be implemented as a program for causing a computer to execute such a voice processing method. In addition, the present disclosure may also be implemented as a computer-readable non-transitory recording medium recording such a program. Further, the program here includes an application program for causing a general information terminal to function as the voice processing system in the above-described embodiment.

[0131] In addition, modes obtained by applying various modifications that occur to those skilled in the art to each embodiment, or modes achieved by arbitrarily combining the components and functions in each embodiment without departing from the gist of the present disclosure, are also included in the present disclosure.

[0132] (Summary)

[0133] As described above, the voice processing system 1 according to the first mode includes a first input I / F 10, a second input I / F 11, and a processor 12. The processor 12 is an example of a signal processing circuit. The first input I / F 10 acquires a first voice signal Sig1 via a communication line. The second input I / F 11 acquires a second voice signal Sig2 based on the voice collected by the microphone 2. The processor 12 outputs an output voice signal Sig3 based on the first voice signal Sig1 and the second voice signal Sig2 to the speaker 3. The processor 12 includes, in the output voice signal Sig3, a signal with a component corresponding to the first voice signal Sig1 reduced when both the first condition and the second condition are satisfied. The first condition is that both the first voice signal Sig1 and the second voice signal Sig2 include voice signals based on the voice emitted by the same person, and the second condition is that the second voice signal Sig2 is clear.

[0134] Thereby, in an environment where there are multiple users at the same location, when the direct voice from another user is clear among the direct voice from other users and the voice via the communication line, the user mainly hears the direct voice, and thus it is easy to hear the voice emitted by other users. That is, it has the advantage of not easily impairing the comfort of the conversation even in an environment where there are multiple users at the same location.

[0135] In addition, in the voice processing system 1 according to the second mode, in the first mode, when at least one of the first condition and the second condition is not satisfied, the processor 12 includes the first voice signal Sig1 in the output voice signal Sig3 and does not include the second voice signal Sig2 in the output voice signal Sig3.

[0136] Thereby, in an environment where there are multiple users at the same location, the user mainly hears the voice via the communication line, which is easier to hear than the direct voice, among the direct voice from other users and the voice via the communication line, and thus it is easy to hear the voice emitted by other users. That is, it has the advantage of not easily impairing the comfort of the conversation even in an environment where there are multiple users at the same location.

[0137] In addition, in the voice processing system 1 according to the third mode, in the first or second mode, the processor 12 determines whether the first condition is satisfied based on the correlation between the components corresponding to vowels in the first voice signal Sig1 and the components corresponding to vowels in the second voice signal Sig2.

[0138] Accordingly, compared with the method based on the similarity between the waveform of the first voice signal Sig1 and the waveform of the second voice signal Sig2, it has the advantage of easily determining whether the voice signals include voices uttered by the same person.

[0139] In addition, in the voice processing system 1 according to the fourth mode, in any one of the first to third modes, the processor 12 determines whether the second condition is satisfied based on the components corresponding to vowels in the second voice signal Sig2.

[0140] Accordingly, it has the following advantages: the determination is made based on the components corresponding to vowels, which are an index of the ease of hearing of human voices, so it is easy to determine whether the second voice signal Sig2 is clear.

[0141] In addition, in the voice processing system 1 according to the fifth mode, in any one of the first to fourth modes, when both the first condition and the second condition are satisfied, the processor 12 includes the second voice signal Sig2 in the output voice signal Sig3.

[0142] Accordingly, it has the following advantages: by taking in the voices around the user, it is easier for the user to hear the direct voices from other users.

[0143] In addition, in the voice processing system 1 according to the sixth mode, in the second mode, when at least one of the first condition and the second condition is not satisfied, the processor 12 further includes a voice signal with a phase opposite to that of the second voice signal Sig2 in the output voice signal Sig3.

[0144] Accordingly, it has the following advantages: by removing the noise around the user that includes the direct voices from other users, it is easier for the user to hear the voices from other users via the communication line.

[0145] In addition, in the voice processing method of the seventh mode, a first voice signal Sig1 via a communication line is obtained (S102: Yes), a second voice signal Sig2 based on the voice collected by the microphone 2 is obtained (S101: Yes), and when both the first condition and the second condition are satisfied (S106: Yes, S108: Yes), a signal with the component corresponding to the first voice signal Sig1 reduced is included in the output voice signal Sig3 and output to the speaker 3 (S109). The first condition is that both the first voice signal Sig1 and the second voice signal Sig2 are voice signals based on the voice emitted by the same person, and the second condition is that the second voice signal Sig2 is clear.

[0146] Accordingly, in an environment where there are multiple users at the same location, when the direct voice from other users and the direct voice in the voice via the communication line are clear, the user mainly hears the direct voice, so it is easy to hear the voice emitted by other users. That is, it has the advantage of not easily impairing the comfort of the conversation even in an environment where there are multiple users at the same location.

[0147] Moreover, the program according to the eighth mode causes one or more processors to execute the voice processing method according to the seventh mode.

[0148] Accordingly, in an environment where there are multiple users at the same location, when the direct voice from other users and the direct voice in the voice via the communication line are clear, the user mainly hears the direct voice, so it is easy to hear the voice emitted by other users. That is, it has the advantage of not easily impairing the comfort of the conversation even in an environment where there are multiple users at the same location.

[0149] Industrial Applicability

[0150] The voice processing system and the like of the present disclosure can be applied to a system and the like for processing the voice emitted from a speaker.

[0151] Description of Reference Numerals

[0152] 1 Voice processing system

[0153] 10 First input I / F

[0154] 11 Second input I / F

[0155] 12 Processor

[0156] 121 Clarity calculation unit

[0157] 122 Clarity determination unit

[0158] 123 First feature amount calculation unit

[0159] 124 Second feature amount calculation unit

[0160] 125 Speech identity determination unit

[0161] 126 Output voice determination unit

[0162] 127 Output voice control unit

[0163] 128 External sound input switching unit

[0164] 129 ANC control unit

[0165] 13 Memory

[0166] 2 Microphone

[0167] 3 Speaker

[0168] 100 Conference system

[0169] Sig1 First voice signal

[0170] Sig2 Second voice signal

[0171] Sig3 Output voice signal

[0172] V1, V2 Voice

Claims

1. A voice processing system, wherein, Comprising: A first input interface for obtaining a first voice signal via a communication line; A second input interface for obtaining a second voice signal based on the voice collected by a microphone; And A signal processing circuit for outputting an output voice signal based on the first voice signal and the second voice signal to a speaker, The signal processing circuit includes a signal with a reduced component corresponding to the first voice signal in the output voice signal when both the first condition and the second condition are satisfied. The first condition is that both the first voice signal and the second voice signal include voice signals based on the voice emitted by the same person, and the second condition is that the second voice signal is clear.

2. The voice processing system according to claim 1, wherein The signal processing circuit includes the first voice signal in the output voice signal and does not include the second voice signal in the output voice signal when at least one of the first condition and the second condition is not satisfied.

3. The voice processing system according to claim 1 or 2, wherein The signal processing circuit determines whether the first condition is satisfied based on the correlation between the component corresponding to the vowel in the first voice signal and the component corresponding to the vowel in the second voice signal.

4. The voice processing system according to claim 1 or 2, wherein The signal processing circuit determines whether the second condition is satisfied based on the component corresponding to the vowel in the second voice signal.

5. The voice processing system according to claim 1 or 2, wherein The signal processing circuit includes the second voice signal in the output voice signal when both the first condition and the second condition are satisfied.

6. The voice processing system according to claim 2, wherein The signal processing circuit further includes a voice signal with a phase opposite to that of the second voice signal in the output voice signal when at least one of the first condition and the second condition is not satisfied.

7. A voice processing method, wherein Obtaining a first voice signal via a communication line, Obtaining a second voice signal based on the voice collected by a microphone, When both the first condition and the second condition are satisfied, including a signal with a reduced component corresponding to the first voice signal in the output voice signal and outputting it to a speaker. The first condition is that both the first voice signal and the second voice signal include voice signals based on the voice emitted by the same person, and the second condition is that the second voice signal is clear.

8. A program for causing one or more processors to execute the voice processing method according to claim 7.

Citation Information

Patent Citations

  • Voice communication device and voice communication method

    JP2012108587A