Sound processing method and sound processing device

The sound processing method addresses the issue of acoustic interference in remote conferencing by canceling local space characteristics and adding desired acoustics, providing a preferred acoustic environment with aligned audio and visual directions.

JP2025138101APending Publication Date: 2025-09-25YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024036935
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing audio communication devices output audio signals with localized sound images in a virtual space, but users are affected by the acoustic characteristics of their own space, making it difficult to sense the acoustic space at the far end.

Method used

A sound processing method that acquires a sound signal, separates and cancels the acoustic characteristics of the local space, and adds desired acoustic characteristics to create a preferred acoustic environment for remote conferencing.

Benefits of technology

Enables remote conferencing in a preferred acoustic environment, unaffected by the user's local space, with aligned audio and visual directions, and a sense of width and depth of sound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025138101000001_ABST
    Figure 2025138101000001_ABST
Patent Text Reader

Abstract

To provide a sound processing method that enables a user to hold a remote conference in a preferred acoustic environment without being affected by the acoustic characteristics of the user's own space.SOLUTION: A sound processing method comprises: acquiring a first sound signal including voice of a speaker in a first space; acquiring a second sound signal including at least a direct sound component of the speaker's voice from the first sound signal; performing at least one of a cancellation process that cancels acoustic characteristics in the second space and an addition process that adds acoustic characteristics in a desired space to the second sound signal; and outputting the second sound signal that has been subjected to at least one of the cancellation process and the addition process in the second space.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound processing method and a sound processing device used in a remote conference. [Background technology]

[0002] Patent Document 1 discloses an audio communication device that performs sound image localization processing on audio signals so that audio signals from N (N is a natural number of 2 or more) speakers are localized at different positions in a virtual space. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-047223 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the audio communication device in Patent Document 1 outputs N audio signals with sound images localized at different positions in a virtual space, and background noise. Therefore, with the audio communication device in Patent Document 1, the user is inevitably affected by the acoustic characteristics of their own space (the space in which they are located), making it difficult for them to sense the acoustic space at the far end.

[0005] An object of the present invention is to provide a sound processing method that enables a user to hold a remote conference in a preferred acoustic environment without being affected by the acoustic characteristics of the user's own space. [Means for solving the problem]

[0006] A sound processing method according to one embodiment of the present invention includes acquiring a first sound signal including a speaker's voice in a first space, acquiring a second sound signal from the first sound signal including at least a direct sound component of the speaker's voice, performing at least one of a cancellation process that cancels acoustic characteristics in the second space and an addition process that adds acoustic characteristics in a desired space to the second sound signal, and outputting the second sound signal that has been subjected to at least one of the cancellation process and the addition process in the second space. [Effects of the Invention]

[0007] According to the sound processing method of the present invention, a remote conference can be held in a preferred acoustic environment without being affected by the acoustic characteristics of the local space. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram showing an example of the configuration of a sound processing system 1A according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of the configuration of sound processing devices 10 and 20 according to a first embodiment. [Figure 3] 1 is a functional block diagram illustrating an example of a sound processing method according to a first embodiment. [Figure 4] 4 is a flowchart illustrating an example of a sound processing method according to the first embodiment. [Figure 5] 3A and 3B are diagrams illustrating the relationship between a direct sound signal and an indirect sound signal of an audio signal. [Figure 6] 1 is a functional block diagram illustrating an example of a sound processing method according to a first embodiment. [Figure 7] FIG. 10 is a block diagram showing an example of the configuration of a sound processing system 1B according to a first modification. [Figure 8] FIG. 10 is a functional block diagram showing an example of a sound processing method according to Modification 2. [Figure 9] FIG. 11 is a functional block diagram showing an example of a sound processing method according to Modification 3. [Figure 10] FIG. 11 is a functional block diagram showing an example of a sound processing method according to Modification 4. [Figure 11] FIG. 13 is a functional block diagram showing an example of a sound processing method according to Modification 5. [Figure 12] FIG. 13 is a functional block diagram showing an example of a sound processing method according to Modification 6. [Figure 13] FIG. 10 is a functional block diagram illustrating an example of a sound processing method according to a second embodiment. [Figure 14] 10 is a diagram showing input / output characteristics in the conversion unit 323 by a solid line. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, a signal processing device according to an embodiment of the present invention will be described with reference to the drawings. In each drawing, the same parts are assigned the same reference numerals. From Modification 1 onwards, a description of matters common to the first embodiment will be omitted, and only the differences will be described. In particular, similar actions and effects resulting from similar configurations will not be mentioned in each embodiment.

[0010] First Embodiment FIG. 1 is a block diagram showing an example of the configuration of a sound processing system 1A according to the first embodiment.

[0011] As shown in Fig. 1, the sound processing system 1A is composed of two sound processing devices 10 and 20, and two personal computers (PCs) 11 and 21. The sound processing device 10 is connected to a microphone MIC1 and three speakers SP1-SP3. The sound processing device 20 is connected to a microphone MIC2 and three speakers SP4-SP6. The PCs 11 and 21 are each connected to a communication network 5. As a result, the PCs 11 and 21 communicate data with each other through the network 5.

[0012] Next, the configurations of the conference rooms ROOMa and ROOMb in which the sound processing system 1A is installed will be described. The conference rooms ROOMa and ROOMb are spaces where audio conferences are held. The sound processing device 10 and PC 11 are installed in the conference room ROOMa. The speaker 100a holds an audio conference using speakers SP1-SP3, a microphone MIC1, the sound processing device 10, and the PC 11. The sound processing device 20 and PC 21 are installed in the conference room ROOMb. The speaker 100b holds an audio conference using speakers SP4-SP6, a microphone MIC2, the sound processing device 20, and the PC 21. In the following description, the conference room ROOMa will be referred to as the far-end side, and the conference room ROOMb will be referred to as the near-end side. The conference room ROOMa is an example of a first space of the present invention. The conference room ROOMb is an example of a second space of the present invention.

[0013] FIG. 2 is a block diagram showing an example of the configuration of the sound processing devices 10 and 20 according to the first embodiment.

[0014] The sound processing device 10 includes a processor 101, a memory 102, an interface (I / F) 103, and an audio interface (audio I / F) 104. The sound processing device 20 includes a processor 201, a memory 202, an interface (I / F) 203, and an audio interface (audio I / F) 204. The sound processing device 20 has the same structure and functions as the sound processing device 10, so a description thereof will be omitted.

[0015] The memory 102 is a storage medium that stores an operation program for the processor 101. The processor 101 reads the operation program from the memory 102 and performs various operations. The program does not have to be stored in the memory 102. For example, the program may be stored in a storage medium of an external device such as a server. In this case, the processor 101 simply reads the program from the server and executes it each time.

[0016] The processor 101 receives a sound signal acquired by the microphone MIC1 via the audio I / F 104. The processor 101 performs predetermined sound processing on the sound signal acquired by the microphone MIC1 and outputs the sound signal to the I / F 103.

[0017] The I / F 103 is a communication I / F such as USB, HDMI (registered trademark), or Bluetooth (registered trademark). The I / F 103 is connected to the PC 11 via, for example, USB. The I / F 103 outputs the sound signal acquired from the processor 101 to the PC 11.

[0018] The PC 11 transmits the sound signal processed by the sound processing device 10 to the PC 21 in the conference room ROOMb via the network 5.

[0019] With this configuration, the sound signal acquired in the conference room ROOMa is transmitted to the conference room ROOMb.

[0020] Furthermore, the PC 11 acquires the sound signal processed by the sound processing device 20 from the PC 21 via the network 5. The PC 11 transmits the sound signal acquired from the PC 21 to the sound processing device 10 via the I / F 103.

[0021] The processor 101 performs predetermined sound processing on the sound signal acquired from the PC 21, and outputs the sound signal to the speakers SP1-SP3.

[0022] With this configuration, the sound signal acquired in the conference room ROOMb is transmitted to the conference room ROOMa.

[0023] With the above configuration, a speaker 100a in a conference room ROOMa can hold a remote conference with a speaker 100b in a conference room ROOMb.

[0024] (Specific details of sound processing) The sound processing system 1A processes sound signals in a remote conference as follows: From here on, we will assume a situation in which speakers 100a and 100b are holding a remote conference, and explain a sound processing method for outputting a sound signal acquired in a conference room ROOMa to a conference room ROOMb.

[0025] Fig. 3 is a functional block diagram showing an example of a sound processing method according to the first embodiment. Fig. 4 is a flowchart showing an example of a sound processing method according to the first embodiment. Fig. 5 is a diagram showing the relationship between a direct sound signal and an indirect sound signal of an audio signal.

[0026] 3, the processor 101 is functionally composed of a sound signal acquisition unit 1011 and a separation unit 1012. The processor 201 is functionally composed of a cancellation processing unit 2013, an attachment unit 2014, and an output unit 2015. The sound signal acquisition unit 1011 and the separation unit 1012 are programs executed by the processor 101, and respectively realize sound signal acquisition processing and sound source separation processing. The cancellation processing unit 2013, the attachment unit 2014, and the output unit 2015 are programs executed by the processor 201, and respectively realize cancellation processing, attachment processing, and output processing.

[0027] The sound signal acquisition unit 1011 acquires a first sound signal from the MIC 1 (S001). The first sound signal is a sound signal that includes a direct sound component of the voice of the speaker 100a (hereinafter simply referred to as a "direct sound component") and an indirect sound component of the voice of the speaker 100a (hereinafter simply referred to as an "indirect sound component"). Note that the first sound signal may also include a noise component other than a human voice (hereinafter simply referred to as a "noise component").

[0028] Next, the separation unit 1012 separates the second sound signal from the first sound signal acquired by the sound signal acquisition unit 1011 (S002). The second sound signal is a sound signal that includes at least a direct sound component.

[0029] For example, the separator 1012 includes an FIR filter and obtains a second sound signal by extracting only the direct sound component from the first sound signal.

[0030] As shown in Fig. 5, consider that speaker 100a speaks and sounds S1 and S2 are picked up by microphone MIC1. In this case, sounds S1 and S2 are picked up by microphone MIC1 as direct sounds Sd1 and Sd2, but other parts of the sounds are reflected by the walls and desks of conference room ROOMa and picked up as indirect sounds Si1 and Si2.

[0031] Now, consider that speaker 100a utters voice S1 at a first time and voice S2 at a second time. In this case, the waveform of the voices picked up by microphone MIC1 is as shown in Fig. 5. That is, voice S1 output at the first time is detected at time t1, and voice S2 output at the second time is detected later at time t2. In both cases, the waveforms have a high peak value at the time of reception and decay over time. This is because direct sound is picked up by microphone MIC1 via the shortest path from speaker 100a and is picked up directly in front of microphone MIC1, so it has a high peak value and is detected early.

[0032] In contrast, indirect sound travels through various paths before reaching microphone MIC1 from speaker 100a, and is therefore picked up later than direct sound. In addition, the power decreases as the path lengthens, resulting in a waveform with an attenuated peak value.

[0033] As a result, the waveform in Fig. 5 can be considered to be a waveform obtained by combining the waveforms of the direct sounds Sd1 and Sd2 and the waveforms of the indirect sounds Si1 and Si2. Therefore, the separation unit 1012 extracts the sound signal consisting of the direct sounds Sd1 and Sd2 as a direct sound component (second sound signal). The separation unit 1012 is an FIR filter for canceling the acoustic characteristics of the conference room ROOMa. The separation unit 1012 acquires the acoustic characteristics of the conference room ROOMa in advance as an impulse response and obtains an inverse filter of the impulse response. The separation unit 1012 extracts the second sound signal by convolving the inverse filter with the input first sound signal.

[0034] The separator 1012 transmits the second sound signal to the cancellation processor 2013 .

[0035] The cancellation processing unit 2013 performs cancellation processing on the second sound signal (S003). The cancellation processing here refers to processing for canceling the acoustic characteristics of the conference room ROOMb.

[0036] For example, a user of the sound processing system 1A acquires an impulse response of the conference room ROOMb in advance and stores it in the memory 202 of the sound processing device 20. The cancellation processor 2013 reads the impulse response of the conference room ROOMb from the memory 202 and calculates an inverse filter of the impulse response. Next, the cancellation processor 2013 performs cancellation processing by convolving the inverse filter with the second sound signal.

[0037] Next, the cancellation processor 2013 transmits the second sound signal that has been subjected to the cancellation process to the adding unit 2014 and the output unit 2015, respectively.

[0038] The imparting unit 2014 performs an imparting process to impart acoustic characteristics of a desired space to the second sound signal after the cancellation process (S004). The desired space is, for example, a space preferred by the speaker 100b (for example, a large conference room).

[0039] For example, the PC 21 accepts a selection of a preferred space from the speaker 100b and transmits acoustic characteristics associated with the space to the sound processing device 20. The assigning unit 2014 assigns the acoustic characteristics selected by the speaker 100b to the second sound signal. Alternatively, for example, the assigning unit 2014 may assign acoustic characteristics corresponding to a virtual space set by the speaker 100a to the second sound signal. Specifically, the assigning unit 2014 acquires information about the virtual space set by the speaker 100a (e.g., a music hall) from the PC 11 and assigns acoustic characteristics associated with the virtual space to the second sound signal. In this case, the acoustic characteristics of the virtual space may be generated by the assigning unit 2014 or may be stored in advance in the memory 202. The assigning unit 2014 outputs the second sound signal to which the acoustic characteristics of the predetermined space have been assigned to the output unit 2015 (S005), and the process ends (END).

[0040] The output unit 2015 outputs the second sound signal (direct sound component) received from the cancellation processing unit 2013 to the speaker SP4, which is a monaural speaker. The output unit 2015 also outputs the second sound signal (component to which desired acoustic characteristics have been added) received from the addition unit 2014 to the speakers SP5 and SP6, which are stereo speakers.

[0041] As described above, in the sound processing method according to the first embodiment, it is possible to replace the acoustic characteristics of a sound signal acquired at the far-end side with acoustic characteristics preferred by the user and output the signal to the near-end side. More specifically, the acoustic characteristics are replaced by extracting only the direct sound component from the sound signal collected at the far-end side, performing cancellation processing on the extracted direct sound component, and outputting a sound signal obtained by further convolving the desired acoustic characteristics to the near-end side. This allows the user of the sound processing system 1A to have a new customer experience, in which the user perceives themselves as if they are holding a remote conference in their preferred acoustic environment without being affected by the acoustic characteristics of the far-end side and their own space. Furthermore, by outputting a second sound signal that has not been subjected to the addition processing from the monaural speaker, the user at the near-end side can perceive only the direct sound component of the voice of the speaker at the far-end side at the position of the monaural speaker (for example, directly in front of the monaural speaker). For example, when a video of a speaker on the far-end side is displayed on a display (not shown) on the near-end side, the direction (visual direction) of the video of the speaker displayed on the display and the direction (auditory direction) of the speaker's voice are aligned, making it easier for the user on the near-end side to hear the voice of the speaker on the far-end side. Furthermore, by outputting the second sound signal to which the addition processing has been applied from stereo speakers, the user on the near-end side can also perceive a sense of width to the left and right, and can hear indirect sound with a wider width.

[0042] Note that the speakers installed in conference room ROOMb are not limited to the above example. For example, the speaker installed in conference room ROOMb may be only a monaural speaker SP4. In this case, the output unit 2015 mixes the second sound signal (direct sound component) received from the cancellation processing unit 2013 and the second sound signal (component to which desired acoustic characteristics have been added) received from the adding unit 2014, and outputs the result to the speaker SP4. Also, the speakers installed in conference room ROOMb may be only stereo speakers SP5 and SP6. In this case, the output unit 2015 distributes the second sound signal (direct sound component) received from the cancellation processing unit 2013 to the speakers SP5 and SP6 and outputs the result. Also, the output unit 2015 distributes the second sound signal (component to which desired acoustic characteristics have been added) received from the adding unit 2014 to the speakers SP5 and SP6 and outputs the result.

[0043] Furthermore, the programs executed by the sound processing devices 10 and 20 are not limited to the above examples. Fig. 6 is a functional block diagram showing an example of a sound processing method according to the first embodiment. As shown in Fig. 6, for example, the sound processing device 10 may be configured to execute all of the sound signal acquisition processing, separation processing, cancellation processing, addition processing, and output processing. And, of course, the sound processing device 20 may be configured to execute all of the sound signal acquisition processing, separation processing, cancellation processing, addition processing, and output processing.

[0044] The adding unit 2014 may also convolve the acoustic characteristics of the conference room ROOMa into the second sound signal (direct sound component) on which the cancellation process has been performed. For example, a user of the sound processing system 1A stores the acoustic characteristics of the conference room ROOMa in advance in the memory 202. The adding unit 2014 acquires the acoustic characteristics of the conference room ROOMa from the memory 202 and convolves the acoustic characteristics into the second sound signal on which the cancellation process has been performed.

[0045] As a result, the sound processing system 1A can completely eliminate the influence of the acoustic characteristics of the near-end space in the near-end space and reproduce the acoustic environment of the far-end side. Specifically, cancellation processing is performed on all sound signals output to the near-end side by the sound processing system 1A. Therefore, the user on the near-end side is not affected by indirect sounds generated when the sound signal output from the speaker is reflected by the walls, desks, etc. of the near-end space. Therefore, the user of the sound processing system 1A can have a customer experience of having a conversation while feeling only the reverberation of the acoustic environment of the far-end side without being affected by the acoustic characteristics of his or her own space. Furthermore, since the user on the near-end side has a conversation while feeling the acoustic environment of the far-end space, the user can have a conversation while imagining how his or her voice sounds to the user on the far-end side. Therefore, the user of the sound processing system 1A can have a new customer experience of having a conversation while appropriately adjusting his or her own voice reproduced on the far-end side.

[0046] Variation 1 7 is a block diagram showing an example of the configuration of a sound processing system 1B according to Modification 1. The sound processing system 1B differs from the sound processing system 1A in that it further includes a cloud device 6 connected to the network 5.

[0047] The cloud device 6 includes a memory and a processor (not shown). The memory and the processor included in the cloud device 6 have the same functions as the memory 102 and the processor 101 included in the sound processing device 10.

[0048] The cloud device 6 executes a program in place of the sound processing devices 10 and 20. Specifically, the processor included in the cloud device 6 has the functions of a separation unit, a cancellation processing unit, and an attachment unit.

[0049] In this way, the sound processing system 1B of Modification 1 performs all sound signal processing on the cloud, which enables users to hold remote conferences even in spaces where no sound processing device with sound signal processing capabilities is installed.

[0050] In this modification, the number of spaces in which a remote conference is held is not limited to two. For example, a remote conference may be held by connecting PCs installed in three different spaces to the network 5.

[0051] If there are three or more spaces connected in a remote conference, it is preferable that the cloud device 6 set up the same virtual space for the PCs installed in each space. This allows the cloud device 6 to assign acoustic characteristics associated with the set common virtual space to the sound signals acquired from each space. Therefore, all users in each space can share the same acoustic environment even if they are in different spaces.

[0052] Variation 2 8 is a functional block diagram showing an example of a sound processing method according to Modification 2. The sound processing device 10 of the sound processing system 1C according to Modification 2 differs from the sound processing device 10 of the sound processing system 1A in that it further has a function of extracting an indirect sound component in addition to a direct sound component from a first sound signal. More specifically, the separation unit 1012 of the sound processing device 10 of the sound processing system 1C separates the direct sound component and the indirect sound component from the first sound signal. Furthermore, the processor 201 of the sound processing system 1C functionally differs from the processor 201 of the sound processing system 1A in that it does not include an addition unit 2014. Note that in this example, the direct sound component and the indirect sound component are the second sound signal in the present invention.

[0053] For example, the separation unit 1012 extracts, as an indirect sound component, a sound signal consisting of indirect sounds Si1 and Si2 shown in Fig. 5. For example, the separation unit 1012 obtains an impulse response by excluding the direct sound component from the impulse response of the conference room ROOMa acquired in advance, generates a pseudo indirect sound based on the obtained impulse response, and extracts the pseudo indirect sound as an indirect sound component.

[0054] The separation unit 1012 transmits the separated direct sound component and indirect sound component to the cancellation processing unit 2013 .

[0055] The cancellation processing unit 2013 performs cancellation processing on the direct sound component and the indirect sound component, and outputs the result to the output unit 2015 .

[0056] The output unit 2015 outputs the direct sound component that has been subjected to the cancellation process to the monaural speaker SP4, and outputs the indirect sound component that has been subjected to the cancellation process to the stereo speakers SP5 and SP6.

[0057] In this way, in the sound processing method of Modification 2, it is possible to reproduce indirect sound generated on the far-end side on the near-end side. Also, since there is no need to perform addition processing on the second sound signal, it is possible to reduce the load on the sound processing device.

[0058] Variation 3 9 is a functional block diagram showing an example of a sound processing method according to Modification 3. The sound processing device 10 of the sound processing system 1D according to Modification 3 differs from the sound processing device 10 of the sound processing system 1C in that it further has a function of separating a noise component from the first sound signal in addition to a direct sound component and an indirect sound component. More specifically, the separation unit 1012 of the sound processing device 10 of the sound processing system 1D separates the direct sound component, the indirect sound component, and the noise component from the first sound signal.

[0059] For example, the separation unit 1012 separates the noise component by removing the direct sound component and the indirect sound component from the first sound signal. The separation unit 1012 removes the noise component by, for example, a spectral subtraction method.

[0060] The separation unit 1012 separates the power spectrum of the noise component of the conference room ROOMa as the noise component in the spectral subtraction method. Alternatively, the separation unit 1012 may separate sounds of sound sources other than speech, such as footsteps of people in the vicinity, as the noise component.

[0061] The separation unit 1012 transmits the separated direct sound component, indirect sound component, and noise component to the cancellation processing unit 2013.

[0062] The cancellation processing unit 2013 performs cancellation processing on the direct sound component, the indirect sound component, and the noise component, and outputs the results to the output unit 2015.

[0063] In this way, the sound processing method of Variation 3 makes it possible to reproduce noise generated on the far-end side on the near-end side. This allows the user on the near-end side to converse while listening to not only the acoustic environment of the far-end side but also the noise generated on the far-end side. Speakers often speak quietly when the surrounding environment is quiet and speak loudly when the surrounding environment is noisy. According to the sound processing method of Variation 3, the speaker on the near-end side can converse while listening to environmental sounds such as the hustle and bustle generated on the far-end side and the footsteps of people around them, and therefore can converse at a volume of voice that corresponds to the noise level on the far-end side.

[0064] Variation 4 FIG. 10 is a functional block diagram showing an example of a sound processing method according to Modification 4. The processor 101 of the sound processing system 1E according to Modification 4 functionally further includes a background information acquisition unit 1016. The sound processing device 10 of the sound processing system 1E differs from the sound processing device 10 of the sound processing system 1A in that it is connected to a camera CAM1 and has a function of acquiring background information on the far-end side based on image information. The background information acquisition unit 1016 is a program executed by the processor 101. The sound processing device 20 of the sound processing system 1E also differs from the sound processing device 20 of the sound processing system 1A in that it further has a function of performing an attachment process on a second sound signal based on image information. More specifically, an attachment unit 2014 of the sound processing device 20 of the sound processing system 1E performs an attachment process on the second sound signal based on image information.

[0065] The camera CAM1 captures an image of the conference room ROOMa and transmits the image information to the background information acquisition unit 1016.

[0066] The background information acquisition unit 1016 detects the size of the conference room ROOMa based on the image information. The background information acquisition unit 1016 transmits information about the size of the conference room ROOMa (hereinafter referred to as size information) to the assignment unit 2014.

[0067] The assigning unit 2014 assigns acoustic characteristics corresponding to the size of the conference room ROOMa to the second sound signal based on the size information. The sound processing device 20 may store acoustic characteristics corresponding to the size of the listening environment (conference room) in advance as a table in the memory 202. In this case, the assigning unit 2014 reads out the acoustic characteristics corresponding to the size information from the memory 202 and assigns them to the second sound signal.

[0068] In this way, in the sound processing method of Variation 4, it is possible to acquire background information on the far-end side based on image information captured by a camera and to generate artificial indirect sound that imitates the acoustic environment of the far-end side. As a result, even if, for example, a speaker on the far-end side is using a microphone with a high S / N ratio such as a pin microphone and only direct sound components can be acquired, it is possible to generate indirect sound from the image information, so that the user on the near-end side can have a conversation while feeling the acoustic environment of the far-end side.

[0069] In the sound processing method of Variation 4, the adding process is performed by the sound processing device on the near-end side. In other words, the sound processing device on the far-end side only needs to transmit the direct sound component acquired by the microphone and the background information acquired by the camera to the sound processing device on the near-end side, thereby reducing the load on the sound processing device due to data transmission. This makes it possible to suppress the occurrence of audio delays in remote conferences.

[0070] Variation 5 11 is a functional block diagram showing an example of a sound processing method according to Modification 5. The processor 101 of the sound processing system 1F according to Modification 5 further includes a microphone control unit 1017. The sound processing device 10 of the sound processing system 1F differs from the sound processing device 10 of the sound processing system 1C in that it is connected to two microphones MIC3 and MIC4 and a camera CAM2 and has a function of controlling the microphones MIC3 and MIC4 based on image information. The microphone control unit 1017 is a program executed by the processor 101. In Modification 5, the microphone MIC3 is a directional microphone and the microphone MIC4 is an omnidirectional microphone.

[0071] The camera CAM2 captures an image of the conference room ROOMa and transmits the image information to the microphone control unit 1017.

[0072] The microphone control unit 1017 detects the speaker 100a based on the image information, and transmits a control signal to the microphone MIC3 so as to orient the microphone MIC3 toward the position where the speaker 100a is located.

[0073] The microphone MIC3 has directivity directed toward the speaker 100a and therefore picks up only the direct sound component, whereas the microphone MIC4 is an omnidirectional microphone and therefore picks up a sound signal including the direct sound component, the indirect sound component, and the noise component.

[0074] The separator 1012 separates the indirect sound component and the noise component by subtracting the sound signal acquired by the microphone MIC3 from the sound signal acquired by the microphone MIC4.

[0075] In this way, the sound processing method of Variation 5 can direct the directivity toward the speaker based on image information captured by the camera, and separate a direct sound component with a high S / N ratio. This makes it easier for the user on the near-end to hear the voice of the user on the far-end. Furthermore, the sound processing method of Variation 5 does not require separation of the direct sound component, so the load on the sound processing device during separation processing can be reduced.

[0076] Variation 6 FIG. 12 is a functional block diagram showing an example of a sound processing method according to Modification 6. The processor 101 of the sound processing system 1G according to Modification 6 further includes a speaker position detection unit 1018. The processor 201 of the sound processing system 1G further includes a sound image localization processing unit 2016. The sound processing device 10 of the sound processing system 1G is different from the sound processing device 10 of the sound processing system 1D in that it is connected to a camera CAM3 and has a function of identifying the position of a speaker based on image information. The sound processing device 20 of the sound processing system 1G is different from the sound processing device 20 of the sound processing system 1D in that it has a processing function of localizing a sound image of a sound signal according to the position of the speaker. The speaker position detection unit 1018 and the sound image localization processing unit 2016 are programs executed by the processor 101 and the processor 201, respectively.

[0077] The camera CAM3 performs framing processing, for example, by panning, tilting, or zooming, so that only the speaker 100a appears in the image. Alternatively, the camera CAM3 performs framing processing so that all participants in the conference room ROOMa appear in the image in addition to the speaker 100a. The image captured by the camera CAM3 is preferably displayed on a display unit (not shown) in the conference room ROOMb.

[0078] The speaker position detection unit 1018 identifies the position of the speaker in the conference room ROOMa based on the video information captured by the camera CAM3. The speaker position detection unit 1018 transmits information indicating the detected speaker position (hereinafter referred to as position information) to the sound image localization processing unit 2016.

[0079] The sound image localization processor 2016 performs sound image localization processing on the second sound signal (direct sound components and indirect sound components) received from the cancellation processor 2013 based on the position information. The sound image localization processing is a process for localizing a sound image so that the sounds output from the speakers SP5 and SP6 appear as if they were generated at a predetermined position. The sound image localization processor 2016 achieves the sound image localization processing by convolving a head-related transfer function (HRTF) with the second sound signal. The sound image localization processor 2016 acquires the HRTF from, for example, a memory provided in the sound processing device 20, a network, or an external storage medium.

[0080] Here, the head-related transfer function will be explained in more detail. The head-related transfer function is a function that expresses the transfer characteristics from the position of a sound source to the user's head (specifically, the user's left ear and right ear). There are two head-related transfer functions: one from the sound source to the left ear and one to the right ear. The sound image localization processing unit 2016 convolves the sound signals corresponding to the L channel and the R channel (second sound signals) with the head-related transfer function to the left ear and the head-related transfer function to the right ear, respectively, and transmits the results to the output unit 2015.

[0081] The output unit 2015 outputs the second sound signal convolved with the head-related transfer function to speakers SP5 and SP6, which are stereo speakers.

[0082] Furthermore, the sound image localization processing unit 2016 may perform panning processing on the second sound signal received from the cancellation processing unit 2013, based on the position information. The panning processing is processing for adjusting the level ratio of the sound signals supplied to the speakers SP5 and SP6, and can change the sound image localization position of the sound signal in the horizontal direction.

[0083] In this way, the sound processing method of Variation 6 makes it possible to change the position where the sound image of the sound signal output on the near-end side is localized. More specifically, by convolving the head-related transfer function with the direct sound component and the indirect sound component extracted from the sound signal picked up on the far-end side, the user on the near-end side can perceive the voice of the speaker on the far-end side as if it were actually coming from the position where the speaker is located. Furthermore, the user on the near-end side can perceive the indirect sound from the far-end side as if it were actually reflected from the wall or the like in the room on the far-end side. Therefore, according to the sound processing system 1G of Variation 6, by outputting the sound signal that has been subjected to sound image localization processing from stereo speakers, it is possible to reproduce the acoustic environment on the far-end side in the space on the near-end side with accuracy that cannot be achieved with a monaural speaker.

[0084] Second Embodiment The sound processing device according to the second embodiment further executes voice clarity processing. FIG. 13 is a functional block diagram showing an example of a sound processing method according to the second embodiment. The processor 201 of the sound processing system 1H according to the second embodiment further functionally includes a voice clarity processing unit 300. The sound processing device 20 of the sound processing system 1H differs from the sound processing device 20 of the sound processing system 1D in that it also has a function of clarifying the voice of the speaker acquired on the far-end side. The voice clarity processing unit 300 is a program executed by the processor 201.

[0085] 13, the voice clarity processing unit 300 has a multiplication processing unit 310, a gain determination unit 320, and a multiplier 330. The direct sound component to be processed by the voice clarity processing unit 300 is supplied to the multiplication processing unit 310 and the gain determination unit 320, respectively.

[0086] Multiplication processing unit 310 extracts three components in the voice band from the direct sound of the speaker's voice and multiplies them by a predetermined gain coefficient, and has band pass filters (BPF) 311, 313, 315, multipliers 312, 314, 316, and adder 317.

[0087] Of these, BPF311 has a passband of approximately 200 to 500 Hz and extracts the first formant component from the speaker's direct sound component. BPF313 has a passband of approximately 2 kHz to 3 kHz and extracts the second formant component from the speaker's direct sound component. BPF315 has a passband of approximately 4 kHz to 12 kHz and extracts harmonic components that are 2 to 4 times the second formant from the speaker's direct sound component.

[0088] Multiplier 312 multiplies the output signal of BPF 311 by a preset gain coefficient and outputs the result. Similarly, multipliers 314 and 316 multiply the output signals of BPFs 313 and 315 by preset gain coefficients, respectively, and output the result. Note that when processing analog signals, multipliers 312, 314, and 316 may each be used as a gain amplifier.

[0089] The adder 317 adds the signals output from the multipliers 312, 314, and 316 together and outputs the result as a sound signal Fa.

[0090] Here, since the passbands of BPFs 311, 313, and 315 are as described above, the passbands are isolated from one another, and the general shape of the frequency characteristics of the three BPFs can be considered to have peaks (mountains) at the center frequencies of each passband, with bottoms (valleys) between these peaks.

[0091] The gain determination unit 320 determines the gain for the sound signal Fa in accordance with the level of the sound signal In acquired by the sound signal acquisition unit 1011, and includes a BPF 321, a level detection unit 322, a conversion unit 323, and a smoothing unit 324. The BPF 321 has a passband of approximately 125 to 175 Hz and extracts a signal indicating the magnitude of the voice component in the sound signal In. This is because it has been experimentally confirmed that this band is highly likely to contain voice frequency components. The passband of the BPF 321 roughly corresponds to the vibration frequency of the vocal cords during conversation. Because voice is produced by the vibration of the vocal cords resonating in the vocal tract, the vibration frequency component of the vocal cords is the source of voice and can be used as one indicator of the magnitude of the voice component.

[0092] The level detection unit 322 detects and outputs the maximum value (maximum value in terms of absolute value) of the signal level output from the BPF 321 at regular intervals. The level detection unit 322 outputs the maximum value of the signal level from the BPF 321, for example, every 5 milliseconds. If there is sufficient processing capacity, the level detection unit 322 may output the absolute value for each sample rather than every 5 milliseconds. Furthermore, the level detection unit 322 is not limited to the maximum value of the signal level output from the BPF 321, and may instead calculate and output an effective value (root mean square value) or an envelope (envelope curve). In any case, the level detection unit 322 may be configured to output a level based on a signal indicating a sound component.

[0093] The conversion unit 323 converts the signal (maximum value) output from the level detection unit 322 into a gain for the sound signal Fa.

[0094] Fig. 14 is a diagram showing the input / output characteristics of conversion unit 323 using a solid line. In Fig. 14, the horizontal axis represents the input level to conversion unit 323, i.e., the signal (maximum value) level output from level detection unit 322, and the vertical axis represents the output level of conversion unit 323. Here, what is output from conversion unit 323 is a gain indicated by the ratio of output level / input level in the conversion characteristics. Note that the dashed line in Fig. 14 indicates the case where the input level becomes the output level as is.

[0095] The input / output characteristics of the conversion unit 323 show that in the region where the input level is equal to or less than the threshold th, the gain is equal to or greater than "1," i.e., above the dashed line, and the gain increases as the input level decreases. On the other hand, in the region where the input level exceeds the threshold th, the gain is smaller than "1," and the gain decreases rapidly as the input level increases. Therefore, in the region where the input level is equal to or less than the threshold th, a gain that increases the output level is output, and in the region where the input value exceeds the threshold th, a gain that decreases the output level is output.

[0096] The conversion unit 323 may be configured to store the characteristics shown in FIG. 14 as a table in advance and read out the gain corresponding to the input level by referring to the table, or may be configured to define the characteristics as a function and substitute the input level into the function to determine the gain by calculation.

[0097] Here, if the level detection unit 322 outputs a maximum value at regular intervals (every 5 milliseconds in the above example), the conversion unit 323 will also output a gain corresponding to the maximum value at the same intervals, and the gain will fluctuate at regular intervals. Therefore, the smoothing unit 324 smoothes the gain that fluctuates at regular intervals in this way and outputs it as gain Ka. Smoothing is achieved by using a moving average, a low-pass filter, or the like.

[0098] The multiplier 330, which is a level adjustment unit, multiplies the sound signal Fa, which is the output of the adder 317, by the gain Ka, which is the output of the gain determination unit 320, and outputs the result to the cancellation processing unit 2013 as a sound signal Fb.

[0099] It is generally believed that the higher the ratio of the peak level of a formant to the bottom level between one formant and another, the clearer the speech (see, for example, Japanese Patent Application Laid-Open No. 2008-145940).

[0100] In the sound processing device 20 according to this embodiment, the sound signal Fa extracted by the multiplication processing unit 310 is a direct sound component, specifically, a first formant component, a second formant component, and harmonic components of the second formant. The level of this sound signal Fa is adjusted by a gain Ka that is the output of the gain determination unit 320, and is output as a sound signal Fb.

[0101] In this configuration, when the level of the voice component contained in the sound signal In is small, the gain Ka is increased so as to raise the output level, and as a result, the peak level of each formant is raised relative to the bottom level, and the sound signal Fb becomes dominant over the sound signal In, thereby making the voice clearer.

[0102] On the other hand, if the level of the audio component is sufficiently large, the gain Ka cannot be increased, and therefore the audio signal Fb added to the audio signal In becomes small (or becomes 0), thereby suppressing the emphasis of the audio.

[0103] In this way, the sound processing method of the second embodiment can clarify the voice of the far-end speaker, which allows the near-end speaker to have a conversation while sensing the acoustic environment of the far-end speaker, and even if the far-end speaker is noisy, the near-end speaker can clearly hear the voice of the far-end speaker, providing a new customer experience.

[0104] Finally, the description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0105] 1A, 1B, 1C, 1D, 1E, 1F, 1G, 1H: sound processing system, ROOMa, ROOMb: conference room, 100a, 100b: speaker, MIC1 to MIC4: microphones, SP1 to SP6: speakers, CAM1 to CAM3: cameras, 5: network, 6: cloud device, 10, 20: sound processing device, 11, 21: PC, 101, 201: processor, 102, 202: memory, 103, 203: I / F, 104, 204: audio I / F, 300: voice clarity processing unit, 310: multiplication processing unit, 311, 313, 315, 321: band pass filters, 312, 314, 316, 330: multipliers, 317: adder, 320: gain determination unit, 322: level detection unit, 323: conversion unit, 324: smoothing unit, 1011: sound signal acquisition unit, 1012: separation unit, 1016: background information acquisition unit, 1017: microphone control unit, 1018: speaker position detection unit, 2013: cancellation processing unit, 2014: addition unit, 2015: output unit, 2016: sound image localization processing unit

Claims

1. obtaining a first sound signal including a voice of a speaker in a first space; obtaining a second sound signal including at least a direct sound component of the speaker's voice from the first sound signal; performing at least one of a cancellation process for canceling acoustic characteristics in a second space and an imparting process for imparting acoustic characteristics in a desired space to the second sound signal; outputting the second sound signal, which has been subjected to at least one of the cancellation processing and the addition processing, in the second space; Sound processing methods.

2. performing a sound source separation process to extract the direct sound component and an indirect sound component other than the direct sound component from the first sound signal, thereby obtaining the second sound signal; The sound processing method according to claim 1 .

3. The sound source separation process further includes separating a noise component from the first sound signal. The sound processing method according to claim 2 .

4. localizing the direct sound component and the indirect sound component at different positions, respectively; The sound processing method according to claim 2 .

5. acquiring background information of the first space; performing the assignment process based on the background information; The sound processing method according to any one of claims 1 to 4.

6. The background information of the first space is acquired based on image information from a camera. The sound processing method according to claim 5 .

7. The background information of the first space is acquired based on information about an arbitrary virtual space. The sound processing method according to claim 5 .

8. performing a process to emphasize the direct sound component; The sound processing method according to any one of claims 1 to 4.

9. performing a process of extracting the direct sound component based on image information from a camera; The sound processing method according to any one of claims 1 to 4.

10. performing predetermined sound processing on the second sound signal based on image information from the camera; The sound processing method according to any one of claims 1 to 4.

11. obtaining a first sound signal including a voice of a speaker in a first space; obtaining a second sound signal including at least a direct sound component of the speaker's voice from the first sound signal; performing at least one of a cancellation process for canceling acoustic characteristics in a second space and an imparting process for imparting acoustic characteristics in a desired space to the second sound signal; outputting the second sound signal, which has been subjected to at least one of the cancellation processing and the addition processing, in the second space; With a processor, Sound processing device.

12. the processor performs a sound source separation process to extract the direct sound component and an indirect sound component other than the direct sound component from the first sound signal, thereby obtaining the second sound signal. The sound processing device according to claim 11 .

13. The processor further separates a noise component from the first sound signal. The sound processing device according to claim 12.

14. The processor further localizes the direct sound component and the indirect sound component at different positions. The sound processing device according to claim 12.

15. The processor further acquires background information of the first space and performs the assignment process based on the background information. The sound processing device according to any one of claims 11 to 14.

16. It also has a camera, The background information of the first space is acquired based on image information of the camera. The sound processing device according to claim 15.

17. The background information of the first space is acquired based on information about an arbitrary virtual space. The sound processing device according to claim 15.

18. The processor further performs processing to emphasize the direct sound component. The sound processing device according to any one of claims 11 to 14.

19. the processor performs processing to extract the direct sound component based on image information from a camera. The sound processing device according to any one of claims 11 to 14.

20. The processor performs predetermined sound processing on the second sound signal based on image information from the camera. The sound processing device according to any one of claims 11 to 14.

Citation Information

Patent Citations

  • Voice communication device

    JP2022047223A