Sound signal processing method and sound signal processing device

The audio signal processing method dynamically adjusts to the environment by estimating room type and adjusting acoustic parameters, effectively addressing the issue of inappropriate gain control and noise suppression in existing devices.

JP7844965B2Active Publication Date: 2026-04-14YAMAHA CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing audio signal processing devices fail to adapt their gain control and noise suppression methods appropriately to the specific environment, either amplifying noise in open spaces or failing to capture quiet voices in closed spaces.

Method used

An audio signal processing method that estimates the room type based on visual input, adjusting acoustic parameters such as AGC and noise reduction accordingly to suit the environment, using image analysis and machine learning to determine whether the space is open or closed.

Benefits of technology

Enables appropriate sound processing by automatically adapting to the environment, ensuring clear voice capture and noise suppression, thereby improving conversation quality without user intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844965000001
    Figure 0007844965000001
  • Figure 0007844965000002
    Figure 0007844965000002
  • Figure 0007844965000003
    Figure 0007844965000003
Patent Text Reader

Abstract

To provide a sound signal processing method capable of performing appropriate sound processing depending on the situation.SOLUTION: A video signal processing method according to an embodiment accepts a sound signal, acquires a first image, estimates room information on the basis of the acquired first image, sets acoustic parameters according to the estimated room information, and performs sound processing based on the set acoustic parameters on the sound signal, and outputs the sound signal after sound processing.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] One embodiment of the present invention relates to a method and apparatus for processing audio signals related to audio signal processing.

Background Art

[0002] Patent Document 1 describes a gain automatic device including a microphone. The gain automatic device detects the level of a user's voice and the level of background noise picked up by the microphone. The gain automatic device sets a gain based on the levels of the user's voice and background noise.

[0003] Patent Document 2 describes a noise gate for suppressing an audio signal. The noise gate calculates the signal level of an input audio signal. The noise gate reduces the gain of an audio signal whose signal level is less than a threshold value.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] The automatic gain device described in Patent Document 1 (hereinafter referred to as Device X) and the noise gate described in Patent Document 2 (hereinafter referred to as Device Y) each perform automatic gain adjustment based on the sound signal. Therefore, Device X and Device Y do not necessarily perform appropriate sound processing according to the situation in which they are used. For example, in a closed space such as a conference room, it is highly likely that everyone in the conference room is a participant in the meeting. Therefore, it is preferable for Device X and Y to amplify the speaker's voice using AGC (Auto Gain Control) so that even a quiet voice of the speaker can be picked up as much as possible. In addition, since it is considered unlikely that everyone in the conference room will make noise, it is unlikely that Device X and Y will pick up noise whose volume has been increased by AGC. On the other hand, for example, in an open space, multiple people with different purposes share the space. Therefore, it is highly likely that people other than the users of Device X and Device Y will make noise. Therefore, it is preferable for Device X and Device Y to suppress noise. However, if Device X and Device Y were to perform AGC in an open space in the same way as in a closed space, it would actually amplify the noise.

[0006] One embodiment of the present invention aims to provide an audio signal processing method that can perform appropriate sound processing depending on the situation. [Means for solving the problem]

[0007] The sound signal processing method relating to the present invention is Receiving an audio signal, Take the first image, Based on the acquired first image, room information is estimated. Set acoustic parameters according to the estimated room information. Sound processing based on the set acoustic parameters is performed on the sound signal. The sound signal that has undergone the aforementioned sound processing is output. [Effects of the Invention]

[0008] According to one embodiment of the sound signal processing method of this invention, it is possible to perform appropriate sound processing depending on the situation. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a block diagram showing an example of the connection between the sound signal processing device 1 and a device different from the sound signal processing device 1. [Figure 2] Figure 2 is a block diagram showing the functional configuration of processor 17. [Figure 3] Figure 3 is a flowchart showing an example of the processing performed by the sound signal processing device 1. [Figure 4] Figure 4 is an example of the first image M1, which shows a closed space. [Figure 5] Figure 5 is an example of the first image M1 showing an open space. [Figure 6] Figure 6 shows the correspondence between room information RI and acoustic parameters SP. [Figure 7] Figure 7 is a block diagram showing the functional configuration of the processor 17b of the sound signal processing device 1b. [Figure 8] Figure 8 is a flowchart showing an example of setting the acoustic parameter SP in the sound signal processing device 1c. [Figure 9] Figure 9 shows the gain adjustment in the sound signal processing device 1d. [Figure 10] Figure 10 is a block diagram showing the functional configuration of the processor 17e of the sound signal processing device 1e. [Figure 11] Figure 11 is a block diagram showing the functional configuration of the processor 17f of the sound signal processing device 1f. [Figure 12] Figure 12 is a block diagram showing the functional configuration of the processor 17h of the sound signal processing device 1h. [Figure 13] Figure 13 is a flowchart showing an example of setting the acoustic parameter SP in the sound signal processing device 1h. [Figure 14] Figure 14 shows an example of image processing in the sound signal processing device 1h.

Best Mode for Carrying Out the Invention

[0010] (First Embodiment) Hereinafter, a sound signal processing method according to the first embodiment will be described with reference to the drawings. FIG. 1 is a block diagram showing an example of the connection between a sound signal processing apparatus 1 and a device (processing apparatus 2) different from the sound signal processing apparatus 1.

[0011] The sound signal processing apparatus 1 is a device for performing a remote conversation by connecting to a processing apparatus 2 such as a PC at a remote location (see FIG. 1). The sound signal processing apparatus 1 is, for example, an information processing apparatus such as a PC. The sound signal processing apparatus 1 executes the sound signal processing method according to the first embodiment.

[0012] As shown in FIG. 1, the sound signal processing apparatus 1 includes an audio interface 11, a general-purpose interface 12, a communication interface 13, a user interface 14, a flash memory 15, a RAM (Random Access Memory) 16, and a processor 17. The processor 17 is, for example, a CPU (Central Processing Unit) or the like.

[0013] The audio interface 11 communicates with an audio device such as a microphone 4 or a speaker 5 via a signal line (see FIG. 1). The microphone 4 acquires the voice of a user (hereinafter referred to as user U) of the sound signal processing apparatus 1. The microphone 4 outputs the acquired voice as a sound signal to the audio interface 11. The audio interface 11, for example, converts a digital sound signal received from the processing apparatus 2 into an analog sound signal. The speaker 5 receives an analog sound signal from the audio interface 11 and outputs a sound based on the received analog sound signal.

[0014] The general-purpose interface 12 is an interface based on a standard such as USB (Universal Serial Bus). The general-purpose interface 12 is connected to the camera 6 as shown in Figure 1. The camera 6 acquires a first image M1 by taking pictures of its surroundings (around the user U). The camera 6 outputs the acquired first image M1 as image data to the general-purpose interface 12.

[0015] The communication interface 13 is a network interface, etc. The communication interface 13 communicates with the processing unit 2 via the communication line 3. The communication line 3 is the Internet or a LAN (Local Area Network), etc. The communication interface 13 and the processing unit 2 communicate wirelessly or via a wired connection.

[0016] The user interface 14 receives input from the user U to the sound signal processing device 1. The user interface 14 is, for example, a keyboard, mouse, or touch panel.

[0017] The flash memory 15 stores various programs. These various programs include, for example, a program for operating the sound signal processing device 1, or an application program for executing sound processing related to the sound signal processing method. However, the flash memory 15 does not necessarily have to store various programs. These various programs may be stored in other devices, such as a server. In this case, the sound signal processing device 1 receives various programs from other devices such as the server.

[0018] The processor 17 executes various operations by reading the program stored in the flash memory 15 into the RAM 16. The processor 17 also performs signal processing related to the sound signal processing method (hereinafter referred to as sound processing P), or processing related to communication between the sound signal processing device 1 and the processing device 2.

[0019] The processor 17 receives an audio signal from the microphone 4 via the audio interface 11. The processor 17 performs sound processing P on the received audio signal. The processor 17 transmits the processed audio signal to the processing unit 2 via the communication interface 13. The processor 17 receives an audio signal from the processing unit 2 via the communication interface 13. The processor 17 transmits the audio signal to the speaker 5 via the audio interface 11. The processor 17 also receives the first image M1 from the camera 6 via the general-purpose interface 12.

[0020] The processing unit 2 is equipped with a speaker (not shown). The speaker of the processing unit 2 outputs sound based on the sound signal received from the sound signal processing unit 1. The user of the processing unit 2 (hereinafter referred to as the interlocutor) listens to the sound output from the speaker of the processing unit 2. The processing unit 2 is equipped with a microphone (not shown). The processing unit 2 transmits the sound signal acquired by the microphone of the processing unit 2 to the sound signal processing unit 1 via the communication interface 13.

[0021] The sound processing P in processor 17 will be described in detail below with reference to the figures. Figure 2 is a block diagram showing the functional configuration of processor 17. Figure 3 is a flowchart showing an example of processing by sound signal processing device 1. Figure 4 is an example of first image M1 showing a closed space. Figure 5 is an example of first image M1 showing an open space. Figure 6 is a diagram showing the correspondence between room information RI and acoustic parameters SP.

[0022] As shown in Figure 2, the processor 17 functionally includes a reception unit 170, an acquisition unit 171, an estimation unit 172, a setting unit 173, a signal processing unit 174, and an output unit 175. The reception unit 170, acquisition unit 171, estimation unit 172, setting unit 173, signal processing unit 174, and output unit 175 perform sound processing P.

[0023] For example, when the processor 17 executes an application program related to sound processing P, it starts sound processing P (Figure 3: START).

[0024] After the start, the acquisition unit 171 acquires an image (hereinafter referred to as the first image M1) (Figure 3: Step S11). The acquisition unit 171 acquires the first image M1 from the camera 6 and outputs it to the estimation unit 172.

[0025] Next, the estimation unit 172 estimates room information RI based on the first image M1 (Figure 3: Step S12). Room information RI is, for example, information indicating the space where user U is located. In this embodiment, information indicating the space where user U is located is, for example, information indicating whether it is a closed space (a space that is not open) or an open space (a space that is open). In other words, in this embodiment, room information RI includes information indicating whether it is an open space or a closed space. A closed space is, for example, an indoor space partitioned by walls, ceilings, etc., such as a conference room. An open space is, for example, a multipurpose space or an open space that is not partitioned by walls, ceilings, etc., such as outdoors.

[0026] The estimation unit 172 estimates room information RI by analyzing the first image M1. The analysis process includes, for example, neural networks (e.g., DNN (Deep Neural Network)). u This is an analysis process using artificial intelligence such as a RAL Network. The estimation unit 172 estimates the room information RI using a trained model that has learned the relationship between the input image and the room information RI through machine learning. Specifically, the estimation unit 172 extracts the features of the first image M1 and outputs them to the trained model. The trained model determines the objects contained in the first image M1 based on, for example, the features contained in the first image M1. Features are, for example, edges or textures in the first image M1. The trained model determines whether the space where the user U is located is a closed space or an open space based on the objects contained in the first image M1.

[0027] In this case, the trained model determines "Room Information RI: Closed Space" when it determines that the first image M1 contains objects specific to a closed space. For example, if camera 6 captures a closed space, the first image M1 is highly likely to capture the boundary B1 between the wall and the ceiling (see Figure 4). Therefore, if the trained model recognizes boundary B1 as an object contained in the first image M1, it determines that the space where user U is located is a closed space. On the other hand, the trained model determines "Room Information RI: Open Space" when it determines that the first image M1 does not contain objects specific to a closed space.

[0028] In the example shown in Figure 4, if camera 6 captures a closed space, there is a high probability that door D is captured in the first image M1. Therefore, if the trained model recognizes door D as an object included in the first image M1, it may determine "Room information RI: Closed space".

[0029] The method by which the sound signal processing device 1 estimates room information RI is not limited to methods using artificial intelligence such as neural networks. For example, the sound signal processing device 1 may estimate room information RI by pattern matching. In this case, the sound signal processing device 1 has pre-recorded template data, such as an image showing a closed space or an image showing an open space. The estimation unit 172 calculates the similarity between the first image M1 and the template data, and estimates room information RI based on the similarity.

[0030] After step S12, the setting unit 173 sets acoustic parameters SP according to the estimated room information RI (Figure 3: step S13). In this embodiment, the acoustic parameters SP are parameters related to AGC or noise reduction. In this embodiment, the setting unit 173 sets acoustic parameters SP suitable for a closed space, or sets acoustic parameters SP suitable for an open space. For example, if the estimation unit 172 estimates "room information RI: closed space", the setting unit 173 sets parameters to turn on AGC and parameters to turn off noise reduction as acoustic parameters SP (see Figure 6). That is, if the estimation unit 172 estimates "room information RI: closed space", the setting unit 173 turns on AGC and turns off noise reduction. On the other hand, if the estimation unit 172 estimates "room information RI: open space", the setting unit 173 turns off AGC and turns on noise reduction (see Figure 6). As described above, in this embodiment, the setting unit 173 sets the acoustic parameter SP based on information indicating whether it is an open space or a closed space.

[0031] The noise reduction in this embodiment is, for example, multi-channel signal processing that outputs a single output signal from the output signals of multiple microphones. In this case, microphone 4 is a microphone array having multiple microphones.

[0032] Note that noise reduction is not limited to the examples shown above. For example, noise reduction may be a noise gate that calculates the signal level of microphone 4 and attenuates the signal level of microphone 4 only when the signal level is below a certain level. Alternatively, noise reduction may be a process that calculates the average power of microphone 4 over a predetermined period (long time) for each frequency and removes noise by filtering, such as a Wiener filter.

[0033] Next, the reception unit 170 receives the sound signal (Figure 3: Step S14). As shown in Figure 2, the reception unit 170 acquires the sound signal SS1 from the microphone 4.

[0034] Next, the signal processing unit 174 performs sound processing on the sound signal SS1 based on the acoustic parameter SP (Figure 3: step S15). For example, if AGC is on, the setting unit 173 automatically increases or decreases the gain of the sound signal SS1 so that the speaker's voice level remains constant (gain adjustment). In other words, in this embodiment, sound processing P includes gain adjustment. On the other hand, if AGC is off in the setting unit 173, the signal processing unit 174 does not perform AGC on the sound signal SS1. Also, if noise reduction is on in the setting unit 173, the setting unit 173 suppresses noise in the sound signal SS1. In other words, in this embodiment, sound processing P includes noise reduction. On the other hand, if noise reduction is off in the setting unit 173, the signal processing unit 174 does not perform noise reduction on the sound signal SS1. Hereinafter, the sound signal after sound processing will be referred to as sound signal SS2.

[0035] Next, the output unit 175 outputs the sound signal SS2 (Figure 3: step S16). Specifically, the output unit 175 outputs the sound signal SS2 to the communication interface 13. The communication interface 13 transmits the sound signal SS2 to the processing unit 2 via the communication line 3. The speaker of the processing unit 2 emits sound based on the sound signal SS2.

[0036] After step S16, the processor 17 determines, for example, whether or not there is a termination instruction for the application program related to sound processing P (Figure 3: step S17). If the processor 17 determines that there is "no termination instruction" (Figure 3: step S17 No), it repeats the processing from step S14 to step S16. This allows the processor 17 to repeatedly perform sound processing based on the acoustic parameter SP that was set initially.

[0037] In step S17, if the processor 17 determines that "termination instruction: present" (Figure 3: Step S17 Yes), it completes the execution of the series of sound processing P (Figure 3: END). The processor 17 may also determine whether or not to complete the execution of sound processing P by a method other than determining whether or not there is a termination instruction in the application program related to sound processing P.

[0038] Note that the processing order shown in Figure 3 is just one example, and the processor 17 does not necessarily have to execute the processes in the order shown in Figure 3. The processor 17 may execute the processes in any order, as long as it has executed the processes in step S13 and step S14 before executing step S15. For example, the processor 17 may perform the processes from step S11 to step S13 (setting the acoustic parameter SP) and the process in step S14 (receiving the sound signal SS1) in parallel.

[0039] (Effects of the first embodiment) The sound signal processing device 1 can perform appropriate sound processing depending on the situation. Specifically, the sound signal processing device 1 automatically estimates the type of space in which the user U is located (whether it is a closed space such as a conference room or an open space). Then, the sound signal processing device 1 automatically sets the acoustic parameters SP based on the estimation result. For example, if the sound signal processing device 1 estimates "Room information RI: Closed space", it automatically turns on AGC and automatically turns off noise reduction. By turning on AGC, the sound signal processing device 1 makes the voice of a speaker located far from the microphone 4 and the voice of a speaker located close to the microphone 4 at a constant level. Also, by turning off noise reduction, the sound signal processing device 1 does not remove the voice of user U, who is located far from the microphone 4, as noise. Therefore, the sound signal processing device 1 automatically sets the acoustic parameters SP to be suitable for a closed space where a speaker may be located far from the microphone 4.

[0040] The sound signal processing device 1 removes sounds that are far from the microphone 4 (for example, steady-state noise or the voice of a person far from the microphone 4) by turning on noise reduction. Furthermore, the sound signal processing device 1 prevents the volume of noise far from the microphone 4 from increasing by turning off AGC (Automatic Gain Control). As a result, the sound signal processing device 1 automatically sets the acoustic parameter SP to be suitable for open spaces where the speaker is only located close to the microphone 4. As described above, the sound signal processing device 1 can perform sound processing appropriately depending on the situation (depending on the space in which the user U is located).

[0041] The sound signal processing device 1 automatically sets the acoustic parameters SP based on the type of space in which the user U is located. Therefore, the user U does not need to manually set the acoustic parameters SP. As a result, errors in setting the acoustic parameters SP by the user U do not occur. Consequently, the user U and the person they are talking to can converse based on sound that has been appropriately processed.

[0042] (Variation 1) The following describes the sound signal processing device 1a (not shown) according to Modification 1. The configuration of the sound signal processing device 1a is the same as the configuration of the sound signal processing device 1 shown in Figure 2. Instead of performing sound processing P on the sound signal received from the microphone 4, the sound signal processing device 1a performs sound processing P on the sound signal received from the processing device 2. For example, if the sound signal processing device 1a estimates "room information RI: closed space," it sets the acoustic parameter SP to decrease the gain of the sound signal received from the processing device 2 to match the low-noise environment. As a result, the sound signal processing device 1a outputs the voice of a distant conversationalist at an appropriate volume to match the listening environment. On the other hand, if the sound signal processing device 1a estimates "room information RI: open space," for example, it sets the acoustic parameter SP to increase the gain of the sound signal received from the processing device 2 to match the high-noise environment. In this case as well, the sound signal processing device 1a outputs the voice of a distant conversationalist at an appropriate volume to match the listening environment. The sound signal processing device 1a can perform appropriate sound processing on the sound signal output to the speaker 5 depending on the situation.

[0043] The sound signal processing device 1a may perform sound processing P on both the sound signal SS1 received from the microphone 4 and the sound signal received from the processing device 2.

[0044] (Modification 2) The sound signal processing device 1b according to Modification 2 will be described below with reference to the figures. Figure 7 is a block diagram showing the functional configuration of the processor 17b of the sound signal processing device 1b.

[0045] The processor 17b of the sound signal processing device 1b functionally includes a setting unit 173b instead of a setting unit 173 (see Figure 7). In addition to the processing of the setting unit 173, the setting unit 173b performs processing to set acoustic parameters SP based on the sound signal SS2 that has undergone sound processing P. For example, the setting unit 173b measures the signal level of noise (steady-state noise) contained in the sound signal SS2 after sound processing. If the setting unit 173b detects noise with a signal level above a predetermined threshold, it turns off AGC and turns on noise reduction. In this way, even if the estimation unit 172 mistakenly estimates an open space as a closed space, the setting unit 173b performs acoustic control (AGC off and noise reduction on) that is suitable for an open space. The interlocutor can converse with user U with improved sound quality thanks to the sound signal processing device 1b.

[0046] (Variation 3) The following describes the sound signal processing device 1c according to Modification 3 with reference to the figures. Figure 8 is a flowchart showing an example of setting the acoustic parameter SP in the sound signal processing device 1c. The configuration of the sound signal processing device 1c is the same as the configuration of the sound signal processing device 1 shown in Figure 2.

[0047] While the sound signal processing unit 1 performs image acquisition, room information RI estimation, and acoustic parameter SP setting once each, the sound signal processing unit 1c performs image acquisition, room information RI estimation, and acoustic parameter SP setting two or more times each. This will be explained in detail below.

[0048] After step S14, the sound signal processing device 1c acquires the nth image (referred to as the nth image Mn) from the camera 6 (Figure 8: step S21). Here, n is any number greater than or equal to 1, and acquiring the nth image Mn means that the processing after S14 is the nth time. In other words, the acquisition unit 171 acquires the second image M2 at a different timing than when the first image M1 was acquired.

[0049] After the acquisition unit 171 acquires the second image M2, the estimation unit 172 of the sound signal processing device 1c estimates the room information RI from the acquired second image M2 (Figure 8: Step S22). The method for estimating the room information RI in the sound signal processing device 1c is the same as the method for estimating the room information RI in the sound signal processing device 1.

[0050] After the estimation unit 172 estimates the room information RI based on the second image M2, the setting unit 173 of the sound signal processing device 1c changes the acoustic parameters SP based on the room information RI estimated from the second image M2 (Figure 8: step S23). In this case, the signal processing unit 174 of the sound signal processing device 1c performs sound processing on the sound signal SS1 based on the changed acoustic parameters SP (Figure 8: step S15), and the output unit 175 of the sound signal processing device 1c outputs the sound signal SS2, which has undergone sound processing based on the changed acoustic parameters SP, to the processing device 2 (Figure 8: step S16).

[0051] After step S16, the processor 17 of the sound signal processing device 1c executes step S17. If the processor 17 determines in step S17 that there is "no termination instruction" (Figure 8: step S17 No), it executes steps S14, S21, S22, S23, S15, and S16 again.

[0052] In step S17, if the processor 17 determines that "termination instruction: present" (Figure 8: Step S17 Yes), it completes the execution of the series of sound processing P (Figure 8: END).

[0053] (Effect of Modification 3) While the sound signal processing device 1 sets the acoustic parameter SP once after the application program related to sound processing P has started, the sound signal processing device 1c sets the acoustic parameter SP two or more times. Therefore, the sound signal processing device 1c can change the acoustic parameter SP in accordance with changes in the space in which the user U is located. For example, the user U may remove a room partition, etc. In this case, the space in which the user U is located changes from a closed space to an open space. At this time, the sound signal processing device 1c automatically changes the acoustic parameter SP. Therefore, the sound signal processing device 1c can perform sound processing with the acoustic parameter SP set appropriately according to the change in situation.

[0054] (Modification 4) The following describes the sound signal processing device 1d according to Modification 4 with reference to the figures. Figure 9 is a diagram showing the gain adjustment in the sound signal processing device 1d. The configuration of the sound signal processing device 1d is the same as the configuration of the sound signal processing device 1 shown in Figure 2.

[0055] The signal processing unit 174 of the sound signal processing device 1d gradually changes the acoustic parameter SP over a predetermined time Pt ​​when changing the acoustic parameter SP. In this modified example, the sound signal processing device 1d gradually changes from AGC off to AGC on over a predetermined time Pt. Specifically, when the sound signal processing device 1d turns on AGC, it determines the target value TV for the gain of the sound signal SS1. The sound signal processing device 1d sets the target value TV as the acoustic parameter SP. At this time, the target value TV may differ from the current value CD of the sound signal SS1. In this case, the gain value of the sound signal SS1 is gradually changed from the current value CD to the target value TV over a predetermined time Pt. In this modified example, for example, the flash memory 15 of the sound signal processing device 1d has the predetermined time Pt ​​recorded in advance.

[0056] In the example shown in Figure 9, the flash memory 15 records a predetermined time Pt ​​as 6 seconds. In this case, the sound signal processing device 1d gradually changes the gain value of the sound signal SS1 over the course of 6 seconds. For example, in Figure 9, the current gain value of the sound signal SS1 is 20 dB, and the target gain value TV of the sound signal SS1 is 5 dB. In this case, the sound signal processing device 1d changes the gain value of the sound signal SS1 from 20 dB to 5 dB over the course of 6 seconds. As a result, the person speaking can converse with the user U without feeling any discomfort from the sound output from the speaker of the processing device 2.

[0057] (Variation 5) The sound signal processing device 1e according to Modification 5 will be described below with reference to the figures. Figure 10 is a block diagram showing the functional configuration of the processor 17e of the sound signal processing device 1e.

[0058] The processor 17e in the sound signal processing device 1e performs reverberation removal or reverberation addition, which are sound processing methods different from AGC or noise reduction. Therefore, the acoustic parameter SP in this modified example is a parameter related to reverberation removal or a parameter related to reverberation addition. The processor 17e functionally includes a setting unit 173e instead of a setting unit 173 (see Figure 10). The setting unit 173e turns reverberation removal on / off or reverberation addition on / off. In other words, in this modified example, the sound processing P includes at least one of reverberation removal or reverberation addition.

[0059] More specifically, the setting unit 173e turns on reverberation removal when the estimation unit 172 estimates "room information RI: closed space". In this case, the sound signal processing unit 1e performs reverberation removal on the sound signal SS1 related to the sound acquired by the microphone 4. The sound signal processing unit 1e transmits the sound signal SS2 after reverberation removal to the processing unit 2. The interlocutor can converse with user U using the sound that has had reverberation removed by the sound signal processing unit 1e. Therefore, the interlocutor can hear only user U's direct sound, making it easier to hear user U's voice.

[0060] On the other hand, if the estimation unit 172 estimates "Room information RI: Open space", the setting unit 173e turns on reverberation addition. In this case, the sound signal processing unit 1e adds reverberation to the sound signal received from the processing unit 2. The speaker 5 emits sound based on the sound signal SS2 to which reverberation has been added. By adding reverberation to the sound signal, user U can have a conversation with an interlocutor that has a sense of presence (for example, as if user U were having a conversation with the interlocutor in a conference room). As described above, the sound signal processing unit 1e can appropriately perform reverberation addition or reverberation removal depending on the situation.

[0061] (Experimental variation 6) The sound signal processing device 1f according to Modification 6 will be described below with reference to the figures. Figure 11 is a block diagram showing the functional configuration of the processor 17f of the sound signal processing device 1f. In the sound signal processing device 1f, components that are the same as those in the sound signal processing device 1 are denoted by the same reference numerals and their explanation is omitted.

[0062] The processor 17f in the sound signal processing unit 1f functionally includes a signal processing unit 174f instead of a signal processing unit 174 (see Figure 11). The signal processing unit 174f removes noise from the sound signal SS1 using a trained model MM1 for noise reduction. The trained model MM1 has learned the process of converting a given input sound signal (hereinafter referred to as the first sound signal) into a sound signal from which noise has been removed (hereinafter referred to as the second sound signal). In other words, the trained model MM1 has learned the relationship between the first sound signal and the second sound signal obtained by removing noise from the first sound signal. The signal processing unit 174f performs sound processing using the trained model MM1. Specifically, the signal processing unit 174f performs sound processing to convert the sound signal SS1 into a sound signal SS3 obtained by removing noise from the sound signal SS1. The signal processing unit 174f transmits the sound signal SS3 to the processing unit 2 via the output unit 175.

[0063] Note that the sound signal processing device 1f does not necessarily have to include the trained model MM1. Other devices such as a server may include the trained model MM1. In this case, the sound signal processing device 1f removes noise from the sound signal SS1 by transmitting the sound signal SS1 to the other device that includes the trained model MM1.

[0064] (Example 7) The following describes the sound signal processing device 1g (not shown) according to Modification 7, with reference to Figures 4 and 5. The configuration of the sound signal processing device 1g is the same as the configuration of the sound signal processing device 1 shown in Figure 2. The sound signal processing device 1g sets acoustic parameters SP based on room information RII other than information indicating whether it is an open space or a closed space.

[0065] Room information RII specifically includes information that describes the room itself, or information that describes the room's usage. Information that describes the room itself includes, for example, the room's size, shape, or materials. Information that describes the room's usage includes, for example, the number of people in the room or the room's equipment (furniture, etc.). Room equipment includes, for example, the number of chairs in the room or the shape of the desk. In other words, in this modified example, room information RII includes at least one of the following: room size, room shape, materials, number of people, number of chairs, or shape of the desk.

[0066] The sound signal processing device 1g estimates the size, shape, or material of a room based on, for example, the first image M1 shown in Figure 4. For example, the sound signal processing device 1g estimates the size, shape, or material of a room using existing object recognition technology. The sound signal processing device 1g sets acoustic parameters SP to suit the size, shape, or material of the room.

[0067] For example, the sound signal processing device 1g sets the acoustic parameter SP to increase or decrease the gain value of the sound signal received from the processing device 2. Specifically, if the sound signal processing device 1g estimates that the room is large, it increases the gain of the sound signal received from the processing device 2. This increases the volume of the sound output from the speaker 5. Therefore, the user U can hear the sound output from the speaker 5 even if they are far away from it. On the other hand, if the sound signal processing device 1g estimates that the room is small, it decreases the gain value of the sound signal received from the processing device 2. This prevents the user U from experiencing discomfort due to loud noises.

[0068] The size, shape, and material of a room are all factors that affect sound reverberation. Therefore, the sound signal processing device 1g can, for example, turn reverberation on or off. Specifically, the sound signal processing device 1g estimates whether a room is prone to reverberation or not, based on its size, shape, or material. If the sound signal processing device 1g estimates that the room is prone to reverberation, it turns on reverberation. In this case, the sound signal processing device 1g adds reverberation to the sound signal received from the processing device 2. As a result, the speaker 5 outputs sound related to the sound signal with added reverberation. Consequently, the sound quality of the sound emitted from the speaker 5 is improved. On the other hand, if the sound signal processing device 1g estimates that the room is prone to reverberation, it turns off reverberation. In this case, the sound signal processing device 1g does not add reverberation to the sound signal received from the processing device 2. Consequently, the sound signal processing device 1g does not perform unnecessary processing. As described above, the sound signal processing device 1g can appropriately switch the reverberation addition on or off depending on the room.

[0069] Furthermore, for example, the sound signal processing device 1g turns reverberation removal on or off based on its estimation of whether the room is prone to reverberation or not. Specifically, if the sound signal processing device 1g estimates that the room is prone to reverberation, it turns on reverberation removal. In this case, the sound signal processing device 1g obtains sound signal SS2 by performing a reverberation removal process on the sound signal SS1 received from the microphone 4. The sound signal processing device 1g transmits the reverberation-removed sound signal SS2 to the processing device 2. As a result, the speaker of the processing device 2 outputs the sound related to the reverberation-removed sound signal SS2. Therefore, the person speaking can easily hear the voice of user U. On the other hand, if the sound signal processing device 1g estimates that the room is not prone to reverberation, it turns off reverberation removal. In this case, the sound signal processing device 1g does not perform a reverberation removal process on the sound signal SS1 received from the microphone 4. Therefore, the sound signal processing device 1g does not perform unnecessary processing. As shown above, the sound signal processing device 1g can appropriately switch reverberation removal on or off depending on the room.

[0070] Furthermore, the sound signal processing device 1g estimates the number of people, the number of chairs, or the shape of the desk using existing object recognition technology, etc. For example, based on the first image M1 in Figure 4, the sound signal processing device 1g determines that "the number of people is 3 (people H1, H2, H3), the number of chairs is 2 (chairs C1, C2), and the shape of the desk (shape of desk E) is rectangular."

[0071] When there are many people in a room or many chairs in a room, the reverberation in the room tends to be weaker. Also, when the shape of the desks in the room is complex, the reverberation in the room tends to be weaker. Therefore, the sound signal processing device 1g estimates whether a room is prone to reverberation or not based on the number of people in the room, the number of chairs in the room, or the shape of the desks. Based on the estimation result of whether the room is prone to reverberation or not, the sound signal processing device 1g turns reverberation addition on / off or reverberation removal on / off.

[0072] For example, if the sound signal processing device 1g estimates that the room is prone to reverberation (e.g., there are few people, few chairs, or the desk has a simple shape), it turns off reverberation addition. In this case, the sound signal processing device 1g does not perform the reverberation addition process on the sound signal received from the processing device 2. Therefore, the sound signal processing device 1g does not perform unnecessary processing. Also, if the sound signal processing device 1g estimates that the room is prone to reverberation, it turns on reverberation removal. In this case, the sound signal processing device 1g obtains the sound signal SS2 by removing reverberation from the sound signal SS1 received from the microphone 4. The sound signal processing device 1g transmits the reverberation-removed sound signal SS2 to the processing device 2. Therefore, the person speaking can easily hear the voice of user U.

[0073] On the other hand, if the sound signal processing unit 1g estimates that the room is unlikely to generate reverberation (for example, if it estimates that there are many people, many chairs, or that the shape of the desk is complex), it turns on reverberation addition. In this case, the sound signal processing unit 1g processes the sound signal received from the processing unit 2 to add reverberation. As a result, the sound quality of the sound emitted from the speaker 5 is improved. Also, if the sound signal processing unit 1g estimates that the room is unlikely to generate reverberation, it turns off reverberation removal. In this case, the sound signal processing unit 1g does not perform reverberation removal on the sound signal SS1 received from the microphone 4. As a result, the sound signal processing unit 1g does not perform unnecessary processing.

[0074] As described above, in this modified example, the setting unit 173 of the sound signal processing device 1g sets acoustic parameters SP according to the size of the room, the shape and material of the room, the number of people, the number of chairs, or the shape of the desk. Therefore, the sound signal processing device 1g performs sound processing based on the acoustic parameters SP that are appropriately set according to the situation.

[0075] The room information RII may include information other than the size of the room, the shape of the room, the materials, the number of people, the number of chairs, or the shape of the desk. For example, the room information RII may include the number of people in the room who are facing the camera 6 and the number of people who are not facing the camera 6. The sound signal processing device 1g determines the number of people facing the camera 6 and the number of people not facing the camera 6, for example, based on artificial intelligence. In the example shown in Figure 5, the sound signal processing device 1g determines that "the number of people facing the camera 6 = 3 people (people H1, H2, H3)" and "the number of people not facing the camera 6 = 1 person (person Q1)". If the sound signal processing device 1g determines that the number of people facing the camera 6 is greater than the number of people not facing the camera 6, it determines that the space where user U is located is a closed space. On the other hand, if the sound signal processing device 1g determines that the number of people facing the camera 6 is less than the number of people not facing the camera 6, it determines that the space where user U is located is an open space.

[0076] The room information RII may include, for example, the prices of furniture placed in the space. The sound signal processing device 1g sets acoustic parameters SP based on, for example, the prices of the furniture. In this case, the sound signal processing device 1g estimates the prices of the furniture captured in the first image M1 using, for example, artificial intelligence. If the sound signal processing device 1g estimates that the furniture is expensive, it sets the acoustic parameters SP so that the speaker 5 does not generate a volume above a certain level. As shown above, the sound signal processing device 1g estimates, for example, whether or not it is acceptable to generate loud noises in a space based on the prices of the furniture. In other words, it can set acoustic parameters SP that are appropriate for the room.

[0077] (Variation 8) The following describes the sound signal processing device 1h according to the modified example 8 with reference to the figures. Figure 12 is a block diagram showing the functional configuration of the processor 17h of the sound signal processing device 1h. Figure 13 is a flowchart showing an example of setting the acoustic parameter SP in the sound signal processing device 1h. Figure 14 is a diagram showing an example of image processing in the sound signal processing device 1h.

[0078] The sound signal processing device 1h differs from the sound signal processing device 1 in that it performs a process to determine whether or not to output sound reflected from the top surface of the desk E.

[0079] As shown in Figure 12, the sound signal processing device 1h functionally includes a direction detection unit 176 in addition to a reception unit 170, an acquisition unit 171, an estimation unit 172, a setting unit 173, a signal processing unit 174, and an output unit 175. The direction detection unit 176 detects the direction F1 from which the sound is coming (Figure 13: step S30). For example, in this modified example, the sound signal processing device 1h is connected to a plurality of microphones (for example, microphone 4 and microphone 4a in Figure 12). The direction detection unit 176 detects the direction F1 by calculating the cross-correlation of the sound signals from the plurality of microphones (for example, the sound signal SS1 acquired from microphone 4 and the sound signal SS1a acquired from microphone 4a in Figure 12).

[0080] After step S30, the estimation unit 172 analyzes the first image M1 (for example, by artificial intelligence analysis similar to that in the first embodiment) to determine whether or not a human head is captured in the first image M1 (Figure 13: step S31).

[0081] If the estimation unit 172 determines that "human head: present" (Figure 13: step S31 Yes), it calculates the direction F2 of the detected human head (Figure 13: step S32). For example, in Figure 14, the estimation unit 172 estimates the direction F2 of the person H3 based on the first image M1.

[0082] After step S32, the estimation unit 172 determines whether or not a desk is captured in the first image M1 (Figure 13: step S33). Specifically, the estimation unit 172 performs a process to determine the presence or absence of a desk, which will be described later. In this case, the estimation unit 172 calculates the position of the desk based on the first image M1. The position of the desk is an example of information indicating the usage status of the room (information indicating the equipment in the room). Therefore, in this modified example, the room information RI includes information indicating the position of the desk.

[0083] If the estimation unit 172 determines that "desk: present" (Figure 13: step S33 Yes), it calculates the direction F3 of the desk (Figure 13: step S34). For example, in Figure 14, a desk E is captured in the first image M1. In this case, the estimation unit 172 calculates the direction F3 in which the desk E is located.

[0084] After step S34, the estimation unit 172 determines whether the direction F1 from which the sound is coming coincides with the direction F2 in which the person's head is located (Figure 13: step S35). For example, in Figure 14, the voice SH2 of person H3 reaches the microphone connected to the sound signal processing device 1h directly. In this case, the estimation unit 172 determines that the direction F1 coincides with the direction F2 in which person H3's head is located.

[0085] If the estimation unit 172 determines that "direction F1 matches direction F2" (Figure 13: step S35 Yes), the setting unit 173 sets the sound from direction F3 not to be output (Figure 13: step S36). This prevents the sound signal processing device 1h from receiving multiple delayed recordings of the user U's voice due to the sound SH3 reflected off the desk E, which would otherwise sound like an echo.

[0086] After step S36, the setting unit 173 forms a sound-collecting beam with high sensitivity in direction F1 (Figure 13: step S37). Specifically, the sound-collecting beam with high sensitivity in direction F1 is formed by delaying and combining the sound-collecting signals from multiple microphones connected to the sound signal processing device 1h by a predetermined delay amount. As a result, the sound signal processing device 1h can clearly acquire the voice SH2 of person H3. As shown above, in this modified example, the setting unit 173 sets acoustic parameters SP according to information indicating the position of the desk (an example of room information).

[0087] If the estimation unit 172 determines "No human head" in step S31 (Figure 13: Step S31 No), the direction detection unit 176 may have detected sound SH1 (such as the voice of a person not captured in the first image M1) coming from an area not captured in the first image M1, or sound from a sound source other than a human voice (for example, the sound of a PC as shown in Figure 14) (see Figure 14). For this reason, if the estimation unit 172 determines "No human head" (Figure 13: Step S31 No), the setting unit 173 sets the system to form a sound-collecting beam with high sensitivity in direction F1 (Figure 13: Step S40). As a result, the sound signal processing device 1h can clearly acquire sound SH1 (such as the voice of a person not captured in the first image M1) coming from an area not captured in the first image M1.

[0088] In step S33, if the estimation unit 172 determines that "desk: none" (Figure 13: step S33 No), it determines whether "the direction F1 from which the sound is coming coincides with the direction F2 in which the person's head is located" (Figure 13: step S38).

[0089] If the estimation unit 172 determines in step S38 that "direction F1 coincides with direction F2" (Figure 13: Step S38 Yes), the setting unit 173 forms a sound-collecting beam with high sensitivity to direction F1 (Figure 13: Step S40). As a result, the sound signal processing device 1h can clearly acquire the sound SH2 that arrived directly from the person, rather than the sound SH3 reflected from the top surface of the desk.

[0090] In step S38, the setting unit 173 terminates processing if the estimation unit 172 determines that "direction F1 does not coincide with direction F2" (Figure 13: Step S38 No). In other words, the sound signal processing device 1h maintains the current state of the sound-collecting beam. If direction F1 does not coincide with direction F2, the sound-collecting beam is pointed in the direction of an area not captured in the first image M1. Therefore, the setting unit 173 maintains the setting of the sound-collecting beam and acquires the sound SH1 (sound of a person not captured in the first image M1) that came from an area not captured in the first image M1.

[0091] In step S35, if the estimation unit 172 determines that "direction F1 does not coincide with direction F2" (Figure 13: step S38 No), it determines whether "direction F1 coincides with direction F3" (Figure 13: step S39).

[0092] In step S39, if the estimation unit 172 determines that "direction F1 coincides with direction F3," it is possible that the speaker's voice is being reflected off the desk E and picked up by the microphone. However, if the room is viewed in plan view, it is also possible that the speaker is located in the same direction as the reflected voice from the desk, and that the microphone is picking up direct sound from that speaker. In this case, if the sound signal processing unit 1h were to perform a process that does not output sound from that direction, it would stop outputting the voice of the speaker located in that direction. As a result, there is a risk that the interlocutor may not be able to hear the speaker's voice. Therefore, if the estimation unit 172 determines that "direction F1 coincides with direction F3" (Figure 13: step S39 Yes), the setting unit 173 sets the system to form a sound-collecting beam with high sensitivity in direction F1 (Figure 13: step S37). This allows the sound signal processing unit 1h to clearly acquire the speaker's voice reflected off the desk E.

[0093] On the other hand, if the direction detection unit 176 determines in step S39 that "direction F1 does not coincide with direction F3" (Figure 13: step S39 No), it terminates processing (Figure 13: END). In other words, the sound signal processing device 1h maintains the current state of the sound-collecting beam. If direction F1 does not coincide with direction F2, and direction F1 does not coincide with direction F3, then the sound-collecting beam is pointed in the direction of an area not captured in the first image M1. Therefore, the setting unit 173 maintains the setting of the sound-collecting beam and acquires the sound SH1 (sound of a person not captured in the first image M1) that came from an area not captured in the first image M1.

[0094] (effect) According to the sound signal processing device 1h, the interlocutor will be able to hear the voice of user U more clearly. The sound signal processing device 1h sets a delay amount (acoustic parameter SP) so as not to pick up sound reflected from the desk. For example, in Figure 14, the sound signal processing device 1h will be less likely to output the voice of person H3 reflected from the desk E. In this case, the sound signal processing device 1h prevents the voice of user U from being picked up multiple times with a delay due to the sound reflected from the desk E, which would sound like an echo. Therefore, the interlocutor will be able to hear the voice of person H3 more clearly.

[0095] (Process to determine whether a desk is present or not) The following describes the process by which the sound signal processing device 1h determines the presence or absence of a desk (hereinafter referred to as process Z). The sound signal processing device 1h determines the presence or absence of a desk by analyzing the color distribution of the first image M1. Specifically, the sound signal processing device 1h divides the first image M1 into multiple regions (for example, 100 x 100 pixels, etc.) as shown by the dashed lines in Figure 14. The sound signal processing device 1h then sequentially applies the following processes (1) to (9) to each divided region.

[0096] (1): Calculate the average RGB value of all pixels in each region (hereinafter referred to as the first average value).

[0097] (2) Calculate the number of regions (hereinafter referred to as the first region) in the first row (the bottom row) of multiple rows in which the RGB values ​​fall within the range where they can be considered the same color. The range where they can be considered the same color is, for example, the range of the median of the first mean values ​​in that row ± α (where α is an arbitrary value). In other words, each region is considered the first region if the range is median - α < first mean value < median + α.

[0098] (3) If the ratio of the number of first regions to the total number of regions in the first row is equal to or greater than the first threshold (for example, 80% or more), it is determined that desk E has been imaged in the first row. If the ratio of the number of first regions to the total number of regions is less than the first threshold, it is determined that desk E has not been imaged.

[0099] (4): If it is determined in (3) that desk E has not been imaged, the process from (2) to (3) is repeated in the row following the row in which the determination was made. For example, if it is determined in the first row that desk E has not been imaged, the process from (2) to (3) is repeated in the second row.

[0100] (5): If it is determined in (3) that desk E has been imaged, the average value of all RGB values ​​in the first region (hereinafter referred to as the second average value) is calculated.

[0101] (6) In the row following the determination that desk E has been imaged, find the number of regions with a color similar to the second mean value (hereinafter referred to as the second region). A similar color is, for example, a color within the range of second mean value ± Δ (where Δ is any value). In other words, each region is considered a second region if it is within the range of second mean value - Δ < first mean value < second mean value + Δ.

[0102] (7) If the ratio of the number of second regions to the total number of regions in that row is equal to or greater than the second threshold (for example, 60% or more), it is determined that desk E has been imaged in that row. The second threshold is less than the first threshold.

[0103] (8): Repeat steps (5) through (7) for the remaining rows.

[0104] (9): If it is determined in (8) that desk E is not imaged in that row, the process of determining the presence or absence of desk E is terminated. As a result, the sound signal processing device 1h determines the range of desk E imaged in the first image M1 (the area in which desk E is imaged).

[0105] (effect) The sound signal processing unit 1h, which performs process Z, determines the presence or absence of desk E on a region-by-region basis, rather than on a pixel-by-pixel basis. In this case, the load on the sound signal processing unit 1h is smaller compared to the case where the presence or absence of desk E is determined on a pixel-by-pixel basis.

[0106] The color of desk E is often the same. In other words, the color of desk E captured in the previous row is likely to be the same as the color of desk E captured in the next row. Therefore, the sound signal processing device 1h, which executes process Z, reflects the calculation result of the previous row (the calculation result of the average RGB value of the first region that was determined to be desk E in the previous row) in the calculation of the next row (whether or not it is the second region). In other words, it determines whether or not desk E is captured in each region of the next row based on the color of desk E identified in the previous row (the average RGB value of the first region that was determined to be desk E in the previous row) (the presence or absence of desk E is determined by whether or not the colors are similar). Consequently, the detection accuracy of desk E in the sound signal processing device 1h is improved.

[0107] Objects appear smaller the further away they are from the image. Therefore, a rectangular desk E is captured in a trapezoidal shape. In the first image M1, the width of the desk E captured in the upper row is smaller than the width of the desk E captured in the lower row. Consequently, the number of areas in which desk E is captured decreases as you move up the row. The sound signal processing device 1h then sets a second threshold that is less than the first threshold (a threshold corresponding to the trapezoidal shape of the captured desk E) and determines whether or not a desk is captured in each row. This improves the detection accuracy of desk E in the sound signal processing device 1h.

[0108] Furthermore, the sound signal processing device 1h may change the processing of the sound beam for each region in which it determines that a desk E exists. For example, sound is more likely to reflect from the center of the desk E than from the edges of the desk E. Therefore, for each region in which it determines that a desk E exists, the sound signal processing device 1h determines whether "the center of the desk E exists or the edges of the desk E exist." For each region in which it determines that "the center of the desk exists," the sound signal processing device 1h performs processing (steps S34, S35, S36, S37, S39) based on the flow shown in Figure 13. On the other hand, the sound signal processing device 1h does not perform sound beam processing for each region in which it determines that "the edges of the desk exist." In this way, the sound signal processing device 1h can appropriately process the sound beam for each region in which it determines that a desk E exists.

[0109] The sound signal processing device 1h may calculate the reflection angle of sound for each region where it is determined that a desk E exists (for example, by analyzing the first image M1), and perform sound beam processing based on the calculated reflection angle. For example, the reflection angle of a standing speaker's voice will be small. Microphones have difficulty picking up sound with a small reflection angle (sound from a direction that does not have directionality). Therefore, the sound signal processing device 1h will not output sound with a small reflection angle (sound that is difficult to pick up). This prevents the interlocutor from feeling that the sound is difficult to hear. On the other hand, the reflection angle of a seated speaker's voice will be large. In this case, the direction of the reflected sound and the direction of the sound that arrived directly from the speaker can be considered to be the same (direction F1 can be considered to be approximately equal to direction F3). For this reason, the sound signal processing device 1h will form a sound pickup beam to pick up sound with a large reflection angle.

[0110] Furthermore, the frequency characteristics of the sound picked up by the microphone may change depending on the direction from which the sound is coming. For example, the frequency characteristics may change due to interference between sound reflected from the desk E and sound that arrives directly from the speaker. Therefore, the sound signal processing device 1h may change the equalizer parameters based on the direction from which the sound is coming for each region in which it has determined that the desk E is present. This allows the sound signal processing device 1h to output sound that is easy for the speaker to hear.

[0111] The sound signal processing device 1h may also determine whether or not to output sound based on the distance between the microphone and the sound reflection position on the desk E (hereinafter referred to as the microphone-reflection position distance). For example, sound reflected at a position close to the microphone can be considered identical to sound that came directly from the speaker (F1 can be considered approximately equal to F3). Therefore, the sound signal processing device 1h calculates the microphone-reflection position distance for each region where it is determined that a desk E exists. If the sound signal processing device 1h determines that the microphone-reflection position distance is short (below an arbitrary threshold set in the sound signal processing device 1h beforehand), it does not perform sound beam processing for that region. This reduces the processing load on the sound signal processing device 1h compared to the case where sound beam processing is performed for all regions where it is determined that a desk E exists.

[0112] Furthermore, the configurations of the sound signal processing devices 1, 1a, 1b, 1c, 1d, 1e, 1f, 1g, and 1h may be combined in any way. [Explanation of Symbols]

[0113] 1,1a,1b,1c,1d,1e,1f,1g,1h: Audio signal processing device 17, 17b, 17e, 17f: Processors 170: Reception Department 171: Acquisition Department 172: Estimation part 173,173b,173e: Setting section 174,174f: Signal Processing Unit 175: Output section M1: First image P: Sound processing RI,RII: Room Information SP: Acoustic parameters SS1, SS2, SS3: Audio signals

Claims

1. Receiving an audio signal, The first image is obtained, Based on the acquired first image, room information is estimated. Set acoustic parameters according to the estimated room information. Sound processing based on the set acoustic parameters is performed on the sound signal. The sound signal that has undergone the aforementioned sound processing is output, The aforementioned room information includes information indicating whether it is an open space or a closed space. If the room information indicates that it is a closed space, set the acoustic parameters to turn on Auto Gain Control and turn off noise reduction. If the room information indicates that it is an open space, the acoustic parameters are set to turn off Auto Gain Control and turn on noise reduction. Audio signal processing method.

2. The acoustic parameters are changed based on the sound signal that has undergone the aforementioned sound processing. The sound signal processing method according to claim 1.

3. The second image is acquired at a different time than the first image was acquired. The room information is estimated from the acquired second image, The acoustic parameters are changed based on the room information estimated from the second image. The sound signal processing method according to claim 1 or claim 2.

4. In the above modification, the acoustic parameters are changed during a predetermined time period. The sound signal processing method according to claim 2 or claim 3.

5. The aforementioned room information includes at least one of the following: room size, room shape, material, number of people, number of chairs, or shape of desk. The acoustic parameters are set according to the size of the room, the shape of the room, the materials used, the number of people, the number of chairs, or the shape of the desk. The sound signal processing method according to any one of claims 1 to 4.

6. The aforementioned sound processing further includes at least one of reverberation removal or reverberation addition. The sound signal processing method according to any one of claims 1 to 5.

7. The aforementioned room information includes information indicating the location of the desk, The acoustic parameters are set according to the information indicating the position of the desk. The sound signal processing method according to any one of claims 1 to 6.

8. The sound processing is performed using a trained model that has been trained by machine learning to determine the relationship between a first sound signal and a second sound signal obtained by removing noise from the first sound signal. The sound signal processing method according to any one of claims 1 to 7.

9. The relationship between the input image and the room information is learned using a pre-trained model acquired through machine learning to estimate the room information. The sound signal processing method according to any one of claims 1 to 8.

10. A reception unit that receives sound signals, An acquisition unit that acquires the first image, An estimation unit that estimates room information based on the acquired first image, A setting unit that sets acoustic parameters according to the estimated room information, A signal processing unit that performs sound processing on the sound signal based on the set acoustic parameters, An output unit that outputs the sound signal that has undergone the aforementioned sound processing, Equipped with, The aforementioned room information includes information indicating whether it is an open space or a closed space. The setting unit sets the acoustic parameters to turn on Auto Gain Control and turn off noise reduction when the room information indicates that the room is a closed space, and sets the acoustic parameters to turn off Auto Gain Control and turn on noise reduction when the room information indicates that the room is an open space. Audio signal processing device.

11. The signal processing unit modifies the acoustic parameters based on the sound signal that has undergone sound processing. The sound signal processing device according to claim 10.

12. The acquisition unit acquires a second image at a different timing than the timing at which the first image was acquired. The estimation unit estimates the room information from the acquired second image, The setting unit modifies the acoustic parameters based on the room information estimated from the second image. The sound signal processing device according to claim 10 or claim 11.

13. The signal processing unit modifies the acoustic parameters during a predetermined time period in the modification. The sound signal processing device according to claim 11 or claim 12.

14. The aforementioned room information includes at least one of the following: room size, room shape, material, number of people, number of chairs, or shape of desk. The setting unit sets the acoustic parameters according to the size of the room, the shape of the room, the material, the number of people, the number of chairs, or the shape of the desk. The sound signal processing device according to any one of claims 10 to 13.

15. The aforementioned sound processing further includes at least one of reverberation removal or reverberation addition. The sound signal processing device according to any one of claims 10 to 14.

16. The aforementioned room information includes information indicating the location of the desk, The setting unit sets the acoustic parameters according to the information indicating the position of the desk. The sound signal processing device according to any one of claims 10 to 15.

17. The signal processing unit performs the sound processing using a trained model that has been machine-learned to determine the relationship between the first sound signal and the second sound signal obtained by removing noise from the first sound signal. The sound signal processing device according to any one of claims 10 to 16.

18. The estimation unit estimates the room information using a trained model that has learned the relationship between the input image and the room information through machine learning. The sound signal processing device according to any one of claims 10 to 17.

Citation Information

Patent Citations

  • Noise gate and sound collecting device

    JP2010122617A

  • Gain automatic setting device and gain automatic setting method

    JP2011151634A

  • Parameter prediction device and parameter prediction method for acoustic signal processing

    JP2018092117A

  • Room acoustics simulation using deep learning image analysis

    JP2022515266A

  • Multi-modal dereverbaration in far-field audio systems

    US20190028829A1