Sound processing apparatus, sound processing method, computer program product, and computer readable storage medium

By acquiring and utilizing the propagation time difference and attenuation factor of the sound signal in the sound processing device and performing correction processing, the echo problem caused by delayed input sound is solved, and effective removal of the sound signal inputted away from the microphone is achieved.

CN120151706APending Publication Date: 2025-06-13CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411800980.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-09
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the case of using the first microphone and the second microphone, the sound input to the second microphone may be input to the first microphone in a delay, resulting in the inability to remove the sound input to the second microphone from the sound signal acquired through the first microphone, thereby resulting in the echoed sound signal.

Method used

By setting a time difference acquisition unit and an attenuation factor acquisition unit in the sound processing device, the time difference and attenuation factor between the sound signal propagation in wireless communication and air are obtained, and the delayed and attenuated sound signal is removed using this information.

Benefits of technology

The sound input to the second microphone far from the first microphone is effectively removed, thereby reducing the state in which the sound of the subject is echoed and improving the quality of the sound signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151706A_ABST
    Figure CN120151706A_ABST
Patent Text Reader

Abstract

The invention relates to a sound processing apparatus, a sound processing method, a computer program product, and a computer readable storage medium. The sound processing apparatus acquires a first sound signal collected at a first location, acquires a second sound signal collected at a second location, transmitted from the second location by wireless communication, and received at the first location, acquiring a difference between a propagation time of a sound signal from the second position to the first position in wireless communication of the sound signal and a propagation time of a sound from the second position to the first position in airborne propagation of the sound, an attenuation factor of a sound in airborne propagation of the sound from the second position to the first position is acquired, and a third sound signal acquired by attenuating the second sound signal based on the attenuation factor and delaying the second sound signal based on the difference is removed from the first sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound processing device, a sound processing method, a computer program product, and a computer-readable storage medium. Background Art

[0002] The following technology has been proposed: when a sound other than a desired sound is input to a microphone (hereinafter referred to as "mic"), the sound signal is regarded as a noise signal and removed (eliminated) to obtain only the desired sound signal. Japanese Patent Application Laid-Open No. 2012-100235 discloses the following technology: a front mic and a rear mic are arranged in a camera, and the sound signal obtained by the rear mic is regarded as noise and removed from the sound signal obtained by the front mic.

[0003] When using a first mic and a second mic, the sound input to the second mic may be input to the first mic with a delay. In this case, even if the technology disclosed in Japanese Patent Application Laid-Open No. 2012-100235 is used, the sound signal of the sound input to the second mic (the sound input to the first mic with a delay) cannot be removed from the sound signal obtained by the first mic. Therefore, if the sound signal obtained by the second mic is acquired through wireless communication and the sound signal is synthesized with the sound signal obtained by the first mic, a sound signal in which the sound input to the second mic is echoed is acquired. Summary of the Invention

[0004] The present invention provides a technology for removing the sound signal of the sound input to a second mic remote from the first mic (the sound input to the first mic with a delay) from the sound signal obtained by the first mic.

[0005] In a first aspect, the present invention provides a sound processing device, including: a first sound acquisition unit configured to acquire a first sound signal collected at a first position; a second sound acquisition unit configured to acquire a second sound signal collected at a second position, transmitted from the second position through wireless communication, and received at the first position; a time difference acquisition unit configured to acquire a difference between a propagation time of the sound signal from the second position to the first position in the wireless communication of the sound signal and a propagation time of the sound from the second position to the first position in the air propagation of the sound; an attenuation factor acquisition unit configured to acquire an attenuation factor of the sound from the second position to the first position in the air propagation of the sound; and a correction unit configured to remove, from the first sound signal, a third sound signal obtained by attenuating the second sound signal based on the attenuation factor and delaying the second sound signal based on the difference.

[0006] The present invention provides, in a second aspect, a sound processing method, the sound processing method comprising: the step of acquiring a first sound signal collected at a first position; the step of acquiring a second sound signal collected at a second position, transmitted from the second position through wireless communication and received at the first position; the step of acquiring the difference between the propagation time of the sound signal from the second position to the first position in the wireless communication of the sound signal and the propagation time of the sound from the second position to the first position in the air propagation of the sound; the step of acquiring the attenuation factor of the sound from the second position to the first position in the air propagation of the sound; and the step of removing, from the first sound signal, a third sound signal obtained by attenuating the second sound signal based on the attenuation factor and delaying the second sound signal based on the difference.

[0007] The present invention provides, in a third aspect, a computer program product comprising a program for causing a computer to execute the respective steps of the above-described sound processing method. The present invention provides, in a fourth aspect, a computer-readable storage medium storing a program for causing a computer to execute the respective steps of the above-described sound processing method.

[0008] Further features of the present invention will become apparent from the following description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a block diagram depicting the configuration of a sound processing system according to Embodiment 1;

[0010] Figure 2 is a flowchart depicting the basic setup process of a slave wireless microphone unit;

[0011] Figure 3 is a flowchart depicting the basic setup process of a master wireless microphone unit;

[0012] Figure 4 is a schematic diagram of a scenario using a slave wireless microphone unit and a master wireless microphone unit;

[0013] Figure 5 is a flowchart depicting sound processing including reverberation correction according to Embodiment 1;

[0014] Figure 6 (A) to Figure 6 (D) of are diagrams showing waveforms of various sound signals;

[0015] Figure 7 is a block diagram depicting the configuration of a shooting system according to Embodiment 2;

[0016] Figure 8 is a flowchart depicting sound processing including reverberation correction according to Embodiment 2; and

[0017] Figure 9 from (A) to Figure 9 The (C) is a diagram showing waveforms of various sound signals. Detailed implementation mode

[0018] <Example 1>

[0019] Example 1 of the present invention will be described. Figure 1 is a block diagram depicting the configuration of the sound processing system according to Example 1. Figure 1 The sound processing system in includes a slave microphone unit (hereinafter referred to as "slave unit") 100 and a master wireless microphone unit (hereinafter referred to as "master unit") 200. For example, the master unit 200 can be used together with a imaging device (such as a digital camera). The scenario of using the sound processing system is not particularly limited, but in Example 1, it is assumed that the sound processing system is used in a shooting scenario, where the slave unit 100 is used at the position of the subject, and the master unit 200 is used at a position away from the subject (for example, at the position of the photographer).

[0020] The slave unit 100 includes a slave microphone unit 101, a slave microprocessing unit (MPU) 102, a slave wireless circuit 103, a slave sound circuit 104, and a speaker unit 108.

[0021] The slave microphone unit 101 includes a circuit for converting the input sound into an electrical signal (sound signal), and for acquiring the sound of the subject.

[0022] The slave MPU 102 controls the operation of the slave unit 100. The slave MPU 102 includes a ROM (not shown) storing a program for controlling the operation of the slave unit 100, a RAM (not shown) storing variables, and an EEPROM (electrically erasable programmable memory) (not shown) storing various parameters. The slave MPU 102 controls the operation of the slave unit 100 by expanding the program stored in the ROM in the RAM and executing the program. For example, the slave MPU 102 sends control signals to the slave wireless circuit 103 and the slave sound circuit 104 based on a predetermined communication system such as SPI communication or I2C communication, and thereby controls the operation of the slave wireless circuit 103 and the operation of the slave sound circuit 104.

[0023] The sound signal obtained from the slave microphone unit 101 is wirelessly transmitted from the wireless circuit 103 via the slave sound circuit 104 to the main wireless circuit 203 of the main unit 200. The transmission (wireless communication) system is, for example, Bluetooth (registered trademark) or Zigbee (registered trademark). The sound signal obtained from the slave microphone unit 101 can be interpreted as a sound signal collected at the position of the subject (the position of the slave unit 100). The transmission from the wireless circuit 103 to the main wireless circuit 203 can be interpreted as a transmission from the position of the subject to the position of the photographer (the position of the main unit 200).

[0024] The slave sound circuit 104 includes a sound acquisition unit 105, a sound level adjustment unit 106, and a sound signal generation unit 107. The sound acquisition unit 105 acquires the sound signal obtained from the slave microphone unit 101 from the slave microphone unit 101. The sound level adjustment unit 106 adjusts the sound level of the sound signal obtained through the sound acquisition unit 105 (the sound signal obtained from the slave microphone unit 101). For example, the sound level adjustment unit 106 adjusts the sound level using the gain set through a register. The sound signal generation unit 107 generates a sound signal output to the speaker unit 108.

[0025] The speaker unit 108 outputs a sound corresponding to the sound signal output from the slave sound circuit 104 (the sound signal generation unit 107).

[0026] The main unit 200 includes a main microphone unit 201, a main MPU 202, a main wireless circuit 203, and a main sound circuit 204.

[0027] The main microphone unit 201 includes a circuit for converting the input sound into an electrical signal (sound signal) and for acquiring ambient sound.

[0028] The main MPU 202 controls the operation of the main unit 200. The main MPU 202 includes a ROM (not shown) that stores a program for controlling the operation of the main unit 200, a RAM (not shown) that stores variables, and an EEPROM (electrically erasable programmable memory) that stores various parameters. The main MPU 202 controls the operation of the main unit 200 by expanding the program stored in the ROM in the RAM and executing the program. For example, the main MPU 202 sends control signals to the main wireless circuit 203 and the main sound circuit 204 based on a predetermined communication system such as SPI communication or I2C communication, and thereby controls the operation of the main wireless circuit 203 and the operation of the main sound circuit 204.

[0029] The main wireless circuit 203 receives the sound signal transmitted from the wireless circuit 103 of the slave unit 100. The communication system of the wireless circuit 103 is the same as the communication system of the main wireless circuit 203.

[0030] The main sound circuit 204 includes a main sound acquisition unit 205, a slave sound acquisition unit 206, a sound level adjustment unit 207, a sound time difference detection unit 208, a sound attenuation factor calculation unit 209, a sound correction unit 210, a sound synthesis unit 211, and a sound frequency detection unit 212.

[0031] The main sound acquisition unit 205 acquires the sound signal obtained through the main microphone unit 201 from the main microphone unit 201. The sound signal obtained through the main microphone unit 201 can be interpreted as the sound signal collected at the position of the photographer (the position of the main unit 200).

[0032] The slave sound acquisition unit 206 acquires the sound signal received through the main wireless circuit 203 from the main wireless circuit 203. Hereinafter, the sound signal acquired through the slave sound acquisition unit 206 (the sound signal acquired by the slave microphone unit 101, transmitted through wireless communication from the slave wireless circuit 103, and received by the main wireless circuit 203) is referred to as the "slave sound signal".

[0033] The sound level adjustment unit 207 adjusts the sound level of the sound signal acquired through the main sound acquisition unit 205 (the sound signal acquired through the main microphone unit 201). For example, the sound level adjustment unit 207 adjusts the sound level using the gain set through the register. Hereinafter, the sound signal acquired by the main sound acquisition unit 205 and adjusted by the sound level adjustment unit 207 (the sound signal acquired by the main microphone unit 201) is referred to as the "main sound signal".

[0034] The sound time difference detection unit 208 acquires the time difference between the main sound signal and the slave sound signal. The sound time difference detection unit 208 includes a circuit for detecting the time difference between the main sound signal and the slave sound signal. Hereinafter, this time difference is referred to as the "sound time difference".

[0035] The sound time difference is the difference between the propagation time of the sound signal from the position of the subject (the position of the slave unit 100) to the position of the photographer (the position of the main unit 200) in wireless communication and the propagation time of the sound from the position of the subject to the position of the photographer in the air propagation of the sound. Hereinafter, the propagation time in wireless communication is referred to as the "wireless propagation time", and the propagation time in air propagation is referred to as the "air propagation time". The wireless propagation time can be interpreted as the time from the transmission of the sound signal from the slave unit 100 (the slave wireless circuit 103) to the arrival of the slave sound signal at the main unit 200 (the main wireless circuit 203). The air propagation time can be interpreted as the time from the emission of the sound from the subject or the slave unit 100 (the speaker unit 108) to the arrival of the sound at the main unit 200 (the main microphone unit 201).

[0036] The sound attenuation factor calculation unit 209 obtains the attenuation factor of sound in air propagation from the position of the subject to the position of the photographer. The sound attenuation factor calculation unit 209 includes a circuit for calculating (detecting) the ratio between the sound level of the main sound signal and the sound level of the slave sound signal as the attenuation factor. Hereinafter, this attenuation factor will be referred to as the "sound attenuation factor".

[0037] The sound correction unit 210 corrects the main sound signal based on the sound time difference obtained by the sound time difference detection unit 208 and the sound attenuation factor obtained by the sound attenuation factor calculation unit 209. In the first embodiment, the sound correction unit 210 delays the slave sound signal based on the sound time difference, and attenuates the slave sound signal based on the sound attenuation factor. Then, the sound correction unit 210 removes (eliminates) the delayed and attenuated slave sound signal from the main sound signal.

[0038] The sound synthesis unit 211 synthesizes the main sound signal and the slave sound signal. When the sound correction unit 210 corrects the main sound signal, the sound synthesis unit 211 synthesizes the corrected main sound signal and the slave sound signal.

[0039] The sound frequency detection unit 212 detects the frequency of the main sound signal or detects the frequency of the slave sound signal. The sound frequency detection unit 212 can also preset (specify) a frequency band through a register, and detect whether the frequency of the sound signal is included in the set frequency band (predetermined frequency band).

[0040] Figure 2 is a flowchart depicting the basic setting process of the unit 100. For example, when the power of the unit 100 is turned on, the Figure 2 basic setting process starts.

[0041] In step S201, the slave MPU 102 initializes the variables and programs stored in the RAM, and performs preparation operations such as supplying power to the slave wireless circuit 103 and the slave sound circuit 104.

[0042] In step S202, the slave MPU 102 sets the communication parameters of the slave wireless circuit 103. The communication parameters of the slave wireless circuit 103 include, for example: MAC address, service set identifier (SSID), data channel, and transmission rate.

[0043] In step S203, sound processing parameters of the slave sound circuit 104 are set from the MPU 102. The sound processing parameters of the slave sound circuit 104 include, for example: gain for adjusting the sound level, frequency characteristics of an equalizer (for sound quality adjustment), a filter for preventing wind noise, and generation / non-generation of a sound signal to be output to the speaker unit 108. The sound processing parameters of the slave sound circuit 104 can be set by changing settings in the registers of the slave sound circuit 104.

[0044] By performing the processes of steps S201 to S203, the settings of the slave unit 100 are completed, and the sound signal acquired from the slave microphone unit 101 can be transmitted to the master unit 200.

[0045] Figure 3 is a flowchart depicting the basic setting process of the master unit 200. For example, when the power of the master unit 200 is turned on, the Figure 3 basic setting process starts.

[0046] In step S301, the main MPU 202 initializes variables and programs stored in the RAM, and performs preparatory operations such as supplying power to the main wireless circuit 203 and the main sound circuit 204.

[0047] In step S302, the main MPU 202 sets communication parameters of the main wireless circuit 203. Similar to the communication parameters of the slave wireless circuit 103, the communication parameters of the main wireless circuit 203 include, for example: MAC address, SSID, data channel, and transmission rate.

[0048] In step S303, the main MPU 202 sets sound processing parameters of the main sound circuit 204. The sound processing parameters of the main sound circuit 204 include: gain for adjusting the sound level, frequency characteristics of an equalizer (for sound quality adjustment), and a filter for preventing wind noise. The sound processing parameters of the main sound circuit 204 can be set by changing settings in the registers of the main sound circuit 204.

[0049] In step S304, the main MPU 202 determines whether the main wireless circuit 203 is connected to the slave wireless circuit 103. If it is determined that the main wireless circuit 203 is connected to the slave wireless circuit 103, the main MPU 202 advances the process to step S305, or if the main MPU 202 determines that the main wireless circuit 203 is not connected to the slave wireless circuit 103, the main MPU 202 advances the process to step S306. For example, if the power supply of the slave unit 100 is turned off, the main wireless circuit 203 is not connected to the slave wireless circuit 103, and the process advances to step S306. The main unit 200 can be used alone without wirelessly connecting the slave unit 100 to the main unit 200, and in this case, the echo correction described below is not required.

[0050] In step S305, the main MPU 202 instructs the main sound circuit 204 to perform sound processing including echo correction. This will be described in detail later with reference to Figure 5 the sound processing including echo correction.

[0051] In step S306, the main MPU 202 instructs the main sound circuit 204 to perform sound processing that does not include echo correction. In the sound processing that does not include echo correction, the sound signal input from the main microphone unit 201 is adjusted, for example, in terms of sound level, equalizer processing (sound quality adjustment), and filtering processing for preventing wind noise. For example, the sound processing that does not include echo correction is continuously performed until the power supply of the main unit 200 is turned off.

[0052] By performing the processing from step S301 to step S306, the setting of the main unit 200 is completed, and it becomes possible to acquire the slave sound signal and the main sound signal. If the main unit 200 is connected to the slave unit 100 after the processing in Figure 3 is completed, the start of the transmission of the sound signal is indicated to the slave unit 100 through the main wireless circuit 203. The main wireless circuit 203 sends a control signal for indicating the start of the transmission of the sound signal to the slave unit 100. The slave wireless circuit 103 receives the control signal for indicating the start of the transmission of the sound signal from the main wireless circuit 203 and sends the control signal to the slave MPU 102. In response to the instruction to start the transmission of the slave sound signal, the slave MPU 102 controls the slave sound circuit 104 and starts acquiring the slave sound signal. Then, the slave MPU 102 controls the slave wireless circuit 103 and transmits the slave sound signal to the main unit 200. Thereafter, if the slave unit 100 is connected to the main unit 200 while the power supply is on, the slave MPU 102 continues to transmit the slave sound signal to the main unit 200.

[0053] Figure 4 is a schematic diagram depicting a scenario of using the slave unit 100 and the main unit 200. In Figure 4In order to acquire the voice (speech) of the subject (e.g., the interviewee) 401, the slave unit 100 is arranged near the mouth of the subject. The master unit 200 is connected to the camera 400 held by the photographer 402 who is away from the subject 401. The master unit 200 receives the slave voice signal from the slave unit 100, synthesizes the slave voice signal with the master voice signal, and outputs the synthesized voice signal to the camera 400. The camera 400 records the voice signal from the master unit 200 together with the moving image captured by the camera 400. The storage medium for storing the synthesized voice signal generated by synthesizing the master voice signal and the slave voice signal is not limited to the camera.

[0054] The voice of the subject 401 input to the slave unit 100 may also be input to the master unit 200 in a state of attenuation and delay through air propagation. In this case, the timing of the voice of the subject 401 represented by the slave voice signal is offset from the timing of the voice of the subject 401 represented by the master voice signal. Therefore, when the slave voice signal is synthesized with the master voice signal, the voice of the subject 401 included in the synthesized voice signal is reverberated.

[0055] Therefore, in Embodiment 1, reverberation correction is performed. Reverberation correction is a process of removing the voice signal of the voice (the voice input to the master unit 200 in a state of attenuation and delay) input to the slave unit 100 from the master voice signal. By performing this reverberation correction, the state in which the voice of the subject 401 is reverberated can be reduced from the synthesized voice signal.

[0056] Figure 5 is a flowchart describing the voice processing including reverberation correction according to Embodiment 1. For example, when receiving the Figure 3 instruction in step S305 in, the master voice circuit 204 starts Figure 5 the voice processing in. For example, the voice processing in is repeated Figure 5 until the wireless connection between the slave unit 100 and the master unit 200 is cancelled, or until the power of the master unit 200 is turned off.

[0057] In step S501, the master voice circuit 204 performs normal voice processing. In the normal voice processing, like the voice processing not including reverberation correction, for example, the voice signal input from the master microphone unit 201 is subjected to voice level adjustment, equalizer processing (voice quality adjustment), and filtering processing for preventing wind noise. The voice level adjustment unit 207 is used to adjust the voice level.

[0058] In addition, in step S501, the main MPU 202 sends a control signal to the slave unit 100 via the main radio circuit 203 to instruct the output of a predetermined sound (a sound having a predetermined frequency) for detecting the sound time difference and the sound attenuation factor. Hereinafter, this predetermined sound is referred to as the "speaker sound". The slave radio circuit 103 receives the control signal and sends it to the slave MPU 102. The slave MPU 102 controls the slave sound circuit 104 according to the control signal to instruct the output of the speaker sound, and causes the speaker unit 108 to output the speaker sound.

[0059] In step S502, the main sound circuit 204 obtains the sound time difference using the sound time difference detection unit 208. In step S503, the main sound circuit 204 obtains the sound attenuation factor using the sound attenuation factor calculation unit 209.

[0060] The reference Figure 6 of (A) and Figure 6 of (B) will describe in detail the method for obtaining the sound time difference and the sound attenuation factor. Figure 6 of (A) and Figure 6 of (B) (and (C) of Figure 6 and (D) of Figure 6 mentioned later) are diagrams showing the waveforms of various sound signals. In Figure 6 from (A) to Figure 6 of (D), the abscissa represents time, and the ordinate represents the sound level.

[0061] In the first embodiment, it is assumed that the sound time difference and the sound attenuation factor are obtained based on the main sound signal and the slave sound signal respectively acquired by the main unit 200 and the slave unit 100 when the slave speaker unit 108 outputs a predetermined sound (for example, a sound having a predetermined frequency). It is assumed that the positions of the slave unit 100 and the main unit 200 in the case of obtaining the sound time difference and the sound attenuation factor are the same as the positions of the slave unit 100 and the main unit 200 in the case of obtaining the sound from the subject.

[0062] Figure 6 of (A) is the waveform of the slave sound signal acquired by the slave unit 100 when the speaker sound is output. Figure 6 of (B) is the waveform of the main sound signal acquired by the main unit 200 when the speaker sound is output.

[0063] First, the method for obtaining the sound time difference will be described.

[0064] Figure 6 The slave sound signal in (A) of Figure 6In (A), the waveform of the speaker sound in the sound signal appears after a delay corresponding to the wireless propagation time from the timing of the speaker sound emission. The wireless propagation time is uniquely determined by the wireless circuit 103 and the main wireless circuit 203, so information representing the wireless propagation time can be provided in advance. In Embodiment 1, it is assumed that information representing the wireless propagation time has been stored in the EEPROM of the main MPU 202 in advance. The wireless propagation time can be interpreted as the delay time from the timing when the sound is emitted from the speaker unit 108 or the subject to the timing when the waveform of the sound appears in the sound signal.

[0065] Figure 6 In (B), the main sound signal is a sound signal delayed by the amount of the above-mentioned air propagation time from the timing of the speaker sound emission. In Figure 6 In the main sound signal of (B), the waveform of the speaker sound appears after a delay corresponding to the air propagation time from the timing of the speaker sound emission. Since the wireless propagation time and the air propagation time are different, in Figure 6 In the main sound signal of (B), the waveform of the speaker sound appears at a timing different from Figure 6 the timing at which the waveform of the speaker sound appears in the slave sound signal of (A). The air propagation time can be interpreted as the delay time from the timing when the sound is emitted from the speaker unit 108 or the subject to the timing when the waveform of the sound appears in the main sound signal.

[0066] The sound frequency detection unit 212 detects the frequency of the speaker sound from the slave sound signal and the main sound signal respectively. Thus, the timing of the speaker sound in the slave sound signal and the timing of the speaker sound in the main sound signal are detected. The sound time difference detection unit 208 calculates the time Δt from the timing of the speaker sound in the slave sound signal to the timing of the speaker sound in the main sound signal as the sound time difference (the difference between the wireless propagation time and the air propagation time).

[0067] Next, a method for obtaining the sound attenuation factor will be described.

[0068] As described above, the sound frequency detection unit 212 detects the frequency of the speaker sound from the slave sound signal and the master sound signal, respectively. Thereby, the time period of the speaker sound in the slave sound signal and the time period of the speaker sound in the master sound signal are detected. Then, the sound attenuation factor calculation unit 209 calculates the ratio Vrx / Vtx of the sound level (amplitude Vrx) of the speaker sound in the master sound signal to the sound level (amplitude Vtx) of the speaker sound in the slave sound signal as the sound attenuation factor. The amplitude Vtx is a value determined by dividing the slave sound signal by the gain Gaintx set in the sound level adjustment unit 106, and the amplitude Vrx is a value determined by dividing the master sound signal by the gain Gainrx set in the sound level adjustment unit 207. By normalizing using the gains Gaintx and Gainrx, the ratio Vrx / Vtx that accurately represents the attenuation factor (attenuation amount) of the sound from the position of the subject to the position of the photographer during the air propagation of the sound can be obtained.

[0069] The method of obtaining the sound time difference and the sound attenuation factor is not limited to the above method, and the sound time difference and the sound attenuation factor can be obtained based on the master sound signal and the slave sound signal obtained when the subject emits his / her voice. In addition, the sound processing in Figure 5 is repeated here, but the acquisition of the sound time difference and the sound attenuation factor does not always have to be repeated. For example, the process of obtaining the sound time difference and the sound attenuation factor can be performed only once, and the obtained sound time difference and sound attenuation factor can be reused in step S505 described later. The main unit 200 can obtain the sound time difference and the sound attenuation factor from the outside.

[0070] Continue Figure 5 with the description in

[0071] In step S504, the main sound circuit 204 uses the sound frequency detection unit 212 to determine whether the slave sound signal is a sound signal in a predetermined frequency band. If it is determined that the slave sound signal is a sound signal in the predetermined frequency band, the main sound circuit 204 advances the process to step S505, or if it is determined that the slave sound signal is not a sound signal in the predetermined frequency band, the main sound circuit 204 advances the process to step S506. In the first embodiment, since the slave unit 100 is used to acquire the sound of the subject, the frequency band of human speech, that is, at least 120 Hz and not exceeding 300 Hz, is used as the predetermined frequency band.

[0072] The reverberation correction will be described in detail. Figure 6The voice signal in (C) is Figure 6 The slave voice signal in (A) and Figure 6 The synthesized voice signal of the main voice signal in (B). Figure 6 The slave voice signal in (A) and Figure 6 The timing of the speaker voice in the main voice signal in (B) is different. Therefore, in Figure 6 the synthesized voice signal in (C), the waveform of the speaker voice appears at two positions. Just like in the case of the speaker voice, in the case of the subject's voice, the waveform of the same voice of the subject also appears at two positions in the synthesized voice signal. To solve this problem, reverberation correction is performed. The reverberation correction is performed using the voice time difference Δt obtained in step S502 and the voice attenuation factor Vrx / Vtx obtained in step S503.

[0073] The voice correction unit 210 calculates the correction value C(t) according to the slave voice signal ftx(t), the voice time difference Δt, and the voice attenuation factor Vrx / Vtx using the following Expression 1. The variable t is the timing (time position). The slave voice signal ftx(t) can be interpreted as the voice level of the slave voice signal. The slave voice signal ftx(t - Δt) can be interpreted as the voice level of the signal generated by delaying the slave voice signal by the voice time difference Δt.

[0074] [Equation 1]

[0075]

[0076] Then, the voice correction unit 210 calculates the reverberation-corrected main voice signal frx2(t) according to the correction value C(t), the gain Gaintx, the gain Gainrx, and the main voice signal frx(t) using the following Expression 2 (correction formula). The main voice signal frx(t) before reverberation correction can be interpreted as the voice level of the main voice signal before reverberation correction. The main voice signal frx2(t) after reverberation correction can be interpreted as the voice level of the main voice signal after reverberation correction.

[0077] [Equation 2]

[0078]

[0079] For example, the echo correction of Expression 2 can be achieved by subtracting the correction value of the slave sound signal based on the time Δt before from the main sound signal using a dedicated hardware block composed of a timing circuit (such as a flip-flop). In the case where the main sound circuit 204 does not include a dedicated hardware block, the correction value based on the slave sound signal is stored in the RAM of the main MPU 202, and the correction value before the time Δt is read from the RAM and subtracted from the main sound signal. Thus, the echo correction in Expression 2 can be achieved.

[0080] Continue Figure 5 In the description of. In step S506, the main sound circuit 204 synthesizes the main sound signal and the slave sound signal using the sound synthesis unit 211. In the case where the echo correction in step S505 has been performed, the echo-corrected main sound signal frx2(t) and the slave sound signal ftx(t) are synthesized using the following Expression 3, and the synthesized sound signal f(t) is obtained. The synthesized sound signal f(t) can be interpreted as the sound level of the synthesized sound signal.

[0081] [Equation 3]

[0082]

[0083] Figure 6 The sound signal in (D) of Figure 6 The slave sound signal in (A) of Figure 6 And the synthesized sound signal of the sound signal generated by performing echo correction on the main sound signal in (B) of Figure 6 In the synthesized sound signal in (D) of Figure 6 The waveform of the speaker sound appears at one place, that is, the problem described in Figure 6 In (C) of

[0084] As described above, according to Embodiment 1, the sound time difference and the sound attenuation ratio are obtained, and the sound signal generated by delaying the slave sound signal based on the sound time difference and attenuating the slave sound signal based on the sound attenuation ratio is removed (eliminated) from the main sound signal. Thus, the sound signal of the sound input to the slave unit far from the main unit (the sound input to the main unit with a delay) can be removed from the sound signal obtained by the main unit. In addition, the acquisition of the synthesized sound signal in which the sound of the subject is echoed can be suppressed.

[0085] <Embodiment 2>

[0086] Embodiment 2 of the present invention will be described. Hereinafter, descriptions of aspects identical to those of Embodiment 1 (e.g., the same configurations and processes as those of Embodiment 1) will be omitted, and only aspects different from those of Embodiment 1 will be described.

[0087] Figure 7 is a block diagram depicting the configuration of the imaging system according to Embodiment 2. Figure 7 The imaging system in includes a slave unit 100, a master unit 200, and a camera 700. The slave unit 100 has the same configuration as that of Embodiment 1. The master unit 200 has substantially the same configuration as that of Embodiment 1. The main sound circuit 204 of the master unit 200 has a different configuration from that of Embodiment 1. Similar to Embodiment 1, the main sound circuit 204 includes a main sound acquisition unit 205, a slave sound acquisition unit 206, a sound level adjustment unit 207, a sound time difference detection unit 208, a sound attenuation factor calculation unit 209, a sound correction unit 210, a sound synthesis unit 211, and a sound frequency detection unit 212. The main sound circuit 204 further includes a distance information acquisition unit 213. The camera 700 includes a distance detection circuit 701, a camera MPU 702, and an imaging circuit 703. The master unit 200 and the camera 700 may be integrated into one unit.

[0088] The distance detection circuit 701 is a circuit that calculates (detects) the distance from the subject to the camera 700 using a known method such as the time-of-flight (TOF) method, and includes a sensor for calculating this distance.

[0089] The imaging circuit 703 includes an imaging element (image sensor). The imaging circuit 703 can calculate (detect) the distance from the subject to the camera 700 by a known method. For example, the imaging circuit 703 calculates the distance from the subject to the camera 700 based on the image acquired by the imaging element. When the camera 700 is turned toward the subject wearing the slave unit 100, the distance to the subject wearing the slave unit 100 can be acquired by the distance detection circuit 701 or the imaging circuit 703.

[0090] The camera MPU 702 controls the operation of the camera 700. The camera MPU 702 includes a ROM (not shown) that stores a program for controlling the operation of the camera 700, a RAM (not shown) that stores variables, and an EEPROM (electrically erasable programmable memory) (not shown) that stores various parameters. The camera MPU 702 controls the operation of the camera 700 by expanding the program stored in the ROM in the RAM and executing the program. For example, the camera MPU 702 sends control signals to the distance detection circuit 701 and the imaging circuit 703 based on a predetermined communication system such as SPI communication or I2C communication, and thereby controls the operations of the distance detection circuit 701 and the imaging circuit 703.

[0091] The camera MPU 702 can communicate with the main MPU 202 via an external I / F (not shown) arranged in the camera 700 and an external I / F (not shown) arranged in the main unit 200. For example, the camera MPU 702 uses a predetermined communication system (such as SPI communication, I2C communication) to send information representing the distance calculated by the distance detection circuit 701 or the imaging circuit 703 (the distance from the subject to the camera 700) to the main MPU 202.

[0092] The distance information acquisition unit 213 of the main unit 200 acquires shooting distance information representing the distance from the subject to the camera 700 from the camera MPU 702 via the main MPU 202. In Embodiment 2, it is assumed that the distance from the subject to the camera 700 is the same as the distance from the subject to the main unit 200. Therefore, the shooting distance information can be interpreted as information representing the distance from the subject to the main unit 200. The method for acquiring the shooting distance information is not particularly limited, and for example, the main unit 200 may include a sensor for acquiring the shooting distance information.

[0093] Figure 8 is a flowchart depicting sound processing including echo correction according to Embodiment 2. For example, when receiving an instruction in Figure 3 step S305 from the main MPU202, the main sound circuit 204 starts Figure 8 the sound processing in Figure 8 The sound processing in Figure 8 is repeated until the wireless connection between the slave unit 100 and the main unit 200 is cancelled, or until the power of the main unit 200 is turned off. The processing in

[0094] Step S801 and Figure 5is the same as step S501 in it. In step S802, the main sound circuit 204 determines whether there is shooting distance information. If it is determined that there is shooting distance information, the main sound circuit 204 performs the processing in steps S803 and S804, or if it is determined that there is no shooting distance information, the main sound circuit 204 performs the processing in steps S805 and S806. Steps S805 and S806 are the same as Figure 5 steps S502 and S503 in it.

[0095] In step S803, the main sound circuit 204 obtains the air propagation time based on the shooting distance information and obtains the sound time difference. Just like in step S805, in order to obtain the sound time difference, the sound time difference detection unit 208 is used.

[0096] The reference Figure 9 (A) to Figure 9 (C) of it will be described in detail the method for obtaining the sound time difference based on the shooting distance information. Figure 9 (A) to Figure 9 (C) of it are diagrams showing the waveforms of various sound signals. In Figure 9 (A) to Figure 9 (C) of it, the abscissa represents time, and the ordinate represents the sound level. Figure 9 The sound signal in (A) of it is the sound signal obtained from the microphone unit 101 and input to the slave sound circuit 104. Figure 9 The sound signal in (B) of it is the sound signal from, and Figure 9 The sound signal in (C) of it is the main sound signal.

[0097] From Figure 9 the timing of the speaker sound in the sound signal of (A) of it to Figure 9 the timing of the speaker sound in the sound signal of (B) of it, the time trf corresponds to the wireless propagation time. As described in Embodiment 1, the wireless propagation time is the propagation time of the sound signal from the position of the subject (the position of the slave unit 100) to the position of the photographer (the position of the master unit 200) in wireless communication. The wireless propagation time is uniquely determined by the wireless circuit 103 and the main wireless circuit 203, so the information representing the wireless propagation time can be provided in advance. In Embodiment 2, it is assumed that the information representing the wireless propagation time has been stored in the EEPROM of the main MPU 202 in advance.

[0098] From Figure 9 the timing of the speaker sound in the sound signal of (A) of it to Figure 9The time tdis of the timing of the speaker sound in the voice signal of (C) corresponds to the air propagation time. The air propagation time is the propagation time of sound in the air from the position of the subject (from the position of unit 100) to the position of the photographer (the position of the main unit 200). The voice time difference detection unit 208 calculates the time tdis by dividing the distance represented by the shooting distance information by the speed of sound. For example, if the distance from the main unit 200 to the subject is 34 m and the speed of sound is 340 m / sec, the time tdis is calculated as 0.1 sec.

[0099] The voice time difference detection unit 208 calculates the voice time difference Δt by subtracting the time trf from the time tdis using the following Expression 4. For example, if the distance from the main unit 200 to the subject is 34 m, the speed of sound is 340 m / sec, and the radio propagation time is 0.01 sec, the voice time difference Δt is calculated as 0.09 sec.

[0100] [Expression 4]

[0101] Δt = tdis - trf…(Expression 4)

[0102] Continue Figure 8 In the description of. In step S804, the main voice circuit 204 obtains the voice attenuation factor based on the shooting distance information. Just like in step S806, in order to obtain the voice attenuation factor, the voice attenuation factor calculation unit 209 is used.

[0103] Refer to Figure 4 A method for obtaining the voice attenuation factor based on the shooting distance information will be described in detail. In Figure 4 , the distance d1 from the subject to the slave unit 100 and the distance d2 from the subject to the main unit 200 are represented. In Embodiment 2, it is assumed that the information representing the distance d1 has been stored in the EEPROM of the main MPU 202 in advance. The distance d2 is the distance represented in the shooting distance information.

[0104] Since sound diffuses spherically, the voice attenuation factor calculation unit 209 calculates the voice attenuation factor Vatt in decibels based on the distances d1 and d2 using the following Expression 5. For example, if the distance d1 is 0.1 m and the distance d2 is 10 m, the voice attenuation factor Vatt is calculated as 40 dB.

[0105] [Expression 5]

[0106]

[0107] The voice attenuation factor calculation unit 209 converts the voice attenuation factor Vatt to the voice attenuation factor Vrx / Vtx using the following Expression 6.

[0108] [Number 6]

[0109]

[0110] Steps S807 to S809 are the same as Figure 5 steps S504 to S506 in

[0111] As described above, according to Embodiment 2, shooting distance information is obtained, and based on this shooting distance information, a sound time difference and a sound attenuation factor are obtained. Then, just like in Embodiment 1, a sound signal generated by delaying the slave sound signal based on the sound time difference and attenuating the slave sound signal based on the sound attenuation factor is removed (eliminated) from the main sound signal. Thus, an effect similar to that of Embodiment 1 is achieved. In addition, by using the shooting distance information, the sound time difference and the sound attenuation factor can be obtained without using the speaker sound, and a state where the subject or the photographer is surprised or uncomfortable due to the speaker sound can be prevented.

[0112] Note that the above various types of control can be processing executed by one piece of hardware (e.g., a processor or a circuit) or otherwise. The processing can be shared among multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) so as to execute the control of the entire device.

[0113] In addition, the above-mentioned processor is a general processor and includes a general-purpose processor and a dedicated processor. Examples of general-purpose processors include a central processing unit (CPU), a microprocessing unit (MPU), and a digital signal processor (DSP), etc. Examples of dedicated processors include a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), and a programmable logic device (PLD), etc. Examples of PLDs include a field-programmable gate array (FPGA) and a complex programmable logic device (CPLD), etc.

[0114] The above embodiments (including modification examples) are only examples. Any configuration obtained by appropriately modifying or changing some configurations of the embodiments within the scope of the subject matter of the present invention is also included in the present invention. The present invention also includes other configurations obtained by appropriately combining various features of the embodiments.

[0115] According to the present invention, a sound signal of sound input to a second microphone far from the first microphone (sound input to the first microphone with a delay) can be removed from the sound signal that can be obtained from the first microphone.

[0116] Other Embodiments

[0117] Embodiments of the present invention can also be implemented by the following method, that is, software (program) that executes the functions of the above embodiments is provided to a system or device through a network or various storage media, and the computer or central processing unit (CPU) or microprocessing unit (MPU) of the system or device reads and executes the program.

[0118] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims will be given the broadest interpretation so as to cover all such modifications and equivalent structures and functions.

Claims

1. A sound processing device, comprising: A first sound acquisition unit configured to acquire a first sound signal collected at a first position; a second sound acquisition unit configured to acquire a second sound signal collected at a second location, transmitted from the second location through wireless communication, and received at the first location; a time difference acquisition unit configured to acquire a difference between a propagation time of a sound signal from the second position to the first position in wireless communication of the sound signal and a propagation time of the sound from the second position to the first position in air propagation of the sound; an attenuation factor acquisition unit configured to acquire an attenuation factor of the sound from the second position to the first position in the air propagation of the sound; as well as A correction unit configured to remove, from the first sound signal, a third sound signal obtained by attenuating the second sound signal based on the attenuation factor and delaying the second sound signal based on the difference. 2 . The sound processing apparatus according to claim 1 , further comprising a synthesis unit configured to synthesize a fourth sound signal generated by removing the third sound signal from the first sound signal with the second sound signal.

3. The sound processing device according to claim 1 or 2, wherein: The attenuation factor acquisition unit acquires the attenuation factor, and the time difference acquisition unit acquires the difference, based on a first sound signal and a second sound signal acquired by emitting a predetermined sound at a second position.

4. The sound processing device according to claim 3, wherein: The time difference acquisition unit acquires a time from a timing of the predetermined sound in the second sound signal to a timing of the predetermined sound in the first sound signal as a difference between a propagation time in the wireless communication and a propagation time in the air propagation.

5. The sound processing device according to claim 3, wherein: The attenuation factor acquisition unit acquires, as the attenuation factor, a ratio of a level of the predetermined sound in the first sound signal to a level of the predetermined sound in the second sound signal.

6. The sound processing device according to claim 1 or 2, wherein: Information about propagation time in the wireless communication is provided in advance, and The time difference acquisition unit acquires the difference by acquiring the propagation time in the air propagation based on the distance from the first position to the second position.

7. The sound processing device according to claim 1 or 2, wherein: The attenuation factor acquisition unit acquires the attenuation factor based on a distance from the first position to the second position. 8 . The sound processing device according to claim 6 , further comprising an information acquisition unit configured to acquire information indicating a distance from the first position to the second position.

9. The sound processing device according to claim 1 or 2, wherein: The third sound signal is a sound signal in a predetermined frequency band.

10. A sound processing method, comprising: A step of acquiring a first sound signal collected at a first position; The step of acquiring a second sound signal collected at a second location, transmitted from the second location via wireless communication, and received at the first location; A step of obtaining a difference between a propagation time of a sound signal from the second position to the first position in wireless communication of the sound signal and a propagation time of the sound from the second position to the first position in air propagation of the sound; A step of obtaining an attenuation factor of sound from the second position to the first position in air propagation of the sound; as well as A step of removing, from the first sound signal, a third sound signal obtained by attenuating the second sound signal based on the attenuation factor and delaying the second sound signal based on the difference.

11. A computer program product comprising a program for causing a computer to execute the steps of the sound processing method according to claim 10.

12. A computer-readable storage medium storing a program for causing a computer to execute each step of the sound processing method according to claim 10.

Citation Information

Patent Citations

  • Sound recording method for imaging device

    JP2012100235A