Audio processing device

The audio processing apparatus addresses the issue of delayed audio signals between separated microphones by calculating and correcting for time differences and attenuation rates, effectively eliminating echoes in synthesized audio.

JP2025093141APending Publication Date: 2025-06-23CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023208695
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

When using separated first and second microphones, audio input to the second microphone can be delayed when received by the first microphone, making it impossible to remove the delayed audio signal from the first microphone's signal, resulting in an echo effect when synthesizing audio signals.

Method used

An audio processing apparatus that includes means to acquire first and second audio signals, calculate the time difference and attenuation rate between wireless and air propagation of audio signals, and correct the first audio signal by removing a delayed and attenuated version of the second audio signal.

Benefits of technology

Effectively removes the delayed audio signal from the first microphone, preventing echo formation in the synthesized audio, thereby improving audio quality in scenarios with separated microphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093141000001_ABST
    Figure 2025093141000001_ABST
Patent Text Reader

Abstract

To provide a technology capable of removing an audio signal of an audio input to a second microphone separated from a first microphone (audio input to the first microphone with a delay) from an audio signal obtained by the first microphone.SOLUTION: An audio processing device according to the present invention includes first audio acquisition means for acquiring a first audio signal collected at a first position, second audio acquisition means for acquiring a second audio signal collected at a second position, transmitted from the second position by wireless communication and received at the first position, time difference acquisition means for acquiring the difference between the propagation time of the wireless communication from the second position to the first position and the propagation time of air propagation from the second position to the first position, attenuation rate acquisition means for acquiring the attenuation rate of audio in air propagation, and correction means for removing from the first audio signal a third audio signal obtained by attenuating the second audio signal on the basis of the attenuation rate and delaying the second audio signal on the basis of the difference in propagation times.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing apparatus.

Background Art

[0002] When audio other than desired audio is input to a microphone (hereinafter referred to as a "mic"), a technique has been proposed to obtain an audio signal of only the desired audio by regarding the audio signal as a noise signal and removing (canceling) it. Patent Document 1 discloses a technique in which mics are provided in front of and behind a camera, and an audio signal obtained by the rear mic is regarded as noise and removed from the audio signal obtained by the front mic.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When using a first mic and a second mic that are separated from each other, the audio input to the second mic may be input to the first mic with a delay. In this case, even if the technique disclosed in Patent Document 1 is used, it is not possible to remove the audio signal of the audio input to the second mic (the audio input to the first mic with a delay) from the audio signal obtained by the first mic. Therefore, when the audio signal obtained by the second mic is acquired by wireless communication and synthesized with the audio signal obtained by the first mic, an audio signal such that the audio input to the second mic has an echo is obtained.

[0005] An object of the present invention is to provide a technique capable of removing an audio signal of audio input to a second microphone away from the first microphone (audio input with a delay to the first microphone) from an audio signal obtained by the first microphone.

Means for Solving the Problems

[0006] A first aspect of the present invention includes: a first audio acquisition means for acquiring a first audio signal collected at a first position; a second audio acquisition means for acquiring a second audio signal collected at a second position and transmitted from the second position by wireless communication and received at the first position; a time difference acquisition means for acquiring a difference between a propagation time of the audio signal in the wireless communication of the audio signal from the second position to the first position and a propagation time of the audio in the air propagation of the audio from the second position to the first position; an attenuation rate acquisition means for acquiring an attenuation rate of the audio in the air propagation of the audio from the second position to the first position; and a correction means for removing, from the first audio signal, a third audio signal obtained by attenuating the second audio signal based on the attenuation rate and delaying the second audio signal based on the difference. The audio processing apparatus is characterized by having the above.

[0007] A second aspect of the present invention includes: a step of acquiring a first audio signal collected at a first position; a step of acquiring a second audio signal collected at a second position and transmitted from the second position by wireless communication and received at the first position; a step of acquiring a difference between a propagation time of the audio signal in the wireless communication of the audio signal from the second position to the first position and a propagation time of the audio in the air propagation of the audio from the second position to the first position; a step of acquiring an attenuation rate of the audio in the air propagation of the audio from the second position to the first position; and a step of removing, from the first audio signal, a third audio signal obtained by attenuating the second audio signal based on the attenuation rate and delaying the second audio signal based on the difference. The audio processing method is characterized by having the above steps.

[0008] The third aspect of the present invention is a program for causing a computer to function as each means of the above-described voice processing apparatus.

Advantages of the Invention

[0009] According to the present invention, it is possible to remove the voice signal of the voice input to the second microphone (the voice input with a delay to the first microphone) from the voice signal obtained by the first microphone, which is away from the first microphone.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0011] <Example 1> Example 1 of the present invention will be described. FIG. 1 is a block diagram showing the configuration of the audio processing system according to Example 1. The audio processing system in FIG. 1 includes a wireless microphone slave unit (hereinafter referred to as "slave unit") 100 and a wireless microphone master unit (hereinafter referred to as "master unit") 200. The wireless microphone master unit 200 can be used together with an imaging device such as a digital camera, for example. The usage scene of the audio processing system is not particularly limited, but in Example 1, it is assumed that the audio processing system is used in a shooting scene, the slave unit 100 is used at the position of the subject, and the master unit 200 is used at a position away from the subject (for example, the position of the photographer).

[0012] The slave unit 100 includes a slave unit microphone unit 101, a slave unit MPU (microprocessing unit) 102, a slave unit wireless circuit 103, a slave unit audio circuit 104, and a speaker unit 108.

[0013] The slave unit microphone unit 101 includes a circuit that converts the input voice into an electrical signal (voice signal), and is used to acquire the voice of the subject.

[0014] The slave unit MPU 102 controls the operation of the slave unit 100. The slave unit MPU 102 has a ROM (not shown) in which a program for controlling the operation of the slave unit 100 is stored, a RAM (not shown) for storing variables, and an EEPROM (electrically erasable and writable memory) (not shown) for storing various parameters. The slave unit MPU 102 controls the operation of the slave unit 100 by expanding and executing the program stored in the ROM in the RAM. For example, the slave unit MPU 102 controls the operation of the slave unit wireless circuit 103 and the operation of the slave unit audio circuit 104 by sending a control signal to the slave unit wireless circuit 103 and the slave unit audio circuit 104 in a predetermined communication method such as SPI communication or I2C communication.

[0015] The slave unit wireless circuit 103 transmits the voice signal obtained by the slave unit microphone unit 101 to the slave unit It is wirelessly transmitted to the parent device wireless circuit 203 of the parent device 200 via the audio circuit 104. The transmission (wireless communication) method is, for example, Bluetooth (registered trademark) or Zigbee (registered trademark). The audio signal obtained by the slave unit microphone unit 101 may be interpreted as an audio signal collected at the position of the subject (the position of the slave unit 100). The transmission from the slave unit wireless circuit 103 to the parent device wireless circuit 203 may be interpreted as a transmission from the position of the subject to the position of the photographer (the position of the parent device 200).

[0016] The slave unit audio circuit 104 includes an audio acquisition unit 105, an audio level adjustment unit 106, and an audio signal generation unit 107. The audio acquisition unit 105 acquires the audio signal obtained by the slave unit microphone unit 101 from the slave unit microphone unit 101. The audio level adjustment unit 106 adjusts the audio level of the audio signal obtained by the audio acquisition unit 105 (the audio signal obtained by the slave unit microphone unit 101). For example, the audio level adjustment unit 106 adjusts the audio level with the gain set in the register. The audio signal generation unit 107 generates an audio signal to be output to the speaker unit 108.

[0017] The speaker unit 108 outputs sound according to the audio signal output from the slave unit audio circuit 104 (audio signal generation unit 107).

[0018] The parent device 200 includes a parent device microphone unit 201, a parent device MPU 202, a parent device wireless circuit 203, and a parent device audio circuit 204.

[0019] The parent device microphone unit 201 includes a circuit that converts the input sound into an electrical signal (audio signal) and is used to acquire ambient sound.

[0020] The host device MPU 202 controls the operation of the host device 200. The host device MPU 202 incorporates a ROM (not shown) storing a program for controlling the operation of the host device 200, a RAM (not shown) for storing variables, and an EEPROM (electrically erasable and writable memory) (not shown) for storing various parameters. The host device MPU 202 controls the operation of the host device 200 by expanding and executing the program stored in the ROM in the RAM. For example, the host device MPU 202 controls the operations of the host device wireless circuit 203 and the host device audio circuit 204 by sending control signals to the host device wireless circuit 203 and the host device audio circuit 204 in a predetermined communication method such as SPI communication or I2C communication.

[0021] The host device wireless circuit 203 receives the audio signal transmitted from the slave device wireless circuit 103 of the slave device 100. The communication method of the slave device wireless circuit 103 is the same as that of the host device wireless circuit 203.

[0022] The host device audio circuit 204 includes a host device audio acquisition unit 205, a slave device audio acquisition unit 206, an audio level adjustment unit 207, an audio time difference detection unit 208, an audio attenuation rate calculation unit 209, an audio correction unit 210, an audio synthesis unit 211, and an audio frequency detection unit 212.

[0023] The host device audio acquisition unit 205 acquires the audio signal obtained by the host device microphone unit 201 from the host device microphone unit 201. The audio signal obtained by the host device microphone unit 201 may be interpreted as an audio signal collected at the position of the photographer (the position of the host device 200).

[0024] The slave device audio acquisition unit 206 acquires the audio signal received by the host device wireless circuit 203 from the host device wireless circuit 203. Hereinafter, the audio signal obtained by the slave device audio acquisition unit 206 (the audio signal obtained by the slave device microphone unit 101, transmitted from the slave device wireless circuit 103 by wireless communication, and received by the host device wireless circuit 203) is referred to as the "slave device audio signal".

[0025] The voice level adjustment unit 207 adjusts the voice level of the voice signal obtained by the master unit voice acquisition unit 205 (the voice signal obtained by the master unit microphone unit 201). For example, the voice level adjustment unit 207 adjusts the voice level with the gain set in the register. Hereinafter, the voice signal obtained by the master unit voice acquisition unit 205 and adjusted by the voice level adjustment unit 207 (the voice signal obtained by the master unit microphone unit 201) is referred to as the "master unit voice signal".

[0026] The voice time difference detection unit 208 acquires the time difference between the master unit voice signal and the slave unit voice signal. The voice time difference detection unit 208 includes a circuit that detects the time difference between the master unit voice signal and the slave unit voice signal. Hereinafter, this time difference is referred to as the "voice time difference".

[0027] The voice time difference is the difference between the propagation time of the voice signal in the wireless communication of the voice signal from the position of the subject (the position of the slave unit 100) to the position of the photographer (the position of the master unit 200) and the propagation time of the voice in the air propagation of the voice from the position of the subject to the position of the photographer. Hereinafter, the propagation time in this wireless communication is referred to as the "wireless propagation time", and the propagation time in this air propagation is referred to as the "air propagation time". The wireless propagation time may be interpreted as the time from when the slave unit voice signal is transmitted from the slave unit 100 (the slave unit wireless circuit 103) until it reaches the master unit 200 (the master unit wireless circuit 203). The air propagation time may be interpreted as the time from when the voice is emitted from the subject or the slave unit 100 (the speaker unit 108) until it reaches the master unit 200 (the master unit microphone unit 201).

[0028] The voice attenuation rate calculation unit 209 acquires the attenuation rate of the voice in the air propagation of the voice from the position of the subject to the position of the photographer. The voice attenuation rate calculation unit 209 includes a circuit that calculates (detects) the ratio of the voice level of the master unit voice signal and the voice level of the slave unit voice signal as the attenuation rate. Hereinafter, this attenuation rate is referred to as the "voice attenuation rate".

[0029] The voice correction unit 210 corrects the master unit voice signal based on the voice time difference obtained by the voice time difference detection unit 208 and the voice attenuation rate obtained by the voice attenuation rate calculation unit 209. In the first embodiment, the voice correction unit 210 removes (cancels) from the master unit voice signal a voice signal obtained by delaying the slave unit voice signal based on the voice time difference and attenuating the slave unit voice signal based on the voice attenuation rate.

[0030] The voice synthesis unit 211 synthesizes the master unit voice signal and the slave unit voice signal. When the master unit voice signal is corrected by the voice correction unit 210, the voice synthesis unit 211 synthesizes the corrected master unit voice signal and the slave unit voice signal.

[0031] The voice frequency detection unit 212 detects the frequency of the master unit voice signal or the frequency of the slave unit voice signal. The voice frequency detection unit 212 can also preset (specify) a frequency band with a register and detect that the frequency of the voice signal is included in the set frequency band (predetermined frequency band).

[0032] FIG. 2 is a flowchart showing the basic setting process of the slave unit 100. For example, when the power of the slave unit 100 is turned on, the basic setting process in FIG. 2 starts.

[0033] In step S201, the slave unit MPU 102 sets the variables and programs stored in the RAM to the initial state and executes preparatory operations such as power supply to the slave unit wireless circuit 103 and the slave unit voice circuit 104.

[0034] In step S202, the slave unit MPU 102 sets the communication parameters of the slave unit wireless circuit 103. The communication parameters of the slave unit wireless circuit 103 include, for example, the MAC address, SSID (Service Set Identifier), data channel, and transmission rate.

[0035] In step S203, the slave unit MPU 102 sets the voice processing parameters of the slave unit voice circuit 104. The voice processing parameters of the slave unit voice circuit 104 include, for example, the gain for adjusting the voice level, the frequency characteristics for the equalizer (voice quality adjustment), the filter for countermeasures against wind noise, and the generation / non-generation of the voice signal for output to the speaker unit 108. The voice processing parameters of the slave unit voice circuit 104 can be set by changing the registers of the slave unit voice circuit 104.

[0036] By performing the processes from step S201 to step S203, the settings of the slave unit 100 are completed, and it becomes possible to transmit the voice signal acquired by the slave unit microphone unit 101 to the master unit 200.

[0037] FIG. 3 is a flowchart showing the basic setting process of the master unit 200. For example, when the power of the master unit 200 is turned on, the basic setting process in FIG. 3 starts.

[0038] In step S301, the master unit MPU 202 sets the variables and programs stored in the RAM to the initial state, and executes preparatory operations such as power supply to the master unit wireless circuit 203 and the master unit voice circuit 204.

[0039] In step S302, the master unit MPU 202 sets the communication parameters of the master unit wireless circuit 203. The communication parameters of the master unit wireless circuit 203 include, for example, the MAC address, SSID, data channel, and transmission rate, similar to the communication parameters of the slave unit wireless circuit 103.

[0040] In step S303, the master unit MPU 202 sets the voice processing parameters of the master unit voice circuit 204. The voice processing parameters of the master unit voice circuit 204 include, for example, the gain for adjusting the voice level, the frequency characteristics for the equalizer (voice quality adjustment), and the filter for countermeasures against wind noise. The voice processing parameters of the master unit voice circuit 204 can be set by changing the registers of the master unit voice circuit 204.

[0041] In step S304, the master unit MPU 202 determines whether the master unit wireless circuit 203 is connected to the slave unit wireless circuit 103. If the master unit MPU 202 determines that the master unit wireless circuit 203 is connected to the slave unit wireless circuit 103, the process proceeds to step S305; if it determines that the master unit wireless circuit 203 is not connected, the process proceeds to step S306. For example, when the power of the slave unit 100 is OFF, the master unit wireless circuit 203 is not connected to the slave unit wireless circuit 103, and the process proceeds to step S306. It is also possible to use only the master unit 200 without performing wireless connection between the slave unit 100 and the master unit 200. In this case, the echo correction described later is unnecessary.

[0042] In step S305, the master unit MPU 202 instructs the master unit voice circuit 204 to execute voice processing including echo correction. Details of the voice processing including echo correction will be described later with reference to FIG. 5.

[0043] In step S306, the master unit MPU 202 instructs the master unit voice circuit 204 to execute voice processing without echo correction. In the voice processing without echo correction, for example, voice level adjustment, equalizer (sound quality adjustment), and filter processing for wind noise countermeasure are performed on the voice signal input from the master unit microphone unit 201. The voice processing without echo correction is continuously performed, for example, until the power of the master unit 200 is turned OFF. 。

[0044] By performing the processing from step S301 to step S306, the setting of the master unit 200 is completed, and it becomes possible to acquire the slave unit voice signal or the master unit voice signal.

[0045] FIG. 4 is a schematic diagram showing the usage scenes of the slave unit 100 and the master unit 200. In FIG. 4, the slave unit 100 is arranged near the mouth of the subject (for example, the interviewer) in order to acquire the voice of the subject. The master unit 200 is connected to the camera held by the photographer who is away from the subject. The master unit 200 receives the slave unit voice signal from the slave unit 100, synthesizes it with the master unit voice signal, and stores it in the camera 100. Note that the storage medium in which the synthesized voice signal obtained by synthesizing the master unit voice signal and the slave unit voice signal is stored is not limited to the camera.

[0046] The voice of the subject input to the slave unit 100 may be attenuated and delayed by air propagation and also input to the master unit 200. In this case, since there is a timing deviation between the voice of the subject represented by the slave unit voice signal and the voice of the subject represented by the master unit voice signal, a synthesized voice signal in which the voice of the subject has an echo is obtained.

[0047] Therefore, in the first embodiment, echo correction is performed. Echo correction is a process of removing the voice signal of the voice input to the slave unit 100 (the voice input to the master unit 200 after attenuation and delay) from the master unit voice signal. By performing echo correction, it is possible to suppress the obtaining of a synthesized voice signal in which the voice of the subject has an echo.

[0048] FIG. 5 is a flowchart showing voice processing including echo correction according to the first embodiment. For example, when the master unit voice circuit 204 receives an instruction from the master unit MPU 202 in step S305 of FIG. 3, it starts the voice processing of FIG. 5. The voice processing of FIG. 5 is repeatedly performed, for example, until the wireless connection between the slave unit 100 and the master unit 200 is released or until the power of the master unit 200 is turned off.

[0049] In step S501, the master unit voice circuit 204 performs normal voice processing. In normal voice processing, for example, similar to voice processing without echo correction, voice level adjustment, equalizer (voice quality adjustment), and filter processing for wind noise countermeasures are performed on the voice signal input from the master unit microphone unit 201. The voice level adjustment unit 207 is used for voice level adjustment.

[0050] In step S502, the master unit voice circuit 204 acquires the voice time difference using the voice time difference detection unit 208. In step S503, the master unit voice circuit 204 acquires the voice attenuation rate using the voice attenuation rate calculation unit 209.

[0051] Using FIGS. 6(A) and 6(B), the method for acquiring the voice time difference and the voice attenuation rate will be described in detail. FIGS. 6(A) and 6(B) (and FIGS. 6(C) and 6(D) to be described later) are graphs showing the waveforms of various voice signals. In FIGS. 6(A) to 6(D), the horizontal axis represents time, and the vertical axis represents the voice level.

[0052] In the first embodiment, it is assumed that the voice time difference and the voice attenuation rate are acquired based on the master unit voice signal and the slave unit voice signal obtained by emitting a predetermined voice (for example, a voice of a predetermined frequency) from the speaker unit 108. Also, it is assumed that the positions of the slave unit 100 and the master unit 200 when acquiring the voice time difference and the voice attenuation rate are the same as the positions of the slave unit 100 and the master unit 200 when acquiring the voice from the subject. Hereinafter, the predetermined voice emitted from the speaker unit 108 is referred to as "speaker voice".

[0053] FIG. 6(A) shows the waveform of the slave unit voice signal obtained by emitting the speaker voice, and FIG. 6(B) shows the waveform of the master unit voice signal obtained by emitting the speaker voice.

[0054] First, the method for acquiring the voice time difference will be described.

[0055] The slave unit voice signal in Fig. 6(A) is a voice signal delayed by the above-mentioned radio propagation time. In the slave unit voice signal of Fig. 6(A), the waveform of the speaker sound appears after being delayed by the radio propagation time from the timing when the speaker sound was emitted. Since the radio propagation time is uniquely determined by the slave unit radio circuit 103 and the master unit radio circuit 203, information indicating the radio propagation time can be prepared in advance. In the first embodiment, it is assumed that information indicating the radio propagation time is stored in advance in the EEPROM of the master unit MPU202. The radio propagation time may be interpreted as the delay time from the timing when the sound was emitted from the speaker unit 108 or the subject to the timing when the waveform of the sound appears in the slave unit voice signal.

[0056] The master unit voice signal in Fig. 6(B) is a voice signal delayed by the above-mentioned air propagation time. In the master unit voice signal of Fig. 6(B), the waveform of the speaker sound appears after being delayed by the air propagation time from the timing when the speaker sound was emitted. Since the radio propagation time and the air propagation time are different, in the master unit voice signal of Fig. 6(B), the waveform of the speaker sound appears at a timing different from the timing when the waveform of the speaker sound appears in the slave unit voice signal of Fig. 6(A). The air propagation time may be interpreted as the delay time from the timing when the sound was emitted from the speaker unit 108 or the subject to the timing when the waveform of the sound appears in the master unit voice signal.

[0057] The voice frequency detection unit 212 detects the frequency of the speaker sound from each of the slave unit voice signal and the master unit voice signal. Thereby, the timing of the speaker sound in the slave unit voice signal and the timing of the speaker sound in the master unit voice signal are detected. The voice time difference detection unit 208 calculates the time Δt from the timing of the speaker sound in the slave unit voice signal to the timing of the speaker sound in the master unit voice signal as the voice time difference (the difference between the radio propagation time and the air propagation time).

[0058] Next, a method for obtaining the voice attenuation rate will be described.

[0059] As described above, the audio frequency detection unit 212 detects the frequency of the speaker sound from each of the handset audio signal and the base unit audio signal. As a result, the period of the speaker sound in the handset audio signal and the period of the speaker sound in the base unit audio signal are detected. Then, the audio attenuation rate calculation unit 209 calculates the ratio Vrx / Vtx of the audio level (amplitude Vrx) of the speaker sound in the base unit audio signal to the audio level (amplitude Vtx) of the speaker sound in the handset audio signal as the audio attenuation rate. It is assumed that the amplitude Vtx is a value obtained by dividing the handset audio signal by the gain Gaintx set in the audio level adjustment unit 106, and the amplitude Vrx is a value obtained by dividing the base unit audio signal by the gain Gainrx set in the audio level adjustment unit 207. By performing normalization using the gains Gaintx and Gainrx, it is possible to obtain the ratio Vrx / Vtx that more accurately represents the attenuation rate (attenuation amount) of the audio in the air propagation of the audio from the position of the subject to the position of the photographer.

[0060] Note that the method for obtaining the audio time difference and the audio attenuation rate is not limited to the above method. For example, the audio time difference and the audio attenuation rate may be obtained based on the base unit audio signal and the handset audio signal obtained when the subject emits a sound. Also, although it has been described that the audio processing in FIG. 5 is repeated, the audio time difference and the audio attenuation rate do not necessarily need to be obtained repeatedly. For example, the process of obtaining the audio time difference and the audio attenuation rate is performed only once, and the obtained audio time difference and audio attenuation rate are repeatedly used in step S505 described later. The base unit 200 may obtain the audio time difference and the audio attenuation rate from the outside.

[0061] Return to the description of FIG. 5. In step S504, the master unit voice circuit 204 determines whether the slave unit voice signal is a voice signal in a predetermined frequency band using the voice frequency detection unit 212. If the master unit voice circuit 204 determines that the slave unit voice signal is a voice signal in a predetermined frequency band, the process proceeds to step S505. If it determines that the slave unit voice signal is not a voice signal in a predetermined frequency band, the process proceeds to step S506. In the first embodiment, since the slave unit 100 is used to acquire the voice of the subject, a frequency band of 120 Hz or more and 300 Hz or less, which is the frequency band of the human voice, is used as the predetermined frequency band.

[0062] In step S505, the master unit voice circuit 204 performs echo correction on the master unit voice signal using the voice correction unit 210. By making the determination in step S504, only the voice signal in the predetermined frequency band in the master unit voice signal becomes the target of echo correction. Note that step S504 may be omitted.

[0063] The echo correction will be described in detail. The voice signal in FIG. 6(C) is a combined voice signal of the slave unit voice signal in FIG. 6(A) and the master unit voice signal in FIG. 6(B). Since the timings of the speaker sounds in the slave unit voice signal in FIG. 6(A) and the master unit voice signal in FIG. 6(B) are different, in the combined voice signal in FIG. 6(C), the waveforms of the same speaker sound appear at two locations. Similarly, in the case of the voice of the subject rather than the speaker sound, the waveforms of the same voice of the subject appear at two locations in the combined voice signal. In order to solve such a problem, echo correction is performed. The echo processing is performed using the voice time difference Δt obtained in step S502 and the voice attenuation rate Vrx / Vtx obtained in step S503.

[0064] The voice correction unit 210 calculates a correction value C(t) from the slave unit voice signal ftx(t), the voice time difference Δt, and the voice attenuation rate Vrx / Vtx according to the following equation 1. The variable t is the timing (time position). The slave unit voice signal ftx(t) may be interpreted as the voice level of the slave unit voice signal.

Equation

[0065] Then, the voice correction unit 210 calculates the parent unit voice signal frx2(t) after echo correction from the correction value C(t), the gains Gaintx and Gainrx, and the parent unit voice signal frx(t) according to the following Equation 2 (correction formula). The parent unit voice signal frx(t) before echo correction may be interpreted as the voice level of the parent unit voice signal before echo correction. The parent unit voice signal frx2(t) after echo correction may be interpreted as the voice level of the parent unit voice signal after echo correction.

Equation

[0066] For example, by using a dedicated hard block constructed of a sequential circuit such as a flip-flop and subtracting the correction value based on the handset voice signal Δt time before from the parent unit voice signal, the echo correction of Equation 2 can be realized. When the parent unit voice circuit 204 does not have a dedicated hard block, the correction value based on the handset voice signal is stored in the RAM of the parent unit MPU202, the correction value Δt minutes before is read from the RAM, and subtracted from the parent unit voice signal. Thereby, the echo correction of Equation 2 can be realized.

[0067] Returning to the description of FIG. 5, in step S506, the parent unit voice circuit 204 synthesizes the parent unit voice signal and the handset voice signal using the voice synthesis unit 211. When the echo correction in step S505 is performed, the parent unit voice signal frx2(t) after echo correction and the handset voice signal ftx(t) are synthesized according to the following Equation 3, and the synthesized voice signal f(t) is obtained. The synthesized voice signal f(t) may be interpreted as the voice level of the synthesized voice signal.

Equation

[0068] The audio signal in FIG. 6(D) is a composite audio signal of the slave unit audio signal in FIG. 6(A) and the audio signal obtained by performing echo correction on the master unit audio signal in FIG. 6(B). In the composite audio signal in FIG. 6(D), the waveform of the speaker sound appears only at one location, and the problem described with reference to FIG. 6(C) is solved. Similarly, in the case of the sound of the subject rather than the speaker sound, the problem described with reference to FIG. 6(C) can be solved.

[0069] As described above, according to the first embodiment, the audio time difference and the audio attenuation rate are obtained, and the audio signal obtained by delaying the slave unit audio signal based on the audio time difference and attenuating the slave unit audio signal based on the audio attenuation rate is removed (cancelled) from the master unit audio signal. By doing so, the audio signal of the audio input to the slave unit away from the master unit (the audio input with a delay to the master unit) can be removed from the audio signal obtained by the master unit. As a result, it is possible to suppress the acquisition of a composite audio signal in which an echo is added to the sound of the subject.

[0070] <Second Embodiment> The second embodiment of the present invention will be described. Hereinafter, the description of the same points as those in the first embodiment (for example, the same configuration and processing as those in the first embodiment) will be omitted, and the points different from those in the first embodiment will be described.

[0071] FIG. 7 is a block diagram showing the configuration of a photographing system according to the second embodiment. The photographing system of FIG. 7 includes a slave unit 100, a master unit 200, and a camera 700. The slave unit 100 has the same configuration as that in the first embodiment. The master unit 200 has substantially the same configuration as that in the first embodiment. The master unit audio circuit 204 of the master unit 200 has a different configuration from that in the first embodiment. The master unit audio circuit 204 includes, as in the first embodiment, a master unit audio acquisition unit 205, a slave unit audio acquisition unit 206, an audio level adjustment unit 207, an audio time difference detection unit 208, an audio attenuation rate calculation unit 209, an audio correction unit 210, an audio synthesis unit 211, and an audio frequency detection unit 212. Further, the master unit audio circuit 204 includes a distance information acquisition unit 213. The camera 700 includes a distance detection circuit 701, a camera MPU 702, and an imaging circuit 703. Note that the master unit 200 and the camera 700 may be an integrated device.

[0072] The distance detection circuit 701 is a circuit that calculates (detects) the distance from the subject to the camera 700 by a known method such as TOF (Time of Flight), and includes a sensor for calculating the distance.

[0073] The imaging circuit 703 includes an image sensor. The imaging circuit 703 can also calculate (detect) the distance from the subject to the camera 700 by a known method. For example, the imaging circuit 703 calculates the distance from the subject to the camera 700 based on the image obtained by the image sensor.

[0074] The camera MPU 702 controls the operation of the camera 700. The camera MPU 702 incorporates a ROM (not shown) storing a program for controlling the operation of the slave unit 700, a RAM (not shown) for storing variables, and an EEPROM (electrically erasable and writable memory) (not shown) for storing various parameters. The camera MPU 702 controls the operation of the camera 700 by expanding and executing the program stored in the ROM in the RAM. For example, the camera MPU 702 controls the operation of the distance detection circuit 701 and the operation of the imaging circuit 703 by sending a control signal to the distance detection circuit 701 and the imaging circuit 703 in a predetermined communication method such as SPI communication or I2C communication.

[0075] In addition, the camera MPU 702 can communicate with the master unit MPU 202 via an external IF (not shown) provided in the camera 700 and an external IF (not shown) provided in the master unit 202. For example, the camera MPU 702 transmits information indicating the distance (the distance from the subject to the camera 700) calculated by the distance detection circuit 701 or the imaging circuit 703 to the master unit MPU 202 in a predetermined communication method such as SPI communication or I2C communication.

[0076] The distance information acquisition unit 213 of the master unit 200 acquires shooting distance information indicating the distance from the subject to the camera 700 from the camera MPU 702 via the master unit MPU 202. In the second embodiment, it is assumed that the distance from the subject to the camera 700 is equal to the distance from the subject to the master unit 200. Therefore, the shooting distance information may be interpreted as information indicating the distance from the subject to the master unit 200. Note that the method for acquiring the shooting distance information is not particularly limited. For example, the master unit 200 may have a sensor for acquiring the shooting distance information.

[0077] FIG. 8 is a flowchart showing audio processing including echo correction according to the second embodiment. For example, when the master unit audio circuit 204 receives an instruction from the master unit MPU 202 at step S305 in FIG. 3, it starts the audio processing in FIG. 5. The audio processing in FIG. 8 is repeatedly performed, for example, until the wireless connection between the slave unit 100 and the master unit 200 is released or until the power of the master unit 200 is turned off.

[0078] Step S801 is the same as step S501 in FIG. 5. In step S802, the master unit audio circuit 204 determines whether there is shooting distance information. If the master unit audio circuit 204 determines that there is shooting distance information, it performs the processes in steps S803 and S804. If it determines that there is no shooting distance information, it performs the processes in steps S805 and S806. Steps S805 and S806 are the same as steps S502 and S503 in FIG. 5.

[0079] In step S803, the master unit audio circuit 204 acquires the air propagation time based on the shooting distance information and obtains the audio time difference. For obtaining the audio time difference, the audio time difference detection unit 208 is used in the same manner as in step S805.

[0080] Using FIGS. 9(A) to 9(C), a method for obtaining the audio time difference based on the shooting distance information will be described in detail. FIGS. 9(A) to 9(C) are graphs showing the waveforms of various audio signals. In FIGS. 9(A) to 9(C), the horizontal axis represents time, and the vertical axis represents the audio level. The audio signal in FIG. 9(A) is the audio signal obtained by the slave unit microphone unit 101 and input to the slave unit audio circuit 104. The audio signal in FIG. 9(B) is the slave unit audio signal, and the audio signal in FIG. 9(C) is the master unit audio signal.

[0081] The time trf from the timing of the speaker sound in the audio signal of FIG. 9(A) to the timing of the speaker sound in the audio signal of FIG. 9(B) corresponds to the radio propagation time. As described in the first embodiment, the radio propagation time is the propagation time of the audio signal in the wireless communication of the audio signal from the position of the subject (the position of the slave unit 100) to the position of the photographer (the position of the master unit 200). And since the radio propagation time is uniquely determined by the slave unit wireless circuit 103 and the master unit wireless circuit 203, the information indicating the radio propagation time can be prepared in advance. In the second embodiment, it is assumed that the information indicating the radio propagation time is stored in advance in the EEPROM of the master unit MPU 202. Since the radio propagation time is uniquely determined by the slave unit wireless circuit 103 and the master unit wireless circuit 203, the information indicating the radio propagation time can be prepared in advance. In the second embodiment, it is assumed that the information indicating the radio propagation time is stored in advance in the EEPROM of the master unit MPU 202.

[0082] The time tdis from the timing of the speaker sound in the audio signal of FIG. 9(A) to the timing of the speaker sound in the audio signal of FIG. 9(C) corresponds to the air propagation time. The air propagation time is the propagation time of the audio from the position of the subject (the position of the slave unit 100) to the position of the photographer (the position of the master unit 200). The audio time difference detection unit 208 calculates the time tdis by dividing the distance indicated by the shooting distance information by the speed of sound. For example, if the distance from the master unit 200 to the subject is 34 m and the speed of sound is 340 m / sec, then 0.1 sec is calculated as the time tdis.

[0083] The voice time difference detection unit 208 calculates the voice time difference Δt by subtracting the time trf from the time tdis according to the following formula 4. For example, when the distance from the master unit 200 to the subject is 34 m, the speed of sound is 340 m / sec, and the radio propagation time is 0.01 sec, 0.09 sec is calculated as the voice time difference Δt.

Number

[0084] Return to the description of FIG. 8. In step S804, the master unit voice circuit 204 acquires the voice attenuation rate based on the shooting distance information. For the acquisition of the voice attenuation rate, the voice attenuation rate calculation unit 209 is used in the same manner as in step S806.

[0085] Using FIG. 4, the method for acquiring the voice attenuation rate based on the shooting distance information will be described in detail. In FIG. 4, the distance d1 from the subject to the slave unit 100 and the distance d2 from the subject to the master unit 200 are shown. In the second embodiment, it is assumed that the information indicating the distance d1 is stored in advance in the EEPROM of the master unit MPU 202. The distance d2 is the distance indicated by the shooting distance information.

[0086] Since sound spreads spherically, the voice attenuation rate calculation unit 209 calculates the voice attenuation rate Vatt in decibels from the distances d1 and d2 according to the following formula 5. For example, when the distance d1 is 0.1 m and the distance d2 is 10 m, 40 dB is calculated as the voice attenuation rate Vatt.

Number

[0087] The voice attenuation rate calculation unit 209 converts the voice attenuation rate Vatt to the voice attenuation rate Vrx / Vtx according to the following formula 5.

Number

[0088] Steps S807 to S809 are the same as steps S504 to S506 in FIG. 5.

[0089] As described above, according to the second embodiment, shooting distance information is obtained, and based on the shooting distance information, an audio time difference and an audio attenuation rate are obtained. Then, similar to the first embodiment, an audio signal obtained by delaying the slave unit audio signal based on the audio time difference and attenuating the slave unit audio signal based on the audio attenuation rate is removed (cancelled) from the master unit audio signal. By doing so, the same effect as that of the first embodiment can be obtained. Furthermore, by using the shooting distance information, it becomes possible to obtain the audio time difference and the audio attenuation rate without using the speaker sound, and it is possible to suppress the occurrence of situations such as the subject or the photographer being surprised or feeling uncomfortable due to the speaker sound.

[0090] Note that the above embodiments (including modified examples) are merely examples, and configurations obtained by appropriately modifying or changing the configurations of the above embodiments within the scope of the gist of the present invention are also included in the present invention. Configurations obtained by appropriately combining the configurations of the above embodiments are also included in the present invention.

[0091] <Other Embodiments> The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0092] The disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) First audio acquisition means for acquiring a first audio signal collected at a first position; Second audio acquisition means for acquiring a second audio signal collected at a second position and transmitted from the second position by wireless communication and received at the first position; A time difference acquisition means for acquiring a difference between the propagation time of the audio signal in the wireless communication of the audio signal from the second position to the first position and the propagation time of the audio in the air propagation of the audio from the second position to the first position; An attenuation rate acquisition means for acquiring an attenuation rate of the audio in the air propagation of the audio from the second position to the first position; A correction means for removing a third audio signal obtained by attenuating the second audio signal based on the attenuation rate and delaying the second audio signal based on the difference from the first audio signal; An audio processing apparatus, characterized by comprising the above. (Configuration 2) A synthesizing means for synthesizing a fourth audio signal obtained by removing the third audio signal from the first audio signal and the second audio signal; Further comprising The audio processing apparatus according to Configuration 1, characterized by the above. (Configuration 3) Based on the first audio signal and the second audio signal obtained by emitting predetermined audio at the second position, the attenuation rate acquisition means acquires the attenuation rate, and the time difference acquisition means acquires the difference; The audio processing apparatus according to Configuration 1 or 2, characterized by the above. (Configuration 4) The time difference acquisition means acquires, as the difference between the propagation time in the wireless communication and the propagation time in the air propagation, the time from the timing of the predetermined audio in the second audio signal to the timing of the predetermined audio in the first audio signal; The audio processing apparatus according to Configuration 3, characterized by the above. (Configuration 5) The attenuation rate acquisition means acquires, as the attenuation rate, a ratio of the level of the predetermined audio in the first audio signal to the level of the predetermined audio in the second audio signal; The audio processing apparatus according to Configuration 3 or 4, characterized by the above. (Configuration 6) Information on the propagation time in the wireless communication is prepared in advance, The time difference acquisition means acquires the propagation time in the air propagation based on the distance from the first position to the second position, and acquires the difference. The voice processing apparatus according to Configuration 1 or 2, characterized by the above. (Configuration 7) The attenuation rate acquisition means acquires the attenuation rate based on the distance from the first position to the second position. The voice processing apparatus according to Configuration 1, 2, or 6, characterized by the above. (Configuration 8) Information acquisition means for acquiring information indicating the distance from the first position to the second position further comprising The voice processing apparatus according to Configuration 6 or 7, characterized by the above. (Configuration 9) The third voice signal is a voice signal in a predetermined frequency band. The voice processing apparatus according to any one of Configurations 1 to 8, characterized by the above. (Method) acquiring a first voice signal collected at a first position; acquiring a second voice signal collected at a second position, transmitted from the second position by wireless communication, and received at the first position; acquiring a difference between the propagation time of the voice signal in the wireless communication of the voice signal from the second position to the first position and the propagation time of the voice in the air propagation of the voice from the second position to the first position; acquiring an attenuation rate of the voice in the air propagation of the voice from the second position to the first position; removing, from the first voice signal, a third voice signal obtained by attenuating the second voice signal based on the attenuation rate and delaying the second voice signal based on the difference; A voice processing method, characterized by comprising the above. (Program) A program for causing a computer to function as each means of the voice processing apparatus according to any one of Configurations 1 to 9.

Explanation of Reference Numerals

[0093] 200: Wireless microphone main unit 204: Main unit voice circuit 205: Main unit voice acquisition unit 206: Sub-unit voice acquisition unit 208: Voice time difference detection unit 209: Voice attenuation rate calculation unit 210: Voice correction unit

Claims

1. a first audio acquisition means for acquiring a first audio signal collected at a first position; a second audio acquisition means for acquiring a second audio signal collected at a second position, transmitted from the second position by wireless communication, and received at the first position; a time difference acquisition means for acquiring a difference between a propagation time of the audio signal in the wireless communication of the audio signal from the second position to the first position and a propagation time of the audio in the air propagation of the audio from the second position to the first position; an attenuation rate acquisition means for acquiring an attenuation rate of the audio in the air propagation of the audio from the second position to the first position; a correction means for removing, from the first audio signal, a third audio signal obtained by attenuating the second audio signal based on the attenuation rate and delaying the second audio signal based on the difference; An audio processing apparatus, characterized by comprising the above.

2. a combining means for combining a fourth audio signal obtained by removing the third audio signal from the first audio signal and the second audio signal; The audio processing apparatus according to claim 1, further comprising:

3. Based on the first audio signal and the second audio signal obtained by emitting a predetermined audio at the second position, the attenuation rate acquisition means acquires the attenuation rate, and the time difference acquisition means acquires the difference. The audio processing apparatus according to claim 1, characterized by the above.

4. The time difference acquisition means acquires, as the difference between the propagation time in the wireless communication and the propagation time in the air propagation, the time from the timing of the predetermined audio in the second audio signal to the timing of the predetermined audio in the first audio signal. The audio processing apparatus according to claim 3, characterized by the above.

5. The attenuation rate acquisition means acquires, as the attenuation rate, a ratio of a level of the predetermined voice in the first audio signal to a level of the predetermined voice in the second audio signal The audio processing apparatus according to claim 3, characterized in that

6. Information on the propagation time in the wireless communication is prepared in advance, The time difference acquisition means acquires the propagation time in the air propagation based on the distance from the first position to the second position, and acquires the difference The audio processing apparatus according to claim 1, characterized in that

7. The attenuation rate acquisition means acquires the attenuation rate based on the distance from the first position to the second position The audio processing apparatus according to claim 1, characterized in that

8. Information acquisition means for acquiring information indicating the distance from the first position to the second position further comprising The audio processing apparatus according to claim 6, characterized in that

9. The third audio signal is an audio signal in a predetermined frequency band The audio processing apparatus according to claim 1, characterized in that

10. acquiring a first audio signal collected at a first position; acquiring a second audio signal collected at a second position, transmitted from the second position by wireless communication, and received at the first position; acquiring a difference between a propagation time of the audio signal in the wireless communication from the second position to the first position and a propagation time of the audio in the air propagation from the second position to the first position; acquiring an attenuation rate of the audio in the air propagation from the second position to the first position; A step of removing a third audio signal obtained by attenuating the second audio signal based on the attenuation rate and delaying the second audio signal based on the difference from the first audio signal An audio processing method characterized by including the above.

11. A program for causing a computer to function as each means of the audio processing apparatus according to any one of Claims 1 to 9.

Citation Information

Patent Citations

  • Sound recording method for imaging device

    JP2012100235A