Audio processing device and control method thereof
The audio processing device in digital cameras uses multiple microphones and adaptive amplification to manage varying drive noise from lenses, reducing distortion and improving audio quality by lens-specific noise reduction.
Patent Information
- Application Number
- JP2021091350
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-05-31
AI Technical Summary
Existing audio processing devices in digital cameras, such as interchangeable-lens cameras, face challenges in managing varying drive noise levels from optical lenses, which can lead to audio signal distortion and quantization noise, especially when amplification is uniform across different lenses.
The audio processing device employs multiple microphones to capture environmental and drive noise, adjusts amplification based on lens type and noise level, using Fourier transforms and inverse transforms to generate and reduce noise data, thereby controlling amplification effectively.
This approach allows for appropriate amplification control, reducing noise and preventing audio signal distortion, enhancing audio quality by addressing the variability in drive noise from different optical lenses.
Smart Images

Figure 0007725244000001 
Figure 0007725244000002 
Figure 0007725244000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio processing device capable of reducing noise contained in audio data. [Background technology]
[0002] A digital camera, which is an example of an audio processing device, can record surrounding audio when recording video data. Digital cameras also have an autofocus function that adjusts the focus on a subject while recording video data by driving an optical lens. Digital cameras also have a zoom function that drives the optical lens while recording video.
[0003] In this way, when an optical lens is driven while recording a video, the driving sound of the optical lens may be included as noise in the audio recorded along with the video. Hereinafter, the driving sound of the optical lens that is recorded as noise will be referred to as driving noise. In response to this, when a digital camera picks up driving noise, it can reduce the driving noise and record the surrounding audio. Patent Document 1 discloses that audio distortion and crackling can be reduced by changing the amount of analog gain according to the volume setting value. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-141571 Summary of the Invention [Problem to be solved by the invention]
[0005] However, for example, in an interchangeable-lens digital camera, the magnitude of drive noise varies depending on the optical lens attached. Furthermore, even during lens operation, there are periods when the drive noise is high and periods when the drive noise is low. If such noise is amplified uniformly, for example, there is a risk of audio signal distortion or quantization noise occurring during audio processing, which could result in a deterioration in the quality of the audio signal.
[0006] Therefore, an object of the present invention is to appropriately control the amplification process for input noise. [Means for solving the problem]
[0007] The audio processing device of the present invention includes a first microphone for acquiring environmental sound, a second microphone for acquiring sound from a noise source that is a driving member included in an optical lens, a first amplifier for amplifying an audio signal input from the first microphone, a second amplifier for amplifying the audio signal input from the second microphone according to an amplification amount, a first conversion means for Fourier transforming the audio signal input from the first amplifier means to generate a first audio signal, a second conversion means for Fourier transforming the audio signal input from the second amplifier means to generate a second audio signal, a generation means for generating noise data using the second audio signal and parameters related to noise from the noise source, and a conversion means for converting the noise data into a signal related to the noise source. and a third transforming means for performing an inverse Fourier transform on the audio signal output from the reduction means, wherein the second amplifying means sets the amplification amount based on the type of the optical lens or the level of the audio signal output from the second amplifying means, and when the second amplifying means sets the amplification amount based on the type of the optical lens, the second amplifying means sets the amplification amount to a first value if the optical lens is a first optical lens, and sets the amplification amount to a second value smaller than the first value if the optical lens is a second optical lens that generates greater noise than the first optical lens. [Effects of the Invention]
[0008] The audio processing device of the present invention can appropriately control the amplification process for input noise. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a perspective view of an imaging device 100 according to a first embodiment. [Figure 2] FIG. 1 is a block diagram showing a configuration of an imaging device 100 according to a first embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of a voice input unit 104 of the imaging device 100 according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing the arrangement of microphones in an audio input unit 104 of an imaging device 100 according to the first embodiment. [Figure 5] FIG. 2 is a block diagram for explaining an amplifier 202 in the first embodiment. [Figure 6] 10 is a flowchart showing control when the amplification amount is changed using a level detection unit 204 in the first embodiment. [Figure 7] 10 is a flowchart showing control when the amount of amplification is changed using the lens control unit 102 in the first embodiment. [Figure 8] 10 is a flowchart showing control when the amplification amount is changed using the level detection unit 204 and the lens control unit 102 in the first embodiment. [Figure 9] FIG. 4 is a diagram illustrating noise parameters in the first embodiment. [Figure 10] 10A and 10B are diagrams illustrating the frequency spectrum of audio and the frequency spectrum of noise parameters when a drive sound occurs in a situation where it is considered that there is no environmental sound in the first embodiment. [Figure 11] FIG. 10 is a diagram showing the frequency spectrum of audio when a drive sound is generated in a situation where environmental sound is present in the first embodiment. [Figure 12]FIG. 2 is a block diagram showing a noise data generation processing method in a noise data generator 206 in the first embodiment. [Figure 13] FIG. 10 is a block diagram showing a noise data generation processing method in the noise data generator 206 in the case where the audio signal collected by the noise microphone 201c is affected by quantization noise in the first embodiment. [Figure 14] FIG. 10 is a diagram showing the frequency spectrum of an audio signal collected by a noise microphone 201c in the first embodiment when the audio signal is affected by quantization noise. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0011] [First Example] <External view of imaging device 100> 1(a) and 1(b) show an example of the appearance of an image capture device 100 as an example of an audio processing device to which the present invention can be applied. Fig. 1(a) is an example of a front perspective view of the image capture device 100. Fig. 1(b) is an example of a rear perspective view of the image capture device 100. In Fig. 1, an optical lens (not shown) is attached to a lens mount 301.
[0012] The display unit 107 displays image data, text information, etc. The display unit 107 is provided on the rear surface of the imaging device 100. The extra-finder display unit 43 is a display unit provided on the top surface of the imaging device 100. The extra-finder display unit 43 displays the settings of the imaging device 100, such as the shutter speed and aperture value. The eyepiece viewfinder 16 is a peer-type viewfinder. The user can check the focus and composition of the optical image of the subject by observing the focusing screen inside the eyepiece viewfinder 16.
[0013] The release switch 61 is an operation member that allows the user to issue shooting instructions. The mode selector switch 60 is an operation member that allows the user to switch between various modes. The main electronic dial 71 is a rotary operation member. By turning this main electronic dial 71, the user can change the settings of the imaging device 100, such as the shutter speed and aperture value. The release switch 61, mode selector switch 60, and main electronic dial 71 are included in the operation unit 112.
[0014] The power switch 72 is an operating member that switches the power of the imaging device 100 on and off. The sub electronic dial 73 is a rotary operating member. The user can use the sub electronic dial 73 to move the selection frame displayed on the display unit 107 and to advance images in playback mode. The cross key 74 is a cross key (four-way key) that can be pressed up, down, left, or right. The imaging device 100 performs processing according to the part (direction) of the cross key 74 that is pressed. The power switch 72, sub electronic dial 73, and cross key 74 are included in the operation unit 112.
[0015] The SET button 75 is a push button. The SET button 75 is mainly used by the user to confirm a selection item displayed on the display unit 107, etc. The LV button 76 is a button used to switch live view (hereinafter referred to as LV) on and off. In video recording mode, the LV button 76 is used to instruct the start and stop of video shooting (recording). The enlarge button 77 is a push button used to turn enlargement mode on and off in live view display in shooting mode, and to change the magnification ratio in enlargement mode. The SET button 75, LV button 76, and enlargement button 77 are included in the operation unit 112.
[0016] In playback mode, the enlarge button 77 functions as a button for increasing the magnification of image data displayed on the display unit 107. The reduce button 78 is a button for decreasing the magnification of image data enlarged and displayed on the display unit 107. The play button 79 is an operation button for switching between shooting mode and playback mode. When the user presses the play button 79 while the imaging device 100 is in shooting mode, the imaging device 100 transitions to playback mode, and image data recorded on the recording medium 110 is displayed on the display unit 107. The reduce button 78 and play button 79 are included in the operation unit 112.
[0017] The quick-return mirror 12 (hereinafter, mirror 12) is a mirror that switches the light beam incident from an optical lens attached to the imaging device 100 so that it is incident on either the eyepiece finder 16 side or the imaging unit 101 side. The mirror 12 is raised and lowered by the control unit 111 controlling an actuator (not shown) during exposure, live view shooting, and video shooting. The mirror 12 is normally positioned so that the light beam is incident on the eyepiece finder 16. When shooting or in live view display, the mirror 12 flips up (mirror up) so that the light beam is incident on the imaging unit 101. The center of the mirror 12 is a half mirror. A portion of the light beam that passes through the center of the mirror 12 is incident on a focus detection unit (not shown) that performs focus detection.
[0018] The communication terminal 10 is a communication terminal for communication between the imaging device 100 and an optical lens 300 attached to the imaging device 100. The terminal cover 40 is a cover for protecting a connector (not shown) such as a connection cable that connects a connection cable with an external device to the imaging device 100. The lid 41 is a lid for a slot that stores the recording medium 110. The lens mount 301 is an attachment portion to which the optical lens 300 (not shown) can be attached.
[0019] The L microphone 201a and the R microphone 201b are microphones for picking up the user's voice, etc. When viewed from the rear of the imaging device 100, the L microphone 201a is placed on the left side and the R microphone 201b is placed on the right side.
[0020] <Configuration of imaging device 100> FIG. 2 is a block diagram showing an example of the configuration of the imaging device 100 according to this embodiment.
[0021] The optical lens 300 is a lens unit that can be attached to or detached from the imaging device 100. For example, the optical lens 300 is a zoom lens or a varifocal lens. The optical lens 300 has an optical lens, a motor for driving the optical lens, and a communication unit that communicates with a lens control unit 102 of the imaging device 100 (described later). The optical lens 300 can focus and zoom on a subject and correct camera shake by moving the optical lens using the motor based on a control signal received by the communication unit.
[0022] The imaging unit 101 has an imaging element for converting an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and an image processing unit for generating and outputting image data or video data from the electrical signal generated by the imaging element. The imaging element is, for example, a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS). In this embodiment, a series of processes for generating image data including still image data and video data in the imaging unit 101 and outputting the image data from the imaging unit 101 is referred to as "photographing." In the imaging device 100, the image data is recorded on a recording medium 110 (described later) in accordance with the DCF (Design rule for Camera File system) standard.
[0023] The lens control unit 102 transmits a control signal to the optical lens 300 via the communication terminal 10 based on the data output from the imaging unit 101 and a control signal output from the control unit 111 described later, and controls the optical lens 300.
[0024] The information acquisition unit 103 detects the tilt of the image capture device 100 and the temperature inside the housing of the image capture device 100. For example, the information acquisition unit 103 detects the tilt of the image capture device 100 using an acceleration sensor or a gyro sensor. Also, for example, the information acquisition unit 103 detects the temperature inside the housing of the image capture device 100 using a temperature sensor.
[0025] The audio input unit 104 generates audio data from audio acquired by a microphone. The audio input unit 104 acquires audio around the imaging device 100 using the microphone, performs analog-to-digital conversion (A / D conversion) on the acquired audio, and performs various audio processes to generate audio data. In this embodiment, the audio input unit 104 has a microphone. A detailed configuration example of the audio input unit 104 will be described later.
[0026] The volatile memory 105 temporarily stores image data generated by the imaging unit 101 and audio data generated by the audio input unit 104. The volatile memory 105 is also used as a temporary storage area for image data displayed on the display unit 107, a work area for the control unit 111, etc.
[0027] The display control unit 106 controls the display unit 107 to display image data output from the imaging unit 101, text for interactive operations, menu screens, and the like. Furthermore, when capturing still images and moving images, the display control unit 106 controls the display unit 107 to sequentially display digital data output from the imaging unit 101, thereby allowing the display unit 107 to function as an electronic viewfinder. For example, the display unit 107 is a liquid crystal display or an organic EL display. Furthermore, the display control unit 106 can also control the image data and video data output from the imaging unit 101, text for interactive operations, menu screens, and the like to be displayed on an external display via the external output unit 115, which will be described later.
[0028] The encoding processing unit 108 can encode the image data and audio data temporarily stored in the volatile memory 105. For example, the encoding processing unit 108 can generate video data by encoding and compressing image data according to the JPEG standard or a RAW image format. For example, the encoding processing unit 108 can generate video data by encoding and compressing video data according to the MPEG2 standard or the H.264 / MPEG4-AVC standard. Furthermore, for example, the encoding processing unit 108 can generate audio data by encoding and compressing audio data according to the AC3AAC standard, the ATRAC standard, or the ADPCM method. Furthermore, the encoding processing unit 108 may encode audio data without data compression, for example, according to the linear PCM method.
[0029] The recording control unit 109 can record data to and read data from the recording medium 110. For example, the recording control unit 109 can record still image data, video data, and audio data generated by the encoding processing unit 108 to and read data from the recording medium 110. The recording medium 110 is, for example, an SD card, a CF card, an XQD memory card, an HDD (magnetic disk), an optical disk, or a semiconductor memory. The recording medium 110 may be configured to be detachable from the imaging device 100, or may be built into the imaging device 100. That is, the recording control unit 109 only needs to have at least a means for accessing the recording medium 110.
[0030] The control unit 111 controls each component of the imaging device 100 via a data bus 116 in accordance with input signals and a program described below. The control unit 111 has a CPU, ROM, and RAM for executing various controls. Note that instead of the control unit 111 controlling the entire imaging device 100, multiple pieces of hardware may share and control the entire imaging device. The ROM of the control unit 111 stores programs for controlling each component. The RAM of the control unit 111 is a volatile memory used for arithmetic processing and the like.
[0031] The operation unit 112 is a user interface for receiving instructions from the user for the imaging device 100. The operation unit 112 has, for example, a power switch 72 for turning the power of the imaging device 100 on or off, a release switch 61 for issuing an instruction to shoot, a playback button for issuing an instruction to play back image data or video data, a mode change switch 60, and the like.
[0032] The operation unit 112 outputs a control signal to the control unit 111 in response to a user operation. The operation unit 112 may also include a touch panel formed on the display unit 107. The release switch 61 has SW1 and SW2. When the release switch 61 is pressed halfway, SW1 is turned on. This accepts a preparation instruction for performing preparatory operations for image capture, such as AF (autofocus) processing, AE (auto exposure) processing, AWB (auto white balance) processing, and EF (pre-flash) processing. When the release switch 61 is pressed fully, SW2 is turned on. This accepts an image capture instruction for performing an image capture operation. The operation unit 112 also includes an operation member (e.g., a button) that can adjust the volume of audio data played from a speaker 114, which will be described later.
[0033] The audio output unit 113 can output audio data to the speaker 114 and the external output unit 115. The audio data input to the audio output unit 113 includes audio data read from the recording medium 110 by the recording control unit 109, audio data output from the nonvolatile memory 117, and audio data output from the encoding processing unit. The speaker 114 is an electro-acoustic transducer that can reproduce audio data.
[0034] The external output unit 115 can output image data, video data, audio data, etc. to an external device. The external output unit 115 is configured with, for example, a video terminal, a microphone terminal, a headphone terminal, etc.
[0035] The data bus 116 is a data bus for transmitting various data such as audio data, video data, and image data, as well as various control signals, to each block of the image capturing device 100 .
[0036] The nonvolatile memory 117 is a nonvolatile memory that stores programs, etc., which are executed by the control unit 111 and will be described later. Also, sound data is recorded in the nonvolatile memory 117. This sound data is, for example, sound data of electronic sounds such as a focusing sound that is output when a subject is focused, an electronic shutter sound that is output when an instruction to take a photograph is given, and an operation sound that is output when the imaging device 100 is operated.
[0037] <Operation of the imaging device 100> The operation of the imaging device 100 of this embodiment will now be described.
[0038] In the image capture device 100 of this embodiment, power is supplied to each component of the image capture device from a power supply (not shown) in response to a user turning on the power by operating the power switch 72. For example, the power supply is a battery such as a lithium ion battery or an alkaline manganese dry battery.
[0039] In response to the supply of power, the control unit 111 determines whether the camera will operate in a shooting mode or a playback mode, for example, based on the state of the mode selector switch 60. In the video recording mode, the control unit 111 records the video data output from the imaging unit 101 and the audio data output from the audio input unit 104 as one piece of video data with audio. In the playback mode, the control unit 111 reads out the image data or video data recorded on the recording medium 110 using the recording control unit 109, and controls the display unit 107 to display it.
[0040] First, the moving image recording mode will be described. In the moving image recording mode, the control unit 111 first transmits a control signal to each component of the imaging device 100 to transition the imaging device 100 to a shooting standby state. For example, the control unit 111 controls the imaging unit 101 and the audio input unit 104 to perform the following operations.
[0041] The imaging unit 101 converts an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and generates video data from the electrical signal generated by the imaging element. The imaging unit 101 then transmits the video data to the display control unit 106, which displays it on the display unit 107. The user can prepare for shooting while viewing the video data displayed on the display unit 107.
[0042] The audio input unit 104 performs A / D conversion on analog audio signals input from multiple microphones, respectively, to generate multiple digital audio signals. The audio input unit 104 then generates audio data for multiple channels from the multiple digital audio signals. The audio input unit 104 transmits the generated audio data to the audio output unit 113, which plays the audio data from the speaker 114. While listening to the audio data played from the speaker 114, the user can use the operation unit 112 to adjust the volume of the audio data to be recorded in the audio-accompanied video data.
[0043] Next, in response to the user pressing the LV button 76, the control unit 111 transmits an instruction signal to start shooting to each component of the image capturing device 100. For example, the control unit 111 controls the image capturing unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.
[0044] The imaging unit 101 converts an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and generates video data from the electrical signal generated by the imaging element. The imaging unit 101 then transmits the video data to a display control unit 106, which displays the video data on a display unit 107. The imaging unit 101 also transmits the generated video data to a volatile memory 105.
[0045] The audio input unit 104 performs A / D conversion on analog audio signals input from multiple microphones to generate multiple digital audio signals. The audio input unit 104 then generates multi-channel audio data from the multiple digital audio signals. The audio input unit 104 then transmits the generated audio data to the volatile memory 105.
[0046] The encoding processing unit 108 reads out and encodes the video data and audio data temporarily recorded in the volatile memory 105. The control unit 111 generates a data stream from the video data and audio data encoded by the control unit 111 and outputs it to the recording control unit 109. The recording control unit 109 records the input data stream as video data with audio on the recording medium 110 in accordance with a file system such as UDF or FAT.
[0047] The components of the image capturing apparatus 100 continue to perform the above operations during video capture.
[0048] Then, in response to the user pressing the LV button 76, the control unit 111 transmits an instruction signal to end shooting to each component of the image capturing device 100. For example, the control unit 111 controls the image capturing unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.
[0049] The image capturing unit 101 stops generating moving image data, and the audio input unit 104 stops generating audio data.
[0050] The encoding processing unit 108 reads and encodes the remaining video data and audio data recorded in the volatile memory 105. The control unit 111 generates a data stream from the video data and audio data encoded by the encoding processing unit 108 and outputs the data stream to the recording control unit 109.
[0051] The recording control unit 109 records the data stream as a file of audio-accompanying moving image data on the recording medium 110 in accordance with a file system such as UDF or FAT. Then, when the input of the data stream stops, the recording control unit 109 completes the audio-accompanying moving image data. Upon completion of the audio-accompanying moving image data, the recording operation of the imaging device 100 stops.
[0052] In response to the stop of the recording operation, the control unit 111 transmits a control signal to each component of the imaging device 100 to transition to a shooting standby state. As a result, the control unit 111 controls the imaging device 100 to return to the shooting standby state.
[0053] Next, the playback mode will be described. In the playback mode, the control unit 111 transmits a control signal to each component of the image capture device 100 to transition to a playback state. For example, the control unit 111 controls the encoding processing unit 108, the recording control unit 109, the display control unit 106, and the audio output unit 113 to perform the following operations.
[0054] The recording control unit 109 reads out the moving image data with audio recorded on the recording medium 110 and transmits the read moving image data with audio to the encoding processing unit 108 .
[0055] The encoding processing unit 108 decodes the audio-accompanying video data into image data and audio data. The encoding processing unit 108 transmits the decoded video data to the display control unit 106 and the decoded audio data to the audio output unit 113.
[0056] The display control unit 106 displays the decoded image data on the display unit 107. The audio output unit 113 reproduces the decoded audio data through the speaker 114.
[0057] As described above, the imaging device 100 of this embodiment can record and play back image data and audio data.
[0058] In this embodiment, the audio input unit 104 performs audio processing such as adjusting the level of an audio signal input from a microphone. In this embodiment, the audio input unit 104 performs this audio processing in response to the start of video recording. Note that this audio processing may be performed after the imaging device 100 is turned on. This audio processing may also be performed in response to the selection of a shooting mode. This audio processing may also be performed in response to the selection of a mode related to audio recording, such as a video recording mode or an audio memo function. This audio processing may also be performed in response to the start of audio signal recording.
[0059] <Configuration of the voice input unit 104> FIG. 3 is a block diagram showing an example of a detailed configuration of the voice input unit 104 in this embodiment.
[0060] In this embodiment, the audio input unit 104 has three microphones: an L microphone 201a, an R microphone 201b, and a noise microphone 201c. The L microphone 201a and the R microphone 201b are each an example of a first microphone. In this embodiment, the image capture device 100 collects environmental sounds using the L microphone 201a and the R microphone 201b, and records the audio signals input from the L microphone 201a and the R microphone 201b in stereo. For example, environmental sounds are sounds generated outside the housing of the image capture device 100 and outside the housing of the optical lens 300, such as the user's voice, animal cries, the sound of rain, and music.
[0061] The noise microphone 201c is an example of a second microphone. The noise microphone 201c is a microphone for acquiring noise, such as driving sounds from a predetermined noise source, generated within the housing of the image capture device 100 and the housing of the optical lens 300. Examples of noise sources include driving components such as ultrasonic motors (USMs) and stepper motors (STMs). The noise is, for example, vibration noise generated by the driving of motors such as USMs and STMs. For example, motors are driven during AF processing to focus on a subject. The image capture device 100 acquires noise, such as driving sounds, generated within the housing of the image capture device 100 and the housing of the optical lens 300 using the noise microphone 201c, and generates noise parameters (described later) using audio data of the acquired noise. In this embodiment, the left microphone 201a, the right microphone 201b, and the noise microphone 201c are omnidirectional microphones. An example of the arrangement of the L microphone 201a, the R microphone 201b, and the noise microphone 201c in this embodiment will be described later with reference to FIG.
[0062] The L microphone 201a, R microphone 201b, and noise microphone 201c each generate an analog audio signal from the captured audio and input it to the amplifier 202. Here, the audio signal input from the L microphone 201a is referred to as Lch, the audio signal input from the R microphone 201b as Rch, and the audio signal input from the noise microphone 201c as Nch.
[0063] The amplifier 202 amplifies the amplitude of the analog audio signals input from the L microphone 201a, R microphone 201b, and noise microphone 201c. In this embodiment, the gain (hereinafter, amplification amount A) for the analog audio signals input from the L microphone 201a, R microphone 201b, and noise microphone 201c is a fixed value (predetermined amount). The gain (hereinafter, amplification amount B) for the analog audio signal input from the noise microphone 201c is changed as appropriate in accordance with a signal from the level detector 204 or the lens controller 102. Details of how the gain for the analog audio signal input from the noise microphone 201c is changed will be described later using FIGS. 5 to 8.
[0064] The A / D conversion unit 203 converts the analog audio signal amplified by the amplification unit 202 into a digital audio signal. The A / D conversion unit 203 outputs the converted digital audio signal to the FFT unit 205. In this embodiment, the A / D conversion unit 203 converts the analog audio signal into a digital audio signal by performing sampling processing with a sampling frequency of 48 kHz and a bit depth of 16 bits.
[0065] The FFT unit 205 performs fast Fourier transform processing on the time-domain digital audio signal input from the A / D conversion unit 203 to convert it into a frequency-domain digital audio signal. In this embodiment, the frequency-domain digital audio signal has a frequency spectrum of 1024 points in a frequency band from 0 Hz to 48 kHz. The frequency-domain digital audio signal also has a frequency spectrum of 513 points in a frequency band from 0 Hz to 24 kHz, which is the Nyquist frequency. In this embodiment, the imaging device 100 performs noise reduction processing using the 513-point frequency spectrum from 0 Hz to 24 kHz of the audio data output from the FFT unit 205.
[0066] Here, the frequency spectrum of the Lch subjected to the fast Fourier transform is represented by 513-point array data of Lch_Before[0] to Lch_Before
[0512] . When these array data are referred to collectively, they are referred to as Lch_Before. Furthermore, the frequency spectrum of the Rch subjected to the fast Fourier transform is represented by 513-point array data of Rch_Before[0] to Rch_Before
[0512] . When these array data are referred to collectively, they are referred to as Rch_Before. Note that Lch_Before and Rch_Before are each an example of first frequency spectrum data.
[0067] The Nch frequency spectrum after the fast Fourier transform is represented by array data of 513 points, Nch_Before[0] to Nch_Before
[0512] . These array data are collectively referred to as Nch_Before. Nch_Before is an example of second frequency spectrum data.
[0068] The noise data generator 206 generates data for reducing noise contained in Lch_Before and Rch_Before based on Nch_Before. In this embodiment, the noise data generator 206 uses noise parameters to generate array data NL[0] to NL
[0512] for reducing noise contained in Lch_Before[0] to Lch_Before
[0512] , respectively. The noise data generator 206 also generates array data NR[0] to NR
[0512] for reducing noise contained in Rch_Before[0] to Rch_Before
[0512] , respectively. The frequency points in the array data NL[0] to NL
[0512] are the same as the frequency points in the array data Lch_Before[0] to Lch_Before
[0512] . Furthermore, the frequency points in the array data of NR[0] to NR
[0512] are the same as the frequency points in the array data of Rch_Before[0] to Rch_Before
[0512] .
[0069] The sequence data NL[0] to NL
[0512] are collectively referred to as NL. The sequence data NR[0] to NR
[0512] are collectively referred to as NR. NL and NR are each an example of third frequency spectrum data.
[0070] Furthermore, when the amplification amount A and the amplification amount B are different, the noise data generation unit 206 corrects the noise parameters generated by the noise data generation unit 206 based on the amplification amount A and the amplification amount B. In this embodiment, the noise data generation unit 206 corrects the noise parameters generated by the noise data generation unit 206 based on the difference between the amplification amount A and the amplification amount B.
[0071] The noise parameter recording unit 207 records noise parameters used by the noise data generating unit 206 to generate NL and NR from Nch_Before. The noise parameter recording unit 207 records multiple types of noise parameters according to the type of noise. The noise parameters used to generate NL from Nch_Before are collectively referred to as PLx. The noise parameters used to generate NR from Nch_Before are collectively referred to as PRx.
[0072] PLx and PRx have the same number of sequences as NL and NR, respectively. For example, PL1 is sequence data from PL1[0] to PL1
[0512] . The frequency points of PL1 are the same as the frequency points of Lch_Before. For example, PR1 is sequence data from PR1[0] to PR1
[0512] . The frequency points of PR1 are the same as the frequency points of Rch_Before. The noise parameters will be described later using Figure 5.
[0073] The noise microphone 201c determines the noise parameters to be used in the noise data generator 206 from the noise parameters recorded in the noise parameter recording unit 207. In this embodiment, the noise parameter recording unit 207 records all of the coefficients for each of the 513 frequency spectrum points as noise parameters. However, it is sufficient that the coefficients for at least the frequency points necessary to reduce noise are recorded, rather than the coefficients for all 513 frequency points. For example, the noise parameter recording unit 207 may record, as noise parameters, coefficients for each frequency spectrum from 20 Hz to 20 kHz, which are considered to be typical audible frequencies, but may not record coefficients for other frequency spectrums. Furthermore, for example, coefficients for frequency spectrums whose coefficient value is zero may not be recorded in the noise parameter recording unit 207 as noise parameters.
[0074] The subtraction processing unit 208 subtracts NL and NR from Lch_Before and Rch_Before, respectively. For example, the subtraction processing unit 208 has an L subtractor 208a that subtracts NL from Lch_Before, and an R subtractor 208b that subtracts NR from Rch_Before. The L subtractor 208a subtracts NL from Lch_Before and outputs 513-point array data of Lch_After[0] to Lch_After
[0512] . The R subtractor 208b subtracts NR from Rch_Before and outputs 513-point array data of Rch_After[0] to Rch_After
[0512] . In this embodiment, the subtraction processing unit 208 performs subtraction processing using a spectral subtraction method.
[0075] The iFFT unit 209 performs an inverse fast Fourier transform (inverse Fourier transform) on the frequency domain digital audio signal input from the subtraction processing unit 208 to convert it into a time domain digital audio signal.
[0076] The audio processing unit 210 performs audio processing on the time domain digital audio signal, such as an equalizer, an auto level controller, and stereo enhancement processing, etc. The audio processing unit 210 outputs the processed audio data to the volatile memory 105.
[0077] In this embodiment, the imaging device 100 has two microphones as the first microphone, but the imaging device 100 may have one microphone or three or more microphones as the first microphone. For example, when the imaging device 100 has one microphone as the first microphone in the audio input unit 104, the imaging device 100 records audio data picked up by the one microphone in monaural format. Also, when the imaging device 100 has three or more microphones as the first microphone in the audio input unit 104, the imaging device 100 records audio data picked up by the three or more microphones in surround format.
[0078] In this embodiment, the L microphone 201a, the R microphone 201b, and the noise microphone 201c are non-directional microphones, but these microphones may also be directional microphones.
[0079] <Arrangement of microphones in the audio input unit 104> Here, an example of the arrangement of the microphones of the voice input unit 104 of this embodiment will be described. Fig. 4 shows an example of the arrangement of the L microphone 201a, the R microphone 201b, and the noise microphone 201c.
[0080] 4 is an example of a cross-sectional view of a portion of the imaging device 100 to which the L microphone 201a, the R microphone 201b, and the noise microphone 201c are attached. This portion of the imaging device 100 is composed of an exterior part 302, a microphone bushing 303, and a fixing part 304.
[0081] The exterior part 302 has holes (hereinafter referred to as microphone holes) for inputting environmental sounds into the microphones. In this embodiment, the microphone holes are formed above the L microphone 201a and the R microphone 201b. On the other hand, the noise microphone 201c is provided to acquire drive sounds generated within the housing of the image capture device 100 and the housing of the optical lens 300, and does not need to acquire environmental sounds. Therefore, in this embodiment, no microphone holes are formed in the exterior part 302 above the noise microphone 201c.
[0082] Drive sounds generated within the housings of the image capture device 100 and the optical lens 300 are picked up by the L microphone 201a and the R microphone 201b through the microphone holes. If drive sounds or the like are generated within the housings of the image capture device 100 and the optical lens 300 when ambient noise is low, the sound picked up by each microphone will mainly be this drive sound. Therefore, the sound level from the noise microphone 201c is higher than the sound levels from the L microphone 201a and the R microphone 201b. In other words, in this case, the relationship between the levels of the audio signals output from each microphone is as follows: Lch≒Rch <Nch Furthermore, when the environmental sound becomes louder, the audio level of the environmental sound from the L microphone 201a and R microphone 201b becomes louder than the audio level of the drive sound generated by the image capture device 100 or the optical lens 300 and output from the noise microphone 201c. Therefore, in this case, the relationship between the levels of the audio signals output from each microphone is as follows: Lch ≒ Rch > Nch In this embodiment, the shape of the microphone hole formed in exterior part 302 is elliptical, but it may be other shapes such as circular or rectangular. Furthermore, the shape of the microphone hole on microphone 201a and the shape of the microphone hole on microphone 201b may be different from each other.
[0083] In this embodiment, the noise microphone 201c is placed close to the L microphone 201a and the R microphone 201b. In addition, in this embodiment, the noise microphone 201c is placed between the L microphone 201a and the R microphone 201b so that it is approximately equidistant from each microphone. As a result, the audio signal generated by the noise microphone 201c from drive sounds and the like generated inside the housing of the image capture device 100 and the housing of the optical lens 300 becomes similar to the audio signals generated by the L microphone 201a and the R microphone 201b from this drive sounds and the like.
[0084] The microphone bushing 303 is a member for fixing the L microphone 201a, the R microphone 201b, and the noise microphone 201c. The fixing portion 304 is a member for fixing the microphone bushing 303 to the exterior portion 302.
[0085] In this embodiment, exterior part 302 and fixed part 304 are made of a molded material such as PC material. Also, exterior part 302 and fixed part 304 may be made of a metal material such as aluminum or stainless steel. Also, in this embodiment, microphone bushing 303 is made of a rubber material such as ethylene propylene diene rubber.
[0086] <Means for setting the amount of amplification> FIG. 5 is an example of a block diagram of the amplifier unit 202 in this embodiment.
[0087] The amplifier 202 is made up of an environmental sound amplifier 2021, a noise amplifier 2022, and an amplification amount storage unit 2023 that stores amplification amount update data.
[0088] The environmental sound amplifier 2021 amplifies the audio signals input from the L microphone 201a and the R microphone 201b. Here, the gain in the environmental sound amplifier 2021 is an amplification amount A.
[0089] The noise amplifier 2022 amplifies the audio signal input from the noise microphone 201c. Here, the gain in the noise amplifier 2022 is an amplification amount B. The noise amplifier 2022 reduces the amplification amount of the audio signal input from the noise microphone 201c in accordance with the detection content detected by the level detector 204 and the lens controller 102.
[0090] The amplification amount storage unit 2023 is a memory that stores amplification amount update data.
[0091] The level detection unit 204 determines the type of the lens attached to the imaging device 100 based on the sound pressure of the audio signal converted by the A / D conversion unit 203 .
[0092] <Amplification amount change processing using a level detection unit> 6 is a flowchart showing an example of the amplification amount change process of the amplifier 202 when using the level detector 204. The process of this flowchart starts when an instruction to start moving image recording is received from the user via the operation unit 112. For example, the control unit 111 starts moving image recording in response to detecting that the release switch 61 has been pressed.
[0093] In step S601, the control unit 111 starts imaging processing by the imaging unit 101 and audio processing by the audio input unit 104. The video obtained by the imaging processing and the audio obtained by the audio processing are sequentially recorded on the recording medium 110.
[0094] In step S602, the voice input unit 104 sets the amplification amount B to the gain stored in the amplification amount storage unit 2023. In this embodiment, at the start of this flowchart, a predetermined value (initial value) of gain is stored in the amplification amount storage unit 2023. In this embodiment, this initial value is equal to the amplification amount A.
[0095] In step S603, the audio input unit 104 determines whether the optical lens 300 is being driven. For example, the audio input unit 104 determines whether the motor of the optical lens 300 is being driven based on a signal input from the lens control unit 102. When the motor of the optical lens 300 is being driven, the signal input from the lens control unit 102 includes, for example, a signal indicating AF, zoom, etc. If it is determined that the optical lens 300 is being driven, the process of step S603 is executed. If it is determined that the optical lens 300 is not being driven, the process of step S608 is executed.
[0096] In step S604, the level detection unit 204 determines whether the amplitude (sound pressure level) of the audio signals input from the L microphone 201a and the R microphone 201b is equal to or greater than a predetermined threshold. If it is determined that the amplitude of the audio signals input from the L microphone 201a and the R microphone 201b is equal to or greater than the predetermined threshold, the process of step S608 is executed. If it is determined that the amplitude of the audio signals input from the L microphone 201a and the R microphone 201b is less than the predetermined threshold, the process of step S605 is executed.
[0097] In step S605, the level detection unit 204 detects the amplitude of the audio signal input from the noise microphone 201c.
[0098] In step S606, the level detection unit 204 determines whether the amplitude of the audio signal detected in step S605 is equal to or greater than a predetermined threshold. If the amplitude of the audio signal detected in step S605 is equal to or greater than the predetermined threshold, the process proceeds to step S607. If the amplitude of the audio signal detected in step S605 is less than the predetermined threshold, the process proceeds to step S608.
[0099] In step S607, the level detection unit 204 calculates the amplification amount B based on the amplitude of the audio signal detected in step S605, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B when the optical lens 300 is not driven is set to the maximum value. Therefore, the amplification amount B calculated in this step is a value smaller than the initial value.
[0100] In step S608, the control unit 111 determines whether to end the moving image recording. For example, when the user presses the release switch 61 or when the remaining capacity of the recording medium 110 is low, the control unit 111 determines to end the moving image recording. If it is determined to end the moving image recording, the processing of this flowchart ends. If it is determined not to end the moving image recording, the processing returns to step S603.
[0101] The method for changing the amount of amplification using the level detection section 204 has been described above.
[0102] <Amplification amount change processing using lens control unit> 7 is a flowchart showing an example of the amplification amount change process of the amplifier unit 202 when using the lens control unit 102. The process of this flowchart starts when an instruction to start video recording is received from the user via the operation unit 112. For example, the control unit 111 starts video recording in response to detecting that the release switch 61 has been pressed.
[0103] In step S701, the lens control unit 102 determines the type of optical lens 300 attached to the imaging device 100. Here, the lens control unit 102 determines whether the determined optical lens 300 is a lens with high drive noise. If the determined optical lens 300 is a lens with high drive noise, the process of step S702 is executed. If the determined optical lens 300 is not a lens with high drive noise, the process of step S703 is executed.
[0104] In step S702, the voice input unit 104 calculates the amplification amount B based on the type of optical lens 300 input from the lens control unit 102, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B for the optical lens 300 with small driving noise is set to the maximum value. Therefore, the amplification amount B calculated in this step is a value smaller than the maximum value.
[0105] In step S703, the level detection unit 204 calculates the amplification amount B based on the type of optical lens detected in step S701, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B is set to a maximum value when an optical lens with low drive noise is attached to the image capture device 100.
[0106] In step S704, the control unit 111 starts imaging processing by the imaging unit 101 and audio processing by the audio input unit 104. The video obtained by the imaging processing and the audio obtained by the audio processing are sequentially recorded on the recording medium 110.
[0107] In step S705, the voice input unit 104 sets the amplification amount B to the gain stored in the amplification amount storage unit 2023.
[0108] In step S706, the control unit 111 determines whether to end the moving image recording. For example, when the user presses the release switch 61 or when the remaining capacity of the recording medium 110 is low, the control unit 111 determines to end the moving image recording. If it is determined to end the moving image recording, the processing of this flowchart ends. If it is determined not to end the moving image recording, the processing returns to step S705.
[0109] The method for changing the amplification amount using the lens control unit 102 has been described above.
[0110] <Amplification amount change processing using level detection unit and lens control unit> 8 is a flowchart showing an example of the amplification amount change process of the amplifier unit 202 when both the level detection unit 204 and the lens control unit 102 are used. The process of this flowchart starts when an instruction to start moving image recording is received from the user via the operation unit 112. For example, the control unit 111 starts moving image recording in response to detecting that the release switch 61 has been pressed.
[0111] In step S801, the lens control unit 102 determines the type of optical lens 300 attached to the imaging device 100. Here, the lens control unit 102 determines whether the determined optical lens 300 is a lens with high drive noise. If the determined optical lens 300 is a lens with high drive noise, the process of step S802 is executed. If the determined optical lens 300 is not a lens with high drive noise, the process of step S803 is executed.
[0112] In step S802, the voice input unit 104 calculates the amplification amount B based on the type of optical lens 300 input from the lens control unit 102, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B for the optical lens 300 with small driving noise is set to the maximum value. Therefore, the amplification amount B calculated in this step is a value smaller than the maximum value.
[0113] In step S803, the level detection unit 204 calculates the amplification amount B based on the type of optical lens detected in step S801, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B is set to a maximum value when an optical lens with low drive noise is attached to the image capture device 100.
[0114] In step S804, the control unit 111 starts imaging processing by the imaging unit 101 and audio processing by the audio input unit 104. The video obtained by the imaging processing and the audio obtained by the audio processing are sequentially recorded on the recording medium 110.
[0115] In step S805, the voice input unit 104 sets the amplification amount B to the gain stored in the amplification amount storage unit 2023.
[0116] In step S806, the audio input unit 104 determines whether the optical lens 300 is being driven. For example, the audio input unit 104 determines whether the motor of the optical lens 300 is being driven based on a signal input from the lens control unit 102. When the motor of the optical lens 300 is being driven, the signal input from the lens control unit 102 includes, for example, a signal indicating AF, zoom, etc. If it is determined that the optical lens 300 is being driven, the process of step S807 is executed. If it is determined that the optical lens 300 is not being driven, the process of step S811 is executed.
[0117] In step S807, the level detection unit 204 determines whether the amplitude (sound pressure level) of the audio signals input from the L microphone 201a and the R microphone 201b is equal to or greater than a predetermined threshold. If it is determined that the amplitude of the audio signals input from the L microphone 201a and the R microphone 201b is equal to or greater than the predetermined threshold, the process of step S811 is executed. If it is determined that the amplitude of the audio signals input from the L microphone 201a and the R microphone 201b is less than the predetermined threshold, the process of step S808 is executed.
[0118] In step S808, the level detection unit 204 detects the amplitude of the audio signal input from the noise microphone 201c.
[0119] In step S809, the level detection unit 204 determines whether the amplitude of the audio signal detected in step S808 is equal to or greater than a predetermined threshold. If the amplitude of the audio signal detected in step S808 is equal to or greater than the predetermined threshold, the process proceeds to step S810. If the amplitude of the audio signal detected in step S808 is less than the predetermined threshold, the process proceeds to step S811.
[0120] In step S810, the level detection unit 204 calculates the amplification amount B based on the amplitude of the audio signal detected in step S808, and records the calculated amplification amount B in the amplification amount storage unit 2023. In this embodiment, the amplification amount B set in step S803 is set to the maximum value. Therefore, the amplification amount B calculated in this step is a value equal to or less than the amplification amount set according to the optical lens 300.
[0121] The method for changing the amount of amplification using the level detection unit 204 and the lens control unit 102 has been described above.
[0122] In this way, the image capture device 100 can appropriately control the amount of amplification for input noise by setting the amount of amplification for noise according to the detected noise level or the optical lens, thereby enabling the image capture device 100 to effectively reduce noise.
[0123] The image capturing apparatus 100 may appropriately determine whether to use the level detection unit 204, the lens control unit 102, or both for the amplification amount change process.
[0124] <Noise parameters> 9 shows an example of noise parameters recorded in the noise parameter recording unit 207. The noise parameters are parameters for correcting audio signals generated by the noise microphone 201c capturing drive sounds generated within the housing of the image capture device 100 and the housing of the optical lens 300. As shown in FIG. 5, in this embodiment, PLx and PRx are recorded in the noise parameter recording unit 207. In this embodiment, the source of the drive sounds will be described as being within the housing of the optical lens 300. The drive sounds generated within the housing of the optical lens 300 are transmitted into the housing of the image capture device 100 via the lens mount 301 and captured by the L microphone 201a, R microphone 201b, and noise microphone 201c.
[0125] The frequency of the drive sound varies depending on the type of drive sound. Therefore, in this embodiment, the image capture device 100 stores multiple noise parameters corresponding to the type of drive sound (noise). Then, noise data is generated using one of these multiple noise parameters. In this embodiment, the image capture device 100 records noise parameters for white noise as constant noise. The image capture device 100 also reduces noise other than constant noise. For example, the image capture device 100 records noise parameters for short-term noise generated by the meshing of gears inside the optical lens 300. Furthermore, for example, the image capture device 100 stores noise parameters for sliding noise inside the housing of the lens 300 as long-term noise.
[0126] Additionally, as quantization noise, noise parameters for noise that occurs when the A / D conversion unit 203 performs A / D conversion on audio data are recorded.
[0127] The image capturing device 100 may record noise parameters for each type of optical lens 300 , and for each temperature inside the housing of the image capturing device 100 and each tilt of the image capturing device 100 detected by the information acquiring unit 103 .
[0128] <How to generate noise data> 10 and 11, the noise data generation process in the noise data generator 206 will be described. Here, the noise data generation process for the Lch data will be described, but the noise data generation method for the Rch data is similar.
[0129] First, a process for generating noise parameters in a situation where it is considered that there is no environmental sound will be described. Fig. 10(a) is an example of the frequency spectrum of Lch_Before when drive sound occurs within the housing of the optical lens 300 in a situation where it is considered that there is no environmental sound. Fig. 10(b) is an example of the frequency spectrum of Nch_Before when drive sound occurs within the housing of the optical lens 300 in a situation where it is considered that there is no environmental sound. The horizontal axis indicates the frequency from the 0th point to the 512th point, and the vertical axis indicates the amplitude of the frequency spectrum.
[0130] Since the situation is such that it can be assumed that there is no environmental sound, the amplitude of the frequency spectrum in the same frequency band is large in Lch_Before and Nch_Before. In addition, because drive sound is generated inside the housing of optical lens 300, the amplitude of each frequency spectrum for the same drive sound tends to be larger in Nch_Before than in Lch_Before.
[0131] FIG. 10(c) shows an example of PLx in this embodiment. In this embodiment, PLx is a coefficient of each frequency spectrum calculated by dividing the amplitude of each frequency spectrum of Lch_Before by the amplitude of each frequency spectrum of Nch_Before. The result of this division is expressed as Lch_Before / Nch_Before. In other words, PLx is the ratio of the amplitudes of Lch_Before and Nch_Before. The noise parameter recording unit 207 records the value of Lch_Before / Nch_Before as the noise parameter PLx. As described above, the amplitude of the frequency spectrum for the same driving sound tends to be larger in Nch_Before than in Lch_Before, so the value of each coefficient of the noise parameter PLx tends to be smaller than 1. However, if the value of Nch_Before[n] is smaller than a predetermined threshold, the noise parameter recording unit 207 records the noise parameter PLx as PLx[n]=0.
[0132] Next, the process of applying the generated noise parameters to Nch_Before will be described. Fig. 11(a) is an example of the frequency spectrum of Lch_Before when drive sound is generated within the housing of the optical lens 300 in a situation where environmental sound is present. Fig. 11(b) is an example of the frequency spectrum of Nch_Before when drive sound is generated within the housing of the optical lens 300 in a situation where environmental sound is present. The horizontal axis represents the frequency from point 0 to point 512, and the vertical axis represents the amplitude of the frequency spectrum.
[0133] 11(c) shows an example of NL when drive sound occurs inside the housing of the optical lens 300 in the presence of environmental sound. The noise data generator 206 multiplies each frequency spectrum of Nch_Before by each coefficient of PLx to generate NL. NL is the frequency spectrum generated in this way.
[0134] 11(d) shows an example of Lch_After when drive sound occurs inside the housing of optical lens 300 in the presence of environmental sound. Subtraction processing unit 208 subtracts NL from Lch_Before to generate Lch_After. Lch_After is the frequency spectrum generated in this way.
[0135] This allows the imaging device 100 to reduce noise caused by driving noise within the housing of the optical lens 300, and record environmental sounds with less noise.
[0136] Here, we will explain the noise data generation process in the noise data generation unit 206 when the amplifier unit 202 amplifies the audio signals input from the L microphone 201a and R microphone 201b and the audio signal input from the noise microphone 201c by different amplification amounts.
[0137] Here, the process of generating noise data for the Lch data will be described, but the method of generating noise data for the Rch data is similar.
[0138] First, a method for creating noise data when the noise amplifier 2022 is updated to the amplification amount stored in the amplification amount storage unit 2023 and a difference occurs between the amplification amount of the noise amplifier 2022 and the amplification amount of the environmental sound amplifier 2021 will be described.
[0139] 12 is a diagram illustrating the details of the noise data generation unit 206. The noise data generation unit 206 includes an amplification amount comparison unit 2061, a PL subtractor 206a, a PR subtractor 206b, and a generation unit 2062. The amplification amount comparison unit 2061 subtracts the amplification amount of the environmental sound amplification unit 2021 from the amplification amount of the noise amplification unit 2022, and uses the PL subtractor 206a to subtract this difference from PLx. For example, if the amplification amount of the noise amplification unit 2022 is set to 0.5 times the initial value, the PL subtractor 206a performs subtraction processing to multiply PLx by 2.0 (= 1 / 0.5). Through this processing, the NL generated by the generation unit 2062 becomes a parameter in which the difference between the amplification amount of the noise amplification unit 2022 and the amplification amount of the environmental sound amplification unit 2021 has been corrected.
[0140] Next, a description will be given of a noise generation method when quantization noise of the A / D conversion unit 203 is present in Nch_Before. Fig. 13 is a diagram illustrating the details of the noise data generation unit 206. The noise data generation unit 206 has a quantization noise comparison unit 2063 and a generation unit 2062.
[0141] In this case, the noise data generator 206 generates noise data using a quantization noise parameter. Here, the quantization noise parameter is, for example, a unique noise spectrum that occurs when an audio signal is converted by the A / D converter 203, and is recorded in the noise parameter recorder 207. In this embodiment, the quantization noise parameter is shown as PN1 in Fig. 9.
[0142] 14A to 14E are diagrams showing details of a method for comparing Nch_Before with a quantization noise parameter. All of the frequency spectra in Fig. 14A to Fig. 14E are processed on the same time axis.
[0143] 14(a) is an example of the frequency spectrum of Nch_Before when drive noise occurs inside the housing of the optical lens 300. The horizontal axis indicates the frequency from point 0 to point 512, and the vertical axis indicates the amplitude of the frequency spectrum.
[0144] FIG. 14(b) shows an example of the frequency spectrum of the quantization noise parameter PN1 that occurs when the audio signal is converted by the A / D converter 203.
[0145] FIG. 14(c) is an example of the frequency spectrum Noise_Lens of the driving sound actually generated inside the housing of the optical lens 300.
[0146] 14(d) is a diagram showing how much of the frequency spectrum of the drive sound actually generated inside the housing of the optical lens 300 is included in the frequency spectrum of Nch_Before, which is subjected to subtraction processing by the subtraction processing unit 208. That is, among Nch_Before, samples that remain grayed out are samples that have been affected by quantization noise, and samples that do not remain grayed out are samples that have captured the drive sound actually generated inside the housing of the optical lens 300.
[0147] The quantization noise comparison unit 2063 sets the noise parameter of a sample affected by quantization noise to 0. For example, the quantization noise comparison unit 2063 compares the amplitude of Nch_Before with the amplitude of the quantization noise recorded in the noise parameter recording unit 207 for each sample, and if the difference in amplitude is equal to or less than a certain threshold, sets the noise parameter of that sample to 0.
[0148] FIG. 14(e) shows the NL generated by the generator 2062 when the noise parameter in the sample affected by the quantization noise is set to 0.
[0149] When the audio data acquired by the noise microphone 201c is affected by the quantization noise of the A / D conversion unit 203, this processing makes it possible to prevent the subtraction processing unit 208 from performing subtraction processing on the band affected by the quantization noise.
[0150] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0151] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.
Claims
1. a first microphone for capturing environmental sounds; a second microphone for acquiring sound from a noise source that is a moving member included in an optical lens; a first amplifier for amplifying an audio signal input from the first microphone; a second amplifier for amplifying the audio signal input from the second microphone according to an amplification amount; a first conversion means for Fourier transforming the audio signal input from the first amplification means to generate a first audio signal; a second conversion means for Fourier transforming the audio signal input from the second amplification means to generate a second audio signal; a generating means for generating noise data using the second audio signal and parameters related to the noise of the noise source; a reduction unit that reduces noise corresponding to the noise source from the first audio signal based on the noise data; a third transforming means for performing an inverse Fourier transform on the audio signal output from the reducing means, the second amplifier sets the amplification amount based on the type of the optical lens or the level of the audio signal output from the second amplifier, When setting the amplification amount based on the type of the optical lens, the second amplification means sets the amplification amount to a first value when the optical lens is a first optical lens, and sets the amplification amount to a second value smaller than the first value when the optical lens is a second optical lens that generates noise greater than that of the first optical lens.
1. A voice processing device comprising:
2. The audio processing device described in Claim 1, characterized in that when the second amplification means sets the amplification amount based on the type of noise source and the level of the audio signal output from the second amplification means, the second amplification means changes the amplification amount from the first value to a value smaller than the first value depending on whether the level of the audio signal amplified according to the first value and output by the second amplification means is above a threshold value.
3. A first microphone for acquiring environmental sounds; a second microphone for acquiring sound from a noise source that is a moving member included in an optical lens; a first amplifier for amplifying an audio signal input from the first microphone; a second amplifier for amplifying the audio signal input from the second microphone according to an amplification amount; a first conversion means for Fourier transforming the audio signal input from the first amplification means to generate a first audio signal; a second conversion means for Fourier transforming the audio signal input from the second amplification means to generate a second audio signal; a generating means for generating noise data using the second audio signal and parameters related to the noise of the noise source; a reduction unit that reduces noise corresponding to the noise source from the first audio signal based on the noise data; a third transforming means for performing an inverse Fourier transform on the audio signal output from the reducing means, the second amplifier sets the amplification amount based on the type of the optical lens or the level of the audio signal output from the second amplifier, When the second amplifier sets the amplification amount based on the level of the audio signal output from the second amplifier, the second amplifier reduces the amplification amount in response to the level of the audio signal output from the first amplifier not being equal to or higher than a predetermined level and the level of the audio signal output from the second amplifier being equal to or higher than a threshold.
1. A voice processing device comprising:
4. 4. The audio processing device according to claim 1, further comprising an imaging unit for imaging a subject through the optical lens.
5. 5. The audio processing device according to claim 1, wherein the generating means corrects the parameter in accordance with the amount of amplification of the second amplifying means.
6. a first AD conversion means for converting the audio signal output from the first amplification means into a digital signal and outputting the digital signal to the first conversion means; 6. The audio processing device according to claim 1, further comprising: second AD conversion means for converting the audio signal output from the second amplification means into a digital signal and outputting the digital signal to the second conversion means.
7. a first microphone for capturing environmental sounds; a second microphone for acquiring sound from a noise source that is a driving member included in an optical lens, a first amplifying step of amplifying the audio signal input from the first microphone; a second amplification step of amplifying the audio signal input from the second microphone according to an amplification amount; a first transforming step of Fourier transforming the audio signal amplified in the first amplifying step to generate a first audio signal; a second transforming step of Fourier transforming the audio signal amplified in the second amplifying step to generate a second audio signal; a generating step of generating noise data using the second audio signal and parameters related to noise from the noise source; a reduction step of reducing noise corresponding to the noise source from the first audio signal based on the noise data; a third transform step of performing an inverse Fourier transform on the audio signal output from the reduction step, In the second amplifying step, the amplification amount is set based on the type of the optical lens or the level of the audio signal output from the second amplifying step; In the second amplification step, when the amplification amount is set based on the type of the optical lens, the amplification amount is set to a first value when the optical lens is a first optical lens, and the amplification amount is set to a second value smaller than the first value when the optical lens is a second optical lens that generates noise larger than that of the first optical lens. A control method comprising:
8. A first microphone for acquiring environmental sounds; a second microphone for acquiring sound from a noise source that is a driving member included in an optical lens, a first amplifying step of amplifying the audio signal input from the first microphone; a second amplification step of amplifying the audio signal input from the second microphone according to an amplification amount; a first transforming step of Fourier transforming the audio signal amplified in the first amplifying step to generate a first audio signal; a second transforming step of Fourier transforming the audio signal amplified in the second amplifying step to generate a second audio signal; a generating step of generating noise data using the second audio signal and parameters related to noise from the noise source; a reduction step of reducing noise corresponding to the noise source from the first audio signal based on the noise data; a third transform step of performing an inverse Fourier transform on the audio signal output from the reduction step, In the second amplifying step, the amplification amount is set based on the type of the optical lens or the level of the audio signal output from the second amplifying step; In the second amplification step, when the amplification amount is set based on the level of the audio signal output from the second amplification step, the amplification amount is reduced in response to the level of the audio signal output from the first amplification step being less than a predetermined level and the level of the audio signal output from the second amplification step being greater than a threshold. A control method comprising:
Citation Information
Patent Citations
Volume control method, audio signal reproduction device, and program
JP2010141571A
Sound recording apparatus and method, and imaging apparatus
JP2011028061A
Voice processing apparatus and electronic camera
JP2011114465A
Voice signal processor
JP2011124959A
Noise reduction device
JP2014044313A