Audio processing device, control method, and program
The audio processing device uses dual microphones and advanced signal processing to detect and reduce short-term noise from lens drive sounds, improving audio quality in digital cameras.
Patent Information
- Application Number
- JP2021087690
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-25
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-05-25
AI Technical Summary
Conventional digital cameras struggle to effectively reduce short-term noise, such as sliding sounds generated by the optical lens drive, due to inaccurate noise pattern detection from noise collected inside the lens housing.
The audio processing device employs two microphones - one for environmental sounds and another for noise sources like the lens drive unit - to detect and reduce short-term noise through Fourier transforms, volume control, and inverse Fourier transforms, specifically targeting noise patterns from the lens drive unit.
This approach effectively reduces short-term noise, enhancing audio quality during video recording by minimizing lens drive-related noise interference.
Smart Images

Figure 0007725236000001 
Figure 0007725236000002 
Figure 0007725236000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio processing device capable of reducing noise contained in audio data. [Background technology]
[0002] A digital camera, which is an example of an audio processing device, can record surrounding audio when recording video data. Digital cameras also have an autofocus function that drives an optical lens to focus on a subject while recording video data. Digital cameras also have a zoom function that drives the optical lens while recording video.
[0003] In this way, when an optical lens is driven while recording a video, the driving sound of the optical lens may be included as noise in the audio recorded along with the video. Therefore, conventional digital cameras have been able to reduce noise, such as sliding sounds generated when the optical lens is driven, and record surrounding audio. Patent Document 1 discloses a digital camera that reduces noise using a spectral subtraction method. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-205527 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in Patent Document 1, the digital camera creates a noise pattern from noise collected by a microphone that records surrounding sounds, so there is a possibility that an accurate noise pattern cannot be obtained from the sliding sounds generated inside the optical lens housing. In this case, there is a risk that the digital camera will not be able to effectively reduce noise contained in the collected sounds, particularly short-term noise that occurs when the drive unit is intermittently driven or when gears collide.
[0006] Therefore, an object of the present invention is to effectively reduce short-term noise. [Means for solving the problem]
[0007] The audio processing device of the present invention includes a first microphone for acquiring environmental sounds, a second microphone for acquiring sounds from a noise source including a drive unit for driving a lens, a first conversion means for Fourier transforming the audio signal acquired by the first microphone to generate a first audio signal, a second conversion means for Fourier transforming the audio signal acquired by the second microphone to generate a second audio signal, a first reduction means for generating noise data based on the second audio signal and performing processing to reduce the noise data from the first audio signal, and a detection means for detecting short-term noise from the noise source based on the second audio signal. The audio signal processing device includes a second reduction means that, when the short-term noise is detected by the detection means, controls the volume of the audio signal output from the first reduction means to reduce the short-term noise from the audio signal output from the first reduction means, and a third transformation means that performs an inverse Fourier transform on the audio signal output from the second reduction means, and the second reduction means controls the volume of the audio signal of the second frame to be reduced when the volume of the audio signal of a second frame following a first frame of the audio signal output from the first reduction means is greater than the volume of the audio signal of the first frame by more than a threshold value. [Effects of the Invention]
[0008] The voice processing device of the present invention can effectively reduce short-term noise. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a perspective view of an imaging device according to a first embodiment. [Figure 2] FIG. 1 is a block diagram showing a configuration of an imaging apparatus according to a first embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of a voice input unit of the imaging device according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing the arrangement of microphones in a voice input unit of the imaging device according to the first embodiment. [Figure 5] 4 is a timing chart showing audio processing units in the first embodiment. [Figure 6] 10 is a flowchart showing the processing content of a short-term noise processing unit in the first embodiment. [Figure 7] 5 is a timing chart showing a method for detecting short-term noise in the short-term noise processor in the first embodiment; [Figure 8] 5 is a timing chart showing short-term noise reduction processes A and B in the short-term noise processor in the first embodiment. [Figure 9] 10 is an example of a frequency spectrum of a short-term noise reduction process C in the short-term noise processor in the first embodiment. [Figure 10] FIG. 4 is a diagram illustrating noise parameters in the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0011] [First Example] <External view of imaging device 100> 1(a) and 1(b) show an example of the appearance of an image capture device 100 as an example of an audio processing device to which the present invention can be applied. Fig. 1(a) is an example of a front perspective view of the image capture device 100. Fig. 1(b) is an example of a rear perspective view of the image capture device 100. In Fig. 1, an optical lens (not shown) is attached to a lens mount 301.
[0012] The display unit 107 displays image data, text information, etc. The display unit 107 is provided on the rear surface of the imaging device 100. The extra-finder display unit 43 is a display unit provided on the top surface of the imaging device 100. The extra-finder display unit 43 displays the settings of the imaging device 100, such as the shutter speed and aperture value. The eyepiece viewfinder 16 is a peer-type viewfinder. The user can check the focus and composition of the optical image of the subject by observing the focusing screen inside the eyepiece viewfinder 16.
[0013] The release switch 61 is an operation member that allows the user to issue shooting instructions. The mode selector switch 60 is an operation member that allows the user to switch between various modes. The main electronic dial 71 is a rotary operation member. By turning this main electronic dial 71, the user can change the settings of the imaging device 100, such as the shutter speed and aperture value. The release switch 61, mode selector switch 60, and main electronic dial 71 are included in the operation unit 112.
[0014] The power switch 72 is an operating member that switches the power of the imaging device 100 on and off. The sub electronic dial 73 is a rotary operating member. The user can use the sub electronic dial 73 to move the selection frame displayed on the display unit 107 and to advance images in playback mode. The cross key 74 is a cross key (four-way key) that can be pressed up, down, left, or right. The imaging device 100 performs processing according to the part (direction) of the cross key 74 that is pressed. The power switch 72, sub electronic dial 73, and cross key 74 are included in the operation unit 112.
[0015] The SET button 75 is a push button. The SET button 75 is mainly used by the user to confirm a selection item displayed on the display unit 107, etc. The LV button 76 is a button used to switch live view (hereinafter referred to as LV) on and off. In video recording mode, the LV button 76 is used to instruct the start and stop of video shooting (recording). The enlarge button 77 is a push button used to turn enlargement mode on and off in live view display in shooting mode, and to change the magnification ratio in enlargement mode. The SET button 75, LV button 76, and enlargement button 77 are included in the operation unit 112.
[0016] In playback mode, the enlarge button 77 functions as a button for increasing the magnification of image data displayed on the display unit 107. The reduce button 78 is a button for decreasing the magnification of image data enlarged and displayed on the display unit 107. The play button 79 is an operation button for switching between shooting mode and playback mode. When the user presses the play button 79 while the imaging device 100 is in shooting mode, the imaging device 100 transitions to playback mode, and image data recorded on the recording medium 110 is displayed on the display unit 107. The reduce button 78 and play button 79 are included in the operation unit 112.
[0017] The quick-return mirror 12 (hereinafter, mirror 12) is a mirror that switches the light beam incident from an optical lens attached to the imaging device 100 so that it is incident on either the eyepiece finder 16 side or the imaging unit 101 side. The mirror 12 is raised and lowered by the control unit 111 controlling an actuator (not shown) during exposure, live view shooting, and video shooting. The mirror 12 is normally positioned so that the light beam is incident on the eyepiece finder 16. When shooting or in live view display, the mirror 12 flips up (mirror up) so that the light beam is incident on the imaging unit 101. The center of the mirror 12 is a half mirror. A portion of the light beam that passes through the center of the mirror 12 is incident on a focus detection unit (not shown) that performs focus detection.
[0018] The communication terminal 10 is a communication terminal for communication between the imaging device 100 and an optical lens 300 attached to the imaging device 100. The terminal cover 40 is a cover for protecting a connector (not shown) such as a connection cable that connects a connection cable with an external device to the imaging device 100. The lid 41 is a lid for a slot that stores the recording medium 110. The lens mount 301 is an attachment portion to which the optical lens 300 (not shown) can be attached.
[0019] The L microphone 201a and the R microphone 201b are microphones for picking up the user's voice, etc. When viewed from the rear of the imaging device 100, the L microphone 201a is placed on the left side and the R microphone 201b is placed on the right side.
[0020] <Configuration of imaging device 100> FIG. 2 is a block diagram showing an example of the configuration of the imaging device 100 according to this embodiment.
[0021] The optical lens 300 is a lens unit that can be attached to or detached from the imaging device 100. For example, the optical lens 300 is a zoom lens or a varifocal lens. The optical lens 300 has an optical lens, a motor for driving the optical lens, and a communication unit that communicates with a lens control unit 102 of the imaging device 100 (described later). The optical lens 300 can focus and zoom on a subject and correct camera shake by moving the optical lens using the motor based on a control signal received by the communication unit.
[0022] The imaging unit 101 has an imaging element for converting an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and an image processing unit for generating and outputting image data or video data from the electrical signal generated by the imaging element. The imaging element is, for example, a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS). In this embodiment, a series of processes for generating image data including still image data and video data in the imaging unit 101 and outputting the image data from the imaging unit 101 is referred to as "photographing." In the imaging device 100, the image data is recorded on a recording medium 110 (described later) in accordance with the DCF (Design rule for Camera File system) standard.
[0023] The lens control unit 102 transmits a control signal to the optical lens 300 via the communication terminal 10 based on the data output from the imaging unit 101 and a control signal output from the control unit 111 described later, and controls the optical lens 300.
[0024] The information acquisition unit 103 detects the tilt of the image capture device 100 and the temperature inside the housing of the image capture device 100. For example, the information acquisition unit 103 detects the tilt of the image capture device 100 using an acceleration sensor or a gyro sensor. Also, for example, the information acquisition unit 103 detects the temperature inside the housing of the image capture device 100 using a temperature sensor.
[0025] The audio input unit 104 generates audio data from audio acquired by a microphone. The audio input unit 104 acquires audio around the imaging device 100 using the microphone, performs analog-to-digital conversion (A / D conversion) on the acquired audio, and performs various audio processes to generate audio data. In this embodiment, the audio input unit 104 has a microphone. A detailed configuration example of the audio input unit 104 will be described later.
[0026] The volatile memory 105 temporarily stores image data generated by the imaging unit 101 and audio data generated by the audio input unit 104. The volatile memory 105 is also used as a temporary storage area for image data displayed on the display unit 107, a work area for the control unit 111, etc.
[0027] The display control unit 106 controls the display unit 107 to display image data output from the imaging unit 101, text for interactive operations, menu screens, and the like. Furthermore, when capturing still images and moving images, the display control unit 106 controls the display unit 107 to sequentially display digital data output from the imaging unit 101, thereby allowing the display unit 107 to function as an electronic viewfinder. For example, the display unit 107 is a liquid crystal display or an organic EL display. Furthermore, the display control unit 106 can also control the image data and video data output from the imaging unit 101, text for interactive operations, menu screens, and the like to be displayed on an external display via the external output unit 115, which will be described later.
[0028] The encoding processing unit 108 can encode the image data and audio data temporarily stored in the volatile memory 105. For example, the encoding processing unit 108 can generate video data by encoding and compressing image data according to the JPEG standard or a RAW image format. For example, the encoding processing unit 108 can generate video data by encoding and compressing video data according to the MPEG2 standard or the H.264 / MPEG4-AVC standard. Furthermore, for example, the encoding processing unit 108 can generate audio data by encoding and compressing audio data according to the AC3AAC standard, the ATRAC standard, or the ADPCM method. Furthermore, the encoding processing unit 108 may encode audio data without data compression, for example, according to the linear PCM method.
[0029] The recording control unit 109 can record data to and read data from the recording medium 110. For example, the recording control unit 109 can record still image data, video data, and audio data generated by the encoding processing unit 108 to and read data from the recording medium 110. The recording medium 110 is, for example, an SD card, a CF card, an XQD memory card, an HDD (magnetic disk), an optical disk, or a semiconductor memory. The recording medium 110 may be configured to be detachable from the imaging device 100, or may be built into the imaging device 100. That is, the recording control unit 109 only needs to have at least a means for accessing the recording medium 110.
[0030] The control unit 111 controls each component of the imaging device 100 via a data bus 116 in accordance with input signals and a program described below. The control unit 111 has a CPU, ROM, and RAM for executing various controls. Note that instead of the control unit 111 controlling the entire imaging device 100, multiple pieces of hardware may share and control the entire imaging device. The ROM of the control unit 111 stores programs for controlling each component. The RAM of the control unit 111 is a volatile memory used for arithmetic processing and the like.
[0031] The operation unit 112 is a user interface for receiving instructions from the user for the imaging device 100. The operation unit 112 has, for example, a power switch 72 for turning the power of the imaging device 100 on or off, a release switch 61 for issuing an instruction to shoot, a playback button for issuing an instruction to play back image data or video data, a mode change switch 60, and the like.
[0032] The operation unit 112 outputs a control signal to the control unit 111 in response to a user operation. The operation unit 112 may also include a touch panel formed on the display unit 107. The release switch 61 has SW1 and SW2. When the release switch 61 is pressed halfway, SW1 is turned on. This accepts a preparation instruction for performing preparatory operations for image capture, such as AF (autofocus) processing, AE (auto exposure) processing, AWB (auto white balance) processing, and EF (pre-flash) processing. When the release switch 61 is pressed fully, SW2 is turned on. This accepts an image capture instruction for performing an image capture operation. The operation unit 112 also includes an operation member (e.g., a button) that can adjust the volume of audio data played from a speaker 114, which will be described later.
[0033] The audio output unit 113 can output audio data to the speaker 114 and the external output unit 115. The audio data input to the audio output unit 113 includes audio data read from the recording medium 110 by the recording control unit 109, audio data output from the nonvolatile memory 117, and audio data output from the encoding processing unit. The speaker 114 is an electro-acoustic transducer that can reproduce audio data.
[0034] The external output unit 115 can output image data, video data, audio data, etc. to an external device. The external output unit 115 is configured with, for example, a video terminal, a microphone terminal, a headphone terminal, etc.
[0035] The data bus 116 is a data bus for transmitting various data such as audio data, video data, and image data, as well as various control signals, to each block of the image capturing device 100 .
[0036] The nonvolatile memory 117 is a nonvolatile memory that stores programs, etc., which are executed by the control unit 111 and will be described later. Also, sound data is recorded in the nonvolatile memory 117. This sound data is, for example, sound data of electronic sounds such as a focusing sound that is output when a subject is focused, an electronic shutter sound that is output when an instruction to take a photograph is given, and an operation sound that is output when the imaging device 100 is operated.
[0037] <Operation of the imaging device 100> The operation of the imaging device 100 of this embodiment will now be described.
[0038] In the image capture device 100 of this embodiment, power is supplied to each component of the image capture device from a power supply (not shown) in response to a user turning on the power by operating the power switch 72. For example, the power supply is a battery such as a lithium ion battery or an alkaline manganese dry battery.
[0039] In response to the supply of power, the control unit 111 determines whether the camera will operate in a shooting mode or a playback mode, for example, based on the state of the mode selector switch 60. In the video recording mode, the control unit 111 records the video data output from the imaging unit 101 and the audio data output from the audio input unit 104 as one piece of video data with audio. In the playback mode, the control unit 111 reads out the image data or video data recorded on the recording medium 110 using the recording control unit 109, and controls the display unit 107 to display it.
[0040] First, the moving image recording mode will be described. In the moving image recording mode, the control unit 111 first transmits a control signal to each component of the imaging device 100 to transition the imaging device 100 to a shooting standby state. For example, the control unit 111 controls the imaging unit 101 and the audio input unit 104 to perform the following operations.
[0041] The imaging unit 101 converts an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and generates video data from the electrical signal generated by the imaging element. The imaging unit 101 then transmits the video data to the display control unit 106, which displays it on the display unit 107. The user can prepare for shooting while viewing the video data displayed on the display unit 107.
[0042] The audio input unit 104 performs A / D conversion on analog audio signals input from multiple microphones, respectively, to generate multiple digital audio signals. The audio input unit 104 then generates audio data for multiple channels from the multiple digital audio signals. The audio input unit 104 transmits the generated audio data to the audio output unit 113, which plays the audio data from the speaker 114. While listening to the audio data played from the speaker 114, the user can use the operation unit 112 to adjust the volume of the audio data to be recorded in the audio-accompanied video data.
[0043] Next, in response to the user pressing the LV button 76, the control unit 111 transmits an instruction signal to start shooting to each component of the image capturing device 100. For example, the control unit 111 controls the image capturing unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.
[0044] The imaging unit 101 converts an optical image of a subject formed on an imaging surface via the optical lens 300 into an electrical signal, and generates video data from the electrical signal generated by the imaging element. The imaging unit 101 then transmits the video data to a display control unit 106, which displays the video data on a display unit 107. The imaging unit 101 also transmits the generated video data to a volatile memory 105.
[0045] The audio input unit 104 performs A / D conversion on analog audio signals input from multiple microphones to generate multiple digital audio signals. The audio input unit 104 then generates multi-channel audio data from the multiple digital audio signals. The audio input unit 104 then transmits the generated audio data to the volatile memory 105.
[0046] The encoding processor 108 reads out the video data and audio data temporarily recorded in the volatile memory 105 and encodes them. The control unit 111 generates a data stream from the video data and audio data encoded by the encoding processor 108 and outputs it to the recording control unit 109. The recording control unit 109 records the input data stream as video data with audio on the recording medium 110 in accordance with a file system such as UDF or FAT.
[0047] The components of the image capturing apparatus 100 continue to perform the above operations during video capture.
[0048] Then, in response to the user pressing the LV button 76, the control unit 111 transmits an instruction signal to end shooting to each component of the image capturing device 100. For example, the control unit 111 controls the image capturing unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.
[0049] The image capturing unit 101 stops generating moving image data, and the audio input unit 104 stops generating audio data.
[0050] The encoding processing unit 108 reads and encodes the remaining video data and audio data recorded in the volatile memory 105. The control unit 111 generates a data stream from the video data and audio data encoded by the encoding processing unit 108 and outputs the data stream to the recording control unit 109.
[0051] The recording control unit 109 records the data stream as a file of audio-accompanying moving image data on the recording medium 110 in accordance with a file system such as UDF or FAT. Then, when the input of the data stream stops, the recording control unit 109 completes the audio-accompanying moving image data. Upon completion of the audio-accompanying moving image data, the recording operation of the imaging device 100 stops.
[0052] In response to the stop of the recording operation, the control unit 111 transmits a control signal to each component of the imaging device 100 to transition to a shooting standby state. As a result, the control unit 111 controls the imaging device 100 to return to the shooting standby state.
[0053] Next, the playback mode will be described. In the playback mode, the control unit 111 transmits a control signal to each component of the image capture device 100 to transition to a playback state. For example, the control unit 111 controls the encoding processing unit 108, the recording control unit 109, the display control unit 106, and the audio output unit 113 to perform the following operations.
[0054] The recording control unit 109 reads out the moving image data with audio recorded on the recording medium 110 and transmits the read moving image data with audio to the encoding processing unit 108 .
[0055] The encoding processing unit 108 decodes the audio-accompanying video data into image data and audio data. The encoding processing unit 108 transmits the decoded video data to the display control unit 106 and the decoded audio data to the audio output unit 113.
[0056] The display control unit 106 displays the decoded image data on the display unit 107. The audio output unit 113 reproduces the decoded audio data through the speaker 114.
[0057] As described above, the imaging device 100 of this embodiment can record and play back image data and audio data.
[0058] In this embodiment, the audio input unit 104 performs audio processing such as adjusting the level of an audio signal input from a microphone. In this embodiment, the audio input unit 104 performs this audio processing in response to the start of video recording. Note that this audio processing may be performed after the imaging device 100 is turned on. This audio processing may also be performed in response to the selection of a shooting mode. This audio processing may also be performed in response to the selection of a mode related to audio recording, such as a video recording mode or an audio memo function. This audio processing may also be performed in response to the start of audio signal recording.
[0059] <Configuration of the voice input unit 104> FIG. 3 is a block diagram showing an example of a detailed configuration of the voice input unit 104 in this embodiment.
[0060] In this embodiment, the audio input unit 104 has three microphones: an L microphone 201a, an R microphone 201b, and a noise microphone 201c. The L microphone 201a and the R microphone 201b are each an example of a first microphone. In this embodiment, the image capture device 100 collects environmental sounds using the L microphone 201a and the R microphone 201b, and records the audio signals input from the L microphone 201a and the R microphone 201b in stereo. For example, environmental sounds are sounds generated outside the housing of the image capture device 100 and outside the housing of the optical lens 300, such as the user's voice, animal cries, the sound of rain, and music.
[0061] The noise microphone 201c is an example of a second microphone. The noise microphone 201c is a microphone for acquiring noise, such as drive sounds from a predetermined noise source, generated within the housing of the image capture device 100 and the housing of the optical lens 300. Examples of noise sources include driving units such as an ultrasonic motor (USM) and a stepper motor (STM). The noise is, for example, vibration noise generated by driving a motor such as a USM or STM. For example, the motor is driven during AF processing to focus on a subject. The image capture device 100 acquires noise, such as drive sounds, generated within the housing of the image capture device 100 and the housing of the optical lens 300 using the noise microphone 201c, and generates noise parameters (described later) using audio data of the acquired noise. In this embodiment, the left microphone 201a, the right microphone 201b, and the noise microphone 201c are omnidirectional microphones. An example of the arrangement of the L microphone 201a, the R microphone 201b, and the noise microphone 201c in this embodiment will be described later with reference to FIG.
[0062] The L microphone 201a, R microphone 201b, and noise microphone 201c each generate an analog audio signal from the captured audio and input it to the A / D conversion unit 202. Here, the audio signal input from the L microphone 201a is referred to as Lch, the audio signal input from the R microphone 201b as Rch, and the audio signal input from the noise microphone 201c as Nch.
[0063] The A / D conversion unit 202 converts the analog audio signals input from the L microphone 201a, R microphone 201b, and noise microphone 201c into digital audio signals. The A / D conversion unit 202 outputs the converted digital audio signals to the FFT unit 203. In this embodiment, the A / D conversion unit 202 converts the analog audio signals into digital audio signals by performing sampling processing with a sampling frequency of 48 kHz and a bit depth of 16 bits.
[0064] The FFT unit 203 performs fast Fourier transform processing on the time-domain digital audio signal input from the A / D conversion unit 202 to convert it into a frequency-domain digital audio signal. In this embodiment, the frequency-domain digital audio signal has a frequency spectrum of 1024 points in a frequency band from 0 Hz to 48 kHz. Furthermore, the frequency-domain digital audio signal has a frequency spectrum of 513 points in a frequency band from 0 Hz to 24 kHz, which is the Nyquist frequency. In this embodiment, the imaging device 100 performs noise reduction processing using the 513-point frequency spectrum from 0 Hz to 24 kHz of the audio data output from the FFT unit 203.
[0065] Here, the frequency spectrum of the Lch subjected to the fast Fourier transform is represented by 513-point array data of Lch_Before[0] to Lch_Before
[0512] . When these array data are referred to collectively, they are referred to as Lch_Before. Furthermore, the frequency spectrum of the Rch subjected to the fast Fourier transform is represented by 513-point array data of Rch_Before[0] to Rch_Before
[0512] . When these array data are referred to collectively, they are referred to as Rch_Before. Note that Lch_Before and Rch_Before are each an example of first frequency spectrum data.
[0066] The Nch frequency spectrum after the fast Fourier transform is represented by array data of 513 points, Nch_Before[0] to Nch_Before
[0512] . These array data are collectively referred to as Nch_Before. Nch_Before is an example of second frequency spectrum data.
[0067] The switching unit 204 switches the path based on control information from the lens control unit 102. In this embodiment, when the optical lens 300 is driven, the switching unit 204 switches the path so that noise reduction processing is performed in the subtraction processing unit A207, which will be described later. Also, when the optical lens 300 is not driven, the switching unit 204 switches the path so that noise reduction processing is not performed in the subtraction processing unit A207.
[0068] The noise data generation unit A205 generates data for reducing noise related to lens driving included in Lch_Before and Rch_Before based on Nch_Before. In this embodiment, the noise data generation unit A205 uses noise parameters to generate array data NLA[0] to NLA
[0512] for reducing noise included in Lch_Before[0] to Lch_Before
[0512] , respectively. The noise data generation unit A205 also generates array data NRA[0] to NRA
[0512] for reducing noise included in Rch_Before[0] to Rch_Before
[0512] , respectively.
[0069] Note that the frequency points in the array data of NLA[0] to NLA
[0512] are the same as the frequency points in the array data of Lch_Before[0] to Lch_Before
[0512] . Also, the frequency points in the array data of NRA[0] to NRA
[0512] are the same as the frequency points in the array data of Rch_Before[0] to Rch_Before
[0512] .
[0070] The array data NLA[0] to NLA
[0512] are collectively referred to as NLA. Furthermore, the array data NRA[0] to NRA
[0512] are collectively referred to as NRA. NLA and NRA are each an example of third frequency spectrum data.
[0071] The noise parameter recording unit 206 records noise parameters used by the noise data generating unit A205 to generate NLA and NRA from Nch_Before. In this embodiment, noise parameters related to lens drive for each lens type, which are noise parameters used by the noise data generating unit A205, are recorded in the noise parameter recording unit 206. In this embodiment, the noise data generating unit A205 does not switch noise parameters while recording audio data.
[0072] The noise parameter recording unit 206 also records noise parameters used by a noise data generating unit B 208 (described later) to generate NLB and NRB from Nch_Before.
[0073] Here, the noise parameters for generating an NLA from Nch_Before are collectively referred to as PLxA, and the noise parameters for generating an NRA from Nch_Before are collectively referred to as PRxA.
[0074] PLxA and PRxA have the same number of sequences as NLA and NRA, respectively. For example, PL1A is sequence data from PL1A[0] to PL1A
[0512] . The frequency points of PL1A are the same as the frequency points of Lch_Before. For example, PR1A is sequence data from PR1A[0] to PR1A
[0512] . The frequency points of PR1A are the same as the frequency points of Rch_Before. The noise parameters will be described later with reference to FIG. 10.
[0075] In this embodiment, the noise parameter recording unit 206 records all of the coefficients for each of the 513 frequency spectrum points as noise parameters. However, the noise parameter recording unit 206 need only record coefficients for at least the frequency points necessary to reduce noise, rather than coefficients for all 513 frequency points. For example, the noise parameter recording unit 206 may record, as noise parameters, coefficients for each frequency spectrum from 20 Hz to 20 kHz, which are considered to be typical audible frequencies, but may not record coefficients for other frequency spectrums. Furthermore, for example, coefficients for frequency spectrums whose coefficient value is zero may not be recorded in the noise parameter recording unit 206 as noise parameters.
[0076] The subtraction processing unit A207 subtracts NLA and NRA from Lch_Before and Rch_Before, respectively. In this embodiment, the subtraction processing unit A207 reduces high-level noise regardless of whether it is short-term noise or long-term noise.
[0077] The subtraction processing unit A207 also has an L subtractor A207a that subtracts the NLA from Lch_Before, and an R subtractor A207b that subtracts the NRA from Rch_Before. The L subtractor A207a subtracts the NLA from Lch_Before and outputs 513 points of array data from Lch_A_After[0] to Lch_A_After
[0512] . The R subtractor A207b subtracts the NRA from Rch_Before and outputs 513 points of array data from Rch_A_After[0] to Rch_A_After
[0512] . In this embodiment, the subtraction processing unit A207 performs subtraction processing using a spectral subtraction method.
[0078] The noise data generating unit B208 generates data for reducing noise contained in Lch_A_After and Rch_A_After based on Nch_Before.
[0079] In this embodiment, the noise data generation unit B208 generates array data of NLB[0] to NLB
[0512] for reducing noise contained in Lch_A_After[0] to Lch_A_After
[0512] , respectively, using noise parameters. Also, the noise data generation unit B208 generates array data of NRB[0] to NRB
[0512] for reducing noise contained in Rch_A_After[0] to Rch_A_After
[0512] , respectively, using noise parameters.
[0080] The frequency points in the array data of NLB[0] to NLB
[0512] are the same as the frequency points in the array data of Lch_A_After[0] to Lch_A_After
[0512] . Also, the frequency points in the array data of NRB[0] to NRB
[0512] are the same as the frequency points in the array data of Rch_A_After[0] to Rch_A_After
[0512] .
[0081] The sequence data of NLB[0] to NLB
[0512] is collectively referred to as NLB. The sequence data of NLB[0] to NRB
[0512] is collectively referred to as NRB. NLB and NRB are examples of (fourth frequency spectrum data).
[0082] In this embodiment, the noise parameter recording unit 206 records a plurality of types of noise parameters corresponding to the types of noise, which are noise parameters used in the noise data generating unit B 208.
[0083] Here, the noise parameters for generating NLB from Nch_Before are collectively referred to as PLxB, and the noise parameters for generating NRB from Nch_Before are collectively referred to as PRxB.
[0084] PLxB and PRxB have the same number of sequences as NLB and NRB, respectively. For example, PL1B is sequence data from PL1B[0] to PL1B
[0512] . The frequency points of PL1B are the same as those of Lch_Before. For example, PR1B is sequence data from PR1B[0] to PR1B
[0512] . The frequency points of PR1B are the same as those of Rch_Before. Noise parameters will be described later with reference to FIG. 10. In this embodiment, the noise parameter recording unit 206 records all coefficients for each of the 513 frequency spectrum points as noise parameters. However, instead of coefficients for all 513 frequency points, the noise parameter recording unit 206 only needs to record coefficients for at least the frequency points necessary to reduce noise. For example, the noise parameter recording unit 206 may record, as noise parameters, coefficients for each of the frequency spectrum from 20 Hz to 20 kHz, which are considered to be typical audible frequencies, but not coefficients for other frequency spectrums. Furthermore, for example, coefficients for frequency spectra whose coefficient values are zero may not be recorded in the noise parameter recording unit 206 as noise parameters.
[0085] The subtraction processing unit B209 subtracts NLB and NRB from Lch_A_After and Rch_A_After, respectively. For example, the subtraction processing unit B209 has an L subtractor B209a that subtracts NLB from Lch_A_AFTER, and an R subtractor B209b that subtracts NRB from Rch_Before. The L subtractor B209a subtracts NLB from Lch_Before and outputs 513-point array data from Lch_After[0] to Lch_After
[0512] . The R subtractor B209b subtracts NRA from Rch_Before and outputs 513-point array data from Rch_After[0] to Rch_After
[0512] . In this embodiment, the subtraction processing unit B209 performs subtraction processing using a spectral subtraction method.
[0086] In this embodiment, the subtraction processing unit B209 subtracts constantly occurring noise, such as microphone floor noise and electrical noise, other than noise generated by lens driving. Note that, in this embodiment, the noise data generation unit B208 generates NLB and NRB based on Nch_Before, but other methods may be used. For example, NLB and NRB may be recorded in the noise parameter recording unit 206, and the subtraction processing unit B209 may read NLB and NRB directly from the noise parameter recording unit 206 without going through the noise data generation unit B208. This is because there is little need to refer to the noise included in Nch_Before due to constantly occurring noise, such as microphone floor noise and electrical noise.
[0087] The short-term noise detector 210 detects short-term noise from Nch_Before. The short-term noise is, for example, short-term noise generated by the meshing of gears inside the optical lens 300. On the other hand, the long-term noise is, for example, sliding noise inside the housing of the optical lens 300. The short-term noise detector 210 may detect short-term noise from Nch_Before, or from Lch_Before or Rch_Before.
[0088] The short-term noise subtraction processor 211 performs noise reduction processing, particularly to reduce short-term noise, on the audio signal input from the subtraction processor A 207. That is, in this embodiment, while the lens is being driven, noise reduction processing is performed by the subtraction processor A 207 and the short-term noise subtraction processor 211 before processing by the subtraction processor B 209.
[0089] The data buffer 212 is a buffer (memory) that temporarily stores data used for the short-term noise subtraction processing unit 211.
[0090] The processing of the short-term noise detector 210, the short-term noise subtractor 211, and the data buffer 212 will be described in detail later.
[0091] The iFFT unit 213 performs an inverse fast Fourier transform (inverse Fourier transform) on the frequency domain digital audio signal input from the subtraction processing unit B209 to convert it into a time domain digital audio signal.
[0092] The audio processing unit 214 performs audio processing on the time domain digital audio signal, such as an equalizer, an auto level controller, and stereo enhancement processing, etc. The audio processing unit 214 outputs the processed audio data to the volatile memory 105.
[0093] In this embodiment, the imaging device 100 has two microphones as the first microphone, but the imaging device 100 may have one microphone or three or more microphones as the first microphone. For example, when the imaging device 100 has one microphone as the first microphone in the audio input unit 104, the imaging device 100 records audio data picked up by the one microphone in monaural format. Also, when the imaging device 100 has three or more microphones as the first microphone in the audio input unit 104, the imaging device 100 records audio data picked up by the three or more microphones in surround format.
[0094] In this embodiment, the L microphone 201a, the R microphone 201b, and the noise microphone 201c are non-directional microphones, but these microphones may also be directional microphones.
[0095] In this embodiment, the subtraction processing unit B209 reduces constant noise, but other methods may be used. For example, if the subtraction processing unit A207 also has the function of the subtraction processing unit B209, the noise reduction process by the subtraction processing unit B209 may not be performed.
[0096] <Arrangement of microphones in the audio input unit 104> Here, an example of the arrangement of the microphones of the voice input unit 104 of this embodiment will be described. Fig. 4 shows an example of the arrangement of the L microphone 201a, the R microphone 201b, and the noise microphone 201c.
[0097] 4 is an example of a cross-sectional view of a portion of the imaging device 100 to which the L microphone 201a, the R microphone 201b, and the noise microphone 201c are attached. This portion of the imaging device 100 is composed of an exterior part 302, a microphone bushing 303, and a fixing part 304.
[0098] The exterior part 302 has holes (hereinafter referred to as microphone holes) for inputting environmental sounds into the microphones. In this embodiment, the microphone holes are formed above the L microphone 201a and the R microphone 201b. On the other hand, the noise microphone 201c is provided to acquire drive sounds generated within the housing of the image capture device 100 and the housing of the optical lens 300, and does not need to acquire environmental sounds. Therefore, in this embodiment, no microphone holes are formed in the exterior part 302 above the noise microphone 201c.
[0099] Drive sounds generated within the housings of the image capture device 100 and the optical lens 300 are picked up by the L microphone 201a and the R microphone 201b through the microphone holes. If drive sounds or the like are generated within the housings of the image capture device 100 and the optical lens 300 when ambient noise is low, the sound picked up by each microphone will mainly be this drive sound. Therefore, the sound level from the noise microphone 201c is higher than the sound levels from the L microphone 201a and the R microphone 201b. In other words, in this case, the relationship between the levels of the audio signals output from each microphone is as follows: Lch≒Rch <Nch Furthermore, when the environmental sound becomes louder, the audio level of the environmental sound from the L microphone 201a and R microphone 201b becomes louder than the audio level of the drive sound generated by the image capture device 100 or the optical lens 300 and output from the noise microphone 201c. Therefore, in this case, the relationship between the levels of the audio signals output from each microphone is as follows: Lch ≒ Rch > Nch In this embodiment, the shape of the microphone hole formed in exterior part 302 is elliptical, but it may be other shapes such as circular or rectangular. Furthermore, the shape of the microphone hole on microphone 201a and the shape of the microphone hole on microphone 201b may be different from each other.
[0100] In this embodiment, the noise microphone 201c is arranged so as to be close to the L microphone 201a and the R microphone 201b. Also, in this embodiment, the noise microphone 201c is arranged between the L microphone 201a and the R microphone 201b. As a result, the audio signal generated by the noise microphone 201c from the drive sound etc. generated in the housing of the imaging device 100 and in the housing of the optical lens 300 becomes a signal similar to the audio signal generated by the L microphone 201a and the R microphone 201b from this drive sound etc.
[0101] The microphone bushing 303 is a member for fixing the L microphone 201a, the R microphone 201b, and the noise microphone 201c. The fixing portion 304 is a member for fixing the microphone bushing 303 to the exterior portion 302.
[0102] In this embodiment, the exterior portion 302 and the fixing portion 304 are constituted by mold members such as PC materials. Also, the exterior portion 302 and the fixing portion 304 may be constituted by metal members such as aluminum or stainless steel. Also, in this embodiment, the microphone bushing 303 is constituted by a rubber material such as ethylene propylene diene rubber.
[0103] <Processing method of the FFT unit 203> The processing performed by the FFT unit 203 will be described using FIG. 5.
[0104] FIG. 5(a) shows an example of an audio signal in the time domain. In this embodiment, the audio signal is a signal recorded with a sampling frequency of 48 kHz and a bit depth of 24 bits.
[0105] FIG. 5(b) shows an example of the unit of the data length of the audio signal processed by the FFT unit 203. In this embodiment, FFT is performed on the audio signal in units of 1024 samples. In this embodiment, an audio signal of 1024 samples is taken as one frame. The FFT unit 203 performs FFT in response to buffering the audio signal for one frame.
[0106] In this embodiment, the audio input unit 104 performs noise reduction processing using the overlap-and-add method. For example, the audio input unit 104 performs noise reduction processing so as to overlap every 512 samples (half a frame).
[0107] Here, the method for describing each frame will be explained. For example, in FIG. 5(b), the audio signal of one frame generated by FFT processing at time T501 is frame data [t]. In this case, the audio signal of the frame generated immediately before (just before) frame data [t] is described as frame data [t-1], and the audio signal of the frame generated immediately after (just after) frame data [t] is described as frame data [t+1]. In this way, each frame data is described based on the audio signal of the frame subjected to FFT processing at a certain time. Furthermore, the frame data includes audio signals of Lch, Rch, and Nch, and the frequency spectrum is stored as array data for each channel. For example, when specifically describing the channels and frequency spectra, in the above example, the audio signal with the nth frequency spectrum of Lch at time T501 is described as frame data L[t][n].
[0108] <Short-term noise reduction processing> The short-term noise reduction process in the short-term noise detector 210 and the short-term noise subtractor 211 will be described with reference to FIG.
[0109] 6(a) is a flowchart showing an example of a process for reducing short-term noise, in which processing of one frame of frame data is explained.
[0110] In step S601, it is determined whether the optical lens 300 is being driven. For example, the switching unit 204 determines whether the optical lens 300 is being driven based on control information input from the lens control unit 102. If it is determined that the optical lens 300 is being driven, the switching unit 204 switches the path so that Lch_Before and Rch_Before are input to the subtraction processing unit A207. If it is determined that the optical lens 300 is not being driven, the switching unit 204 switches the path so that Lch_Before and Rch_Before are input to the subtraction processing unit B209.
[0111] In step S602, the short-term noise detector 210 determines whether or not one frame of frame data contains short-term noise. In this embodiment, the short-term noise detector 210 calculates the sound volume N[t]_Power for one frame from the frame data N[t][0] to
[0512] . If the value of N[t]_Power is less than a predetermined threshold, the process of this flowchart ends. On the other hand, if N[t]_Power is equal to or greater than the predetermined threshold, the process of step S603 is executed.
[0112] The short-term noise detector 210 may calculate N[t]_Power by weighting it for a specific frequency band or for each frequency. In this embodiment, the short-term noise detector 210 calculates it from the frequency spectrum, but it may also calculate it from the amplitude value of the audio signal in the time domain.
[0113] In step S603, the short-term noise detector 210 determines whether short-term noise has been continuously detected. That is, the short-term noise detector 210 determines whether short-term noise is included in a predetermined number of consecutive frames or more. For example, the short-term noise detector 210 determines whether short-term noise has been detected five consecutive times. This is because if short-term noise is detected a predetermined number of times or more consecutively, the noise is no longer considered to be short-term noise but long-term noise. If short-term noise has not been detected a predetermined number of times or more consecutively, the process of step S604 is executed. If short-term noise has been detected a predetermined number of times or more consecutively, the process of this flowchart ends.
[0114] The reason why Nch_Before is used to detect short-term noise is as follows: As described above, the noise picked up by the noise microphone 201c is louder than the noise picked up by the L microphone 201a and the R microphone 201b. In addition, microphone holes are formed above the L microphone 201a and the R microphone 201b, but no microphone hole is formed above the noise microphone 201c. In other words, the environmental sound picked up by the noise microphone 201c is quieter than the environmental sound picked up by the L microphone 201a and the R microphone 201b. In other words, the signal generated from the sound picked up by the noise microphone 201c is a signal with quieter environmental sound and louder noise than the signal generated from the sound picked up by the L microphone 201a and the R microphone 201b. For this reason, it can be said that Nch_Before is an audio signal more suitable for noise detection than Lch_Before and Rch_Before.
[0115] A detailed method for detecting short-term noise will be described later with reference to FIG.
[0116] In steps S604 to S606, the short-term noise subtraction processing unit 211 performs processing to reduce short-term noise. In step S604, the short-term noise subtraction processing unit 211 performs reduction processing A. In step S605, the short-term noise subtraction processing unit 211 performs reduction processing B. In step S606, the short-term noise subtraction processing unit 211 performs reduction processing C. Details of each reduction processing will be described later. Note that in this embodiment, three reduction processing processes, reduction processing A to C, are performed, but only one of the reduction processing processes may be performed. Furthermore, the order in which reduction processing A to C are performed is not limited to this order and may be any order.
[0117] In step S607, upon completion of processing of the frame data, the short-term noise subtraction processing unit 211 stores (records) the frame data L[t] and frame data R[t] in the data buffer 212. Thereafter, the short-term noise subtraction processing unit 211 treats these frame data as frame data L[t-1] and frame data R[t-1], respectively.
[0118] The short-term noise reduction process has been explained above. Now, reduction processes A to C will be explained.
[0119] First, a description will be given of the reduction process A. Fig. 6(b) is a flowchart showing an example of the reduction process A.
[0120] In step S611, the short-term noise subtraction processor 211 determines whether the frame data [t] is greater than the frame data [t-1] by a predetermined value or more. For example, the short-term noise subtraction processor 211 determines whether the value of the frame data L[t][n] is greater than the value of the frame data L[t-1][n-1] by a threshold value P1 (e.g., 6 dB) or more. If it is determined that the frame data [t] is greater than the frame data [t-1] by a predetermined value or more, the process of step S612 is executed. If it is determined that the frame data [t] is not greater than the frame data [t-1] by a predetermined value or more, the process of step S614 is executed. Note that the short-term noise subtraction processor 211 may also use the frame data R[t] to determine whether the frame data L[t] is greater than the frame data L[t-1] by a predetermined value or more.
[0121] In step S612, the short-term noise subtraction processor 211 performs noise reduction processing on the frame data [t]. For example, the short-term noise subtraction processor 211 calculates the value of the frame data L[t][n] to be the value obtained by adding a threshold value P1 to the frame data L[t-1][n], as shown in the following formula 1. [Formula 1]L[t][n]←L[t-1][n]+P1
[0122] In step S613, the short-term noise subtraction processor 211 changes the threshold P1 to a value P1_Low that is smaller than the threshold P1. For example, if the initial value of the threshold P1 is 6 dB, the short-term noise subtraction processor 211 changes the threshold P1 by setting the value P1_Low to 3 dB. That is, in this embodiment, the threshold P1 is changed from 6 dB to 3 dB.
[0123] In step S614, the short-term noise subtraction processor 211 changes the threshold P1 to a value P1_High that is greater than the threshold P1. In this embodiment, the value P1_High is also greater than the value P1_Low. In this embodiment, the value P1_High is assumed to be the same as the threshold P1. That is, in this embodiment, if the threshold P1 is at its initial value in the processing of step S612, the threshold P1 is not changed. On the other hand, if the threshold P1 was changed to the value P1_Low in the processing of step S612, the threshold P1 is returned to its initial value by the processing of this step.
[0124] The above describes the processing of the reduction process A. A timing chart of the processing of this flowchart will be described later with reference to FIG.
[0125] The processing of this flowchart is similar to that of the frame data R[t].
[0126] Next, the reduction process B will be described. Fig. 6(c) is a flowchart showing an example of the reduction process B.
[0127] In step S621, the short-term noise subtraction processing unit 211 stores the frame data [t]. For example, the short-term noise subtraction processing unit 211 stores the frame data L[t] and the frame data R[t] in the data buffer 212.
[0128] In step S622, the short-term noise subtraction processor 211 determines whether the frame data [t] is smaller than the newly input frame data [t+1] by a predetermined value or more. For example, the short-term noise subtraction processor 211 determines whether the value of the frame data L[t][n] is larger than the value of the frame data L[t+1][n+1] by a threshold P2 (e.g., 3 dB) or more. The short-term noise subtraction processor 211 may also use the frame data R[t] to determine whether the frame data R[t] is larger than the frame data R[t+1] by a predetermined value or more. If it is determined that the frame data [t] is smaller than the newly input frame data [t+1] by a predetermined value or more, the process of step S623 is executed. If it is determined that the frame data [t] is not smaller than the newly input frame data [t+1] by a predetermined value or more, the process of this flowchart is terminated.
[0129] In step S623, the short-term noise subtraction unit 211 performs noise reduction processing on the frame data [t]. For example, the short-term noise subtraction unit 211 calculates the value of the frame data L[t][n] to become the frame data L[t-1][n], as shown in the following formula 2. [Formula 2] L[t][n]←L[t-1][n]
[0130] The above describes the processing of the reduction process B. A timing chart of the processing of this flowchart will be described later with reference to FIG.
[0131] The processing of this flowchart is similar to that of the frame data R[t].
[0132] In this way, when reducing short-term noise, the short-term noise subtraction processing unit 211 performs noise reduction by switching between a plurality of thresholds.
[0133] Next, the processing of the reduction process C will be described. Fig. 6(d) is a flowchart showing an example of the reduction process C.
[0134] In step S631, the short-term noise subtraction processor 211 calculates the average value in a specific frequency band of the frame data [t]. The specific frequency band is a frequency band in which noise is easily audibly detected and easily occurs. In this embodiment, the specific frequency band is 1 kHz to 4 kHz. The average value in the specific frequency band of the frame data L[t] is defined as L_ave[t].
[0135] In step S632, the short-term noise subtraction processor 211 determines whether the average value in a specific frequency band of frame data [t] is greater than the average value in a specific frequency band of frame data [t-1]. For example, the short-term noise subtraction processor 211 determines whether L_ave[t] is greater than L_ave[t-1]. If it is determined that the average value in a specific frequency band of frame data [t] is greater than the average value in a specific frequency band of frame data [t-1], the process of step S633 is executed. If it is determined that the average value in a specific frequency band of frame data [t] is not greater than the average value in a specific frequency band of frame data [t-1], the process of this flowchart ends.
[0136] In step S633, the short-term noise subtraction processor 211 performs noise reduction processing so as to bring the average value in a specific frequency band of the frame data [t] closer to the average value in a specific frequency band of the frame data [t-1]. For example, in this embodiment, the short-term noise subtraction processor 211 calculates the value of the frame data L[t][n] as shown in the following formula 3 so that L_ave[t] approaches L_ave[t-1]. [Formula 3] L[t][n]←L[t][n]-(L_ave[t]-L_ave[t-1])
[0137] The above describes the reduction process C. A timing chart of the process of this flowchart will be described later with reference to FIG.
[0138] The processing of this flowchart is similar to that of the frame data R[t].
[0139] <Timing chart of the short-term noise detection unit 210> A method for detecting short-term noise in the short-term noise detector 210 will be described with reference to the timing chart of FIG.
[0140] 7(a) shows an example of a lens control signal. The lens control signal is a signal that the lens control unit 102 uses to instruct the optical lens 300 to drive. In this embodiment, the level of the lens control signal is expressed as two values: High and Low. When the level of the lens control signal is High, the lens control unit 102 is instructing the optical lens 300 to drive. When the level of the lens control signal is Low, the lens control unit 102 is not instructing the optical lens 300 to drive.
[0141] FIG. 7(b) is a graph showing an example of N[t]_Power. The vertical axis represents the value of N[t]_Power. The horizontal axis represents time. When short-term noise occurs, the value of N[t]_Power increases. The short-term noise detector 210 detects the occurrence of short-term noise when the optical lens 300 is driven and N[t]_Power is equal to or greater than a predetermined value. For example, if N[t]_Power is greater than the short-term noise detection threshold between times T701 and T702 and between times T703 and T704, it is determined that short-term noise has occurred. However, if N[t]_Power is equal to or greater than a predetermined value for a certain period, such as period T705, the short-term noise detector 210 treats that period as a period in which no short-term noise is occurring.
[0142] <Short-term noise reduction timing chart> First, reduction processes A and B will be described using the timing chart of Fig. 8. Next, reduction process C will be described using Fig. 9.
[0143] Fig. 8(a) is an example of a lens control signal. Fig. 8(b) is a graph showing an example of N[t]_Power. Fig. 8(a) and Fig. 8(b) are the same as the graphs for the period from time T701 to time T702 in Fig. 7(a) and Fig. 7(b), respectively.
[0144] FIG. 8(c) shows an example of a frequency spectrum after reduction process A. In this embodiment, the frequency spectrum of frame data L[t][n] is shown. The vertical axis indicates the power value of the frequency spectrum. Note that similar processes are performed on frame data L[t] and frame data R[t] at other frequencies.
[0145] The plain area 811 (including the diagonal and hatched areas) is the frequency spectrum input from the short-term noise subtraction processing unit 211 (the frequency spectrum before reduction processing A is performed), and the diagonal area 812 is the frequency spectrum generated by performing reduction processing A.
[0146] The vertical axis indicates L[t][n] for each time t of the characteristic frequency N.
[0147] First, in section T801, the level of the frequency spectrum at time t when short-term noise is detected is greater than the level of the frequency spectrum at time t-1 by at least threshold P1 (=P1_High). Therefore, in reduction process A, the level of the frequency spectrum at time t is attenuated so that the level of the frequency spectrum is greater than the level of the frequency spectrum at time t-1 by P1 (=P1_High), as shown in Equation 4. The diagonal line area 812 (including the shaded area) indicates the frequency spectrum reduced by reduction process A. Note that the value of threshold P1 is changed to P1_Low as a result of reduction process A being performed at time t. [Formula 4] L[t][n]←L[t-1][n]+P1_High
[0148] Furthermore, the level of the frequency spectrum at time t+1 is greater than the level of the frequency spectrum at time t by at least a threshold P1 (=P1_Low). Therefore, in reduction process A, the level of the frequency spectrum at time t+1 is attenuated so that it becomes a frequency spectrum that is greater by P1 (=P1_Low) than the level of the frequency spectrum at time t, as shown in Equation 5. [Formula 5] L[t][n]←L[t-1][n]+P1_Low
[0149] The above process is similarly performed for sections T802 and T804.
[0150] Next, in section T803, the level of the frequency spectrum at time t when the short-term noise is detected is not higher than the level of the frequency spectrum at time t-1 by more than the threshold P1 (=P1_High). Therefore, reduction process A is not performed on the frequency spectrum at time t. Here, since reduction process A is not performed, the value of threshold P1 is not changed.
[0151] Furthermore, the level of the frequency spectrum at time t+1 is greater than the level of the frequency spectrum at time t by at least a threshold P1 (=P1_High). Therefore, in reduction process A, the level of the frequency spectrum at time t+1 is attenuated so that it becomes a frequency spectrum that is greater than the level of the frequency spectrum at time t by P1 (=P1_High), as shown in Equation 6. [Formula 6] L[t][n]←L[t-1][n]+P1_High
[0152] FIG. 8(d) is a diagram showing an example of a frequency spectrum after the reduction process B has been performed. The shaded portion 813 indicates the frequency spectrum generated by performing the reduction process B.
[0153] First, in section T801, the level of the frequency spectrum at time t+1 is not lower than the frequency spectrum at time t when the short-term noise was detected by more than the threshold P2, so reduction process B is not performed on the frequency spectrum at time t.
[0154] Furthermore, the level of the frequency spectrum at time t+2 is lower than the level of the frequency spectrum at time t+1 by at least threshold P2. Therefore, in reduction process B, the level of the frequency spectrum at time t+1 is attenuated to the level of the frequency spectrum at time t, as shown in Equation 7. The shaded portion 813 indicates the frequency spectrum reduced by reduction process B. [Formula 7] L[t+1][n]←L[t][n]
[0155] The above process is similarly performed for sections T803 and T804.
[0156] Next, in section T802, the level of the frequency spectrum at time t+1 is not lower than the frequency spectrum at time t when the short-term noise was detected by more than the threshold P2, so reduction process B is not performed on the frequency spectrum at time t.
[0157] Furthermore, the level of the frequency spectrum at time t+2 is not lower than the level of the frequency spectrum at time t+1 by the threshold P2 or more, and therefore reduction process B is not executed on the frequency spectrum at time t+1 either.
[0158] The reduction processes A and B have been described above using Fig. 8. Next, the reduction process C will be described.
[0159] 9 is a diagram showing an example of frame data L at time t and time t-1, where the vertical axis represents level and the horizontal axis represents frequency.
[0160] Here, a case where short-term noise is detected at time t will be described.
[0161] 9(a) shows an example of frame data L[t-1] of a frequency spectrum immediately before (time t-1) when short-term noise is detected. The short-term noise subtraction processing unit 211 calculates the average value L_ave[t-1] of a specific frequency band at time t-1.
[0162] FIG. 9(b) is an example of frame data L[t-1] of the frequency spectrum at the time (time t) when short-term noise is detected.
[0163] Here, the plain area (including the shaded area) indicates the level of the frequency spectrum input to the short-term noise subtraction processor 211, and the shaded area indicates the level of the frequency spectrum after reduction processing C. The short-term noise subtraction processor 211 calculates the average value L_ave[t] of a specific frequency band at time t. Here, the short-term noise subtraction processor 211 determines that L_ave[t] is greater than L_ave[t-1].
[0164] Therefore, in the subtraction process C, processing is performed so that the average value L_ave[t] approaches the average value L_ave[t-1]. In this embodiment, the short-term noise subtraction processing unit 211 calculates the ratio between the average value L_ave[t] and the average value L_ave[t-1] as shown in Equation 8, and performs processing so that the average value L_ave[t] approaches the average value L_ave[t-1] based on this ratio. [Formula 8] L[t][n]←L[t][n]×(L_ave[t-1] / L_ave[t])
[0165] The reduction process C has been described above.
[0166] In this way, the image capturing apparatus 100 can generate higher quality audio by further reducing short-term noise from the audio signal that has already undergone noise reduction based on the amount of change in noise.
[0167] <Noise parameters> 10 shows an example of noise parameters recorded in the noise parameter recording unit 206 in this embodiment. The noise parameters are parameters for correcting audio signals generated by the noise microphone 201c capturing drive sounds generated within the housing of the image capture device 100 and the housing of the optical lens 300. As shown in FIG. 10, in this embodiment, PLxA, PRxA, PLxB, and PRxB are recorded in the noise parameter recording unit 206. In this embodiment, the source of drive sounds PLxA and PRxA will be described as being within the housing of the optical lens 300. Drive sounds generated within the housing of the optical lens 300 are transmitted into the housing of the image capture device 100 via the lens mount 301 and captured by the L microphone 201a, R microphone 201b, and noise microphone 201c.
[0168] In this embodiment, a plurality of noise parameters corresponding to the type of optical lens 300 are recorded in the noise parameter recording unit 206. This is because the frequency of the drive sound differs depending on the type of optical lens 300. The imaging device 100 generates noise data using the noise parameter corresponding to the type of optical lens 300 from among these plurality of noise parameters.
[0169] Furthermore, since the frequency of the drive sound varies depending on the type of drive sound, in this embodiment, the image capture device 100 records multiple noise parameters corresponding to the type of drive sound (noise). Then, noise data is generated using one of these multiple noise parameters. In this embodiment, the image capture device 100 records noise parameters for white noise as constant noise. The image capture device 100 also records noise parameters for short-term noise generated, for example, by the meshing of gears inside the optical lens 300. The image capture device 100 also records noise parameters for long-term noise, for example, sliding noise inside the housing of the lens 300.
[0170] In this embodiment, the imaging device 100 records noise parameters for constant noise as PLxB and PRxB for each video shooting setting. Constant noise is, for example, white noise, microphone floor noise, and electrical noise. Constant noise changes depending on video shooting settings such as resolution, white balance, color, and frame rate. Note that the average value of the coefficients of PLxA and PRxA is larger than the average value of the coefficients of PLxB and PRxB, because the noise reduced by PLxA and PRxA is louder and more harsh than the noise reduced by PLxB and PRxB.
[0171] [Other Examples] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a recording medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0172] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.
Claims
1. a first microphone for acquiring environmental sounds; a second microphone for acquiring sound from a noise source including a driving unit for driving the lens; a first conversion means for Fourier-transforming the audio signal acquired by the first microphone to generate a first audio signal; a second conversion means for Fourier-transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction means for generating noise data based on the second audio signal and performing a process for reducing the noise data from the first audio signal; detection means for detecting short-term noise from the noise source based on the second audio signal; a second reduction means for controlling the magnitude of the audio signal output from the first reduction means when the short-term noise is detected by the detection means, and performing processing to reduce the short-term noise from the audio signal output from the first reduction means; a third transforming means for performing an inverse Fourier transform on the audio signal output from the second reducing means; and The second reduction means controls the volume of the audio signal of the second frame to be reduced when the volume of the audio signal of the second frame following the first frame of the audio signal output from the first reduction means is greater than the volume of the audio signal of the first frame by a threshold or more.
1. A voice processing device comprising:
2. The audio processing device described in claim 1, characterized in that when the second reduction means reduces the volume of the audio signal of the second frame, it changes the threshold from a first value to a second value smaller than the first value, and when the volume of the audio signal of a third frame following the second frame is greater than the volume of the audio signal of the second frame whose volume has been reduced by the second value or more, it reduces the volume of the audio signal of the third frame.
3. The audio processing device described in claim 1, characterized in that the second reduction means changes the volume of the audio signal of the second frame to a value obtained by adding the threshold to the audio signal of the first frame when the volume of the audio signal of the second frame is greater than the threshold by more than the threshold.
4. A first microphone for acquiring environmental sounds; a second microphone for acquiring sound from a noise source including a driving unit for driving the lens; a first conversion means for Fourier-transforming the audio signal acquired by the first microphone to generate a first audio signal; a second conversion means for Fourier-transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction means for generating noise data based on the second audio signal and performing a process for reducing the noise data from the first audio signal; detection means for detecting short-term noise from the noise source based on the second audio signal; a second reduction means for controlling the magnitude of the audio signal output from the first reduction means when the short-term noise is detected by the detection means, and performing processing to reduce the short-term noise from the audio signal output from the first reduction means; a third transforming means for performing an inverse Fourier transform on the audio signal output from the second reducing means; and The audio processing device is characterized in that the second reduction means reduces the volume of the audio signal of the second frame when the average value of the volume of a specific frequency band of the audio signal of a second frame following the first frame of the audio signal output from the first reduction means is greater than the average value of the volume of the specific frequency band of the audio signal of the first frame.
5. A first microphone for acquiring environmental sounds; a second microphone for acquiring sound from a noise source including a driving unit for driving the lens; a first conversion means for Fourier-transforming the audio signal acquired by the first microphone to generate a first audio signal; a second conversion means for Fourier-transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction means for generating noise data based on the second audio signal and performing a process for reducing the noise data from the first audio signal; detection means for detecting short-term noise from the noise source based on the second audio signal; a second reduction means for controlling the magnitude of the audio signal output from the first reduction means when the short-term noise is detected by the detection means, and performing processing to reduce the short-term noise from the audio signal output from the first reduction means; a third transforming means for performing an inverse Fourier transform on the audio signal output from the second reducing means; a third reduction means for reducing constant noise other than the noise from the noise source from the audio signal output from the second reduction means; 10. A voice processing device comprising:
6. 6. The audio processing device according to claim 1, wherein the detection means detects that the short-term noise has occurred when the magnitude of the second audio signal is equal to or greater than a predetermined value.
7. 7. The audio processing device according to claim 6, wherein the detection means detects that the short-term noise is not occurring when the magnitude of the second audio signal is equal to or greater than the predetermined value for a predetermined number of consecutive frames.
8. 8. The audio processing device according to claim 1, wherein the detection by the detection means is not performed when the drive unit is not in operation.
9. 9. The audio processing device according to claim 1, wherein the driving unit is a motor that drives the lens.
10. A control method for a sound processing device having a first microphone for acquiring environmental sound and a second microphone for acquiring sound from a noise source including a drive unit for driving a lens, the method comprising: a first transforming step of Fourier transforming the audio signal acquired by the first microphone to generate a first audio signal; a second transforming step of Fourier transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction step of generating noise data based on the second audio signal and performing a process of reducing the noise data from the first audio signal; a detecting step of detecting short-term noise from the noise source based on the second audio signal; a second reduction step of controlling the magnitude of the audio signal output from the first reduction step to reduce the short-term noise from the noise source from the audio signal output from the first reduction step when the short-term noise is detected in the detection step; a third transform step of performing an inverse Fourier transform on the audio signal output from the second reduction step; and In the second reduction step, when the magnitude of the audio signal of a second frame following the first frame of the audio signal output from the first reduction step is greater than the magnitude of the audio signal of the first frame by a threshold or more, the magnitude of the audio signal of the second frame is controlled to be reduced. A control method comprising:
11. A control method for an audio processing device having a first microphone for acquiring environmental sound and a second microphone for acquiring sound from a noise source including a drive unit for driving a lens, comprising: a first transforming step of Fourier transforming the audio signal acquired by the first microphone to generate a first audio signal; a second transforming step of Fourier transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction step of generating noise data based on the second audio signal and performing a process of reducing the noise data from the first audio signal; a detecting step of detecting short-term noise from the noise source based on the second audio signal; a second reduction step of controlling the magnitude of the audio signal output from the first reduction step to reduce the short-term noise from the noise source from the audio signal output from the first reduction step when the short-term noise is detected in the detection step; a third transform step of performing an inverse Fourier transform on the audio signal output from the second reduction step; and A control method characterized in that the second reduction step reduces the volume of the audio signal of the second frame when the average value of the volume of a specific frequency band of the audio signal of a second frame following the first frame of the audio signal output from the first reduction step is greater than the average value of the volume of the specific frequency band of the audio signal of the first frame.
12. A control method for an audio processing device having a first microphone for acquiring environmental sound and a second microphone for acquiring sound from a noise source including a drive unit for driving a lens, comprising: a first transforming step of Fourier transforming the audio signal acquired by the first microphone to generate a first audio signal; a second transforming step of Fourier transforming the audio signal acquired by the second microphone to generate a second audio signal; a first reduction step of generating noise data based on the second audio signal and performing a process of reducing the noise data from the first audio signal; a detecting step of detecting short-term noise from the noise source based on the second audio signal; a second reduction step of controlling the magnitude of the audio signal output from the first reduction step to reduce the short-term noise from the noise source from the audio signal output from the first reduction step when the short-term noise is detected in the detection step; a third transform step of performing an inverse Fourier transform on the audio signal output from the second reduction step; a third reduction step of reducing constant noise other than the noise from the noise source from the audio signal output from the second reduction step; A control method comprising:
13. A computer-readable program for causing a computer to function as each of the means of the voice processing device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Noise removing device
JP1992245300A
Photographing device, and program
JP2011097420A
Imaging apparatus, method and program
JP2011205527A
Noise reduction device
JP2014044313A
Signal processor, signal processing method, and program
JP2014085609A