Voice processing device, control method, and program

JP7686439B2Active Publication Date: 2025-06-02CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021072811
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-27
Filing Date
2021-04-22
Publication Date
2025-06-02
Estimated Expiration
2041-04-22

AI Technical Summary

Technical Problem

Existing digital cameras struggle to accurately reduce noise from optical lens driving sounds, as conventional methods like spectral subtraction fail to effectively capture the noise pattern from the lens housing, leading to incomplete noise reduction in recorded audio.

Method used

The audio processing device employs two microphones - one for environmental sound and another for noise source sound - and utilizes Fourier transforms to generate noise data, which is then subtracted from the environmental sound to reduce noise.

Benefits of technology

This approach effectively reduces noise from optical lens driving sounds, ensuring cleaner audio recordings by accurately capturing and subtracting noise patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000026_0000
    Figure 00000026_0000
  • Figure 00000026_0001
    Figure 00000026_0001
  • Figure 00000027_0000
    Figure 00000027_0000
Patent Text Reader

Abstract

To provide a speech processing device capable of effectively reducing noise contained in speech data.SOLUTION: A speech processing device (speech input part 104 of an imaging device) includes: first microphones (L microphone 201a, R microphone 201b) for acquiring environmental sound; a second microphone (noise microphone 201c) for acquiring sound from a noise source; first conversion means (FFT part 203) for generating a first speech signal by performing Fourier transformation for a speech signal from the first microphones; second conversion means (FFT part 203) for generating a second speech signal by performing Fourier transformation for a speech signal from the second microphone; generation means (noise data generation part 204) for generating noise data using the second speech signal and a parameter regarding noise of the noise source; subtraction means (subtraction processing part 207) for subtracting the noise data from the first speech signal; and third conversion means (iFFT part 208) for performing inverse Fourier transformation for a speech signal from the subtraction means.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing apparatus capable of reducing noise included in audio data.

Background Art

[0002] A digital camera, which is an example of an audio processing apparatus, can record ambient audio when recording video data. Also, the digital camera has an autofocus function that focuses on a subject during recording of video data by driving an optical lens. Further, the digital camera has a function of performing zooming by driving the optical lens during recording of video.

[0003] When the optical lens is driven during recording of video in this way, the driving sound of the optical lens may be included as noise in the audio recorded together with the video. Therefore, conventionally, when the digital camera picks up noise such as sliding noise generated when the optical lens is driven, it can reduce the noise and record ambient audio. In Patent Document 1, a digital camera that reduces noise by the spectral subtraction method is disclosed.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in Patent Document 1, since the digital camera creates a noise pattern from the noise collected by a microphone that records ambient audio, there is a possibility that an accurate noise pattern cannot be obtained from the sliding noise generated inside the housing of the optical lens. In this case, there was a possibility that the digital camera could not effectively reduce the noise included in the picked-up audio.

[0006] Therefore, the present invention aims to effectively reduce noise. [Means for solving the problem]

[0007] The audio processing device of the present invention is characterized by comprising: a first microphone for acquiring ambient sound; a second microphone for acquiring sound from a noise source; a first conversion means for generating a first audio signal by performing a Fourier transform on the audio signal from the first microphone; a second conversion means for generating a second audio signal by performing a Fourier transform on the audio signal from the second microphone; a generation means for generating noise data using the second audio signal and parameters related to the noise of the noise source; a subtraction means for subtracting the noise data from the first audio signal; and a third conversion means for performing an inverse Fourier transform on the audio signal from the subtraction means. [Effects of the Invention]

[0008] The audio processing device of the present invention can effectively reduce noise. [Brief explanation of the drawing]

[0009] [Figure 1] This is a perspective view of the imaging device in the first embodiment. [Figure 2] This is a block diagram showing the configuration of the imaging device in the first embodiment. [Figure 3] This is a block diagram showing the configuration of the audio input section of the imaging device in the first embodiment. [Figure 4] This figure shows the arrangement of microphones in the audio input section of the imaging device in the first embodiment. [Figure 5] This figure shows the noise parameters in the first embodiment. [Figure 6] This figure shows the frequency spectrum of the audio and the frequency spectrum of the noise parameter in the first embodiment when driving noise occurs in a situation where ambient noise can be considered absent. [Figure 7]This figure shows the frequency spectrum of the sound when an operating sound is generated in the presence of ambient noise in the first embodiment. [Figure 8] This is a block diagram showing the configuration of the noise parameter selection unit in the first embodiment. [Figure 9] This is a timing chart related to the audio noise reduction process in the first embodiment. [Figure 10] This is a block diagram showing the configuration of the audio input section of the imaging device in the second embodiment. [Figure 11] This is a block diagram showing the configuration of the audio input section of the imaging device in the third embodiment. [Figure 12] This figure shows the noise parameters in the third embodiment. [Figure 13] This is a timing chart related to the audio noise reduction process in the third embodiment. [Modes for carrying out the invention]

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0011] [First example] <External view of imaging device 100> Figures 1(a) and 1(b) show an example of an external view of an imaging device 100, which is an example of an audio processing device to which the present invention can be applied. Figure 1(a) is an example of a front perspective view of the imaging device 100. Figure 1(b) is an example of a rear perspective view of the imaging device 100. In Figure 1, an optical lens (not shown) is mounted on the lens mount 301.

[0012] The display unit 107 displays image data, character information, etc. The display unit 107 is provided on the back surface of the imaging device 100. The external finder display unit 43 is a display unit provided on the upper surface of the imaging device 100. The external finder display unit 43 displays the set values of the imaging device 100 such as the shutter speed and aperture value. The eyepiece finder 16 is a peephole type finder. The user can confirm the focus and composition of the optical image of the subject by observing the focusing screen inside the eyepiece finder 16.

[0013] The release switch 61 is an operation member for the user to give a shooting instruction. The mode switch 60 is an operation member for the user to switch various modes. The main electronic dial 71 is a rotary operation member. The user can change the set values of the imaging device 100 such as the shutter speed and aperture value by turning the main electronic dial 71. The release switch 61, the mode switch 60, and the main electronic dial 71 are included in the operation unit 112.

[0014] The power switch 72 is an operation member for switching on and off the power of the imaging device 100. The sub - electronic dial 73 is a rotary operation member. The user can move the selection frame displayed on the display unit 107 and perform image advancement in the playback mode using the sub - electronic dial 73. The cross - key 74 is a cross - key (4 - direction key) where the upper, lower, left, and right parts can be pushed in respectively. The imaging device 100 executes processing according to the pressed part (direction) of the cross - key 74. The power switch 72, the sub - electronic dial 73, and the cross - key 74 are included in the operation unit 112.

[0015] The SET button 75 is a push button. The SET button 75 is mainly used for the user to determine the selected item displayed on the display unit 107 and the like. The LV button 76 is a button used to switch between on and off of the live view (hereinafter, LV). The LV button 76 is used to give instructions for starting and stopping video shooting (recording) in the video recording mode. The zoom button 77 is a push button for turning on and off the zoom mode and changing the magnification during the zoom mode in the live view display of the shooting mode. The SET button 75, the LV button 76, and the zoom button 77 are included in the operation unit 112.

[0016] The zoom button 77 functions as a button for increasing the magnification of the image data displayed on the display unit 107 in the playback mode. The reduction button 78 is a button for reducing the magnification of the image data enlarged and displayed on the display unit 107. The playback button 79 is an operation button for switching between the shooting mode and the playback mode. When the user presses the playback button 79 during the shooting mode of the imaging device 100, the imaging device 100 shifts to the playback mode and displays the image data recorded on the recording medium 110 on the display unit 107. The reduction button 78 and the playback button 79 are included in the operation unit 112.

[0017] The quick return mirror 12 (hereinafter, mirror 12) is a mirror for switching the incident light beam from the optical lens mounted on the imaging device 100 to either the eyepiece finder 16 side or the imaging unit 101 side. The mirror 12 is moved up and down by controlling an actuator (not shown) by the control unit 111 during exposure, live view shooting, and video shooting. The mirror 12 is normally arranged to make the light beam incident on the eyepiece finder 16. The mirror 12 jumps upward (mirror up) so that the light beam is incident on the imaging unit 101 when shooting is performed and when live view display is performed. Also, the central portion of the mirror 12 is a half mirror. A part of the light beam transmitted through the central portion of the mirror 12 is incident on a focus detection unit (not shown) for performing focus detection.

[0018] The communication terminal 10 is a communication terminal for communication between the optical lens 300 mounted on the imaging device 100 and the imaging device 100. The terminal cover 40 is a cover that protects connectors (not shown) such as connection cables that connect external devices to the imaging device 100. The lid 41 is a lid for the slot that houses the recording medium 110. The lens mount 301 is a mounting part on which an optical lens 300 (not shown) can be attached.

[0019] The left microphone 201a and the right microphone 201b are microphones for capturing the user's voice and other sounds. When viewed from the rear of the imaging device 100, the left microphone 201a is positioned on the left and the right microphone 201b is positioned on the right.

[0020] <Configuration of imaging device 100> Figure 2 is a block diagram showing an example of the configuration of the imaging device 100 in this embodiment.

[0021] The optical lens 300 is a lens unit that can be attached to and detached from the imaging device 100. For example, the optical lens 300 is a zoom lens or a varifocal lens. The optical lens 300 has an optical lens, a motor for driving the optical lens, and a communication unit that communicates with the lens control unit 102 of the imaging device 100, which will be described later. Based on the control signals received by the communication unit, the optical lens 300 moves the optical lens with the motor, thereby enabling focusing and zooming on the subject, as well as correcting camera shake.

[0022] The imaging unit 101 includes an image sensor for converting an optical image of a subject formed on the imaging surface via an optical lens 300 into an electrical signal, and an image processing unit for generating and outputting image data or video data from the electrical signal generated by the image sensor. The image sensor is, for example, a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor). In this embodiment, the series of processes in which the imaging unit 101 generates image data including still image data and video data and outputs it from the imaging unit 101 is called "shooting". In the imaging device 100, the image data is recorded on a recording medium 110, which will be described later, in accordance with the DCF (Design rule for Camera File system) standard.

[0023] The lens control unit 102 transmits control signals to the optical lens 300 via the communication terminal 10 based on the data output from the imaging unit 101 and the control signals output from the control unit 111 (described later), thereby controlling the optical lens 300.

[0024] The information acquisition unit 103 detects the tilt of the imaging device 100 and the temperature inside the housing of the imaging device 100. For example, the information acquisition unit 103 detects the tilt of the imaging device 100 using an acceleration sensor or a gyro sensor. Also, for example, the information acquisition unit 103 detects the temperature inside the housing of the imaging device 100 using a temperature sensor.

[0025] The audio input unit 104 generates audio data from the sound acquired by the microphone. The audio input unit 104 acquires sound from the surrounding area of ​​the imaging device 100 using the microphone, performs analog-to-digital conversion (A / D conversion) and various audio processing on the acquired sound, and generates audio data. In this embodiment, the audio input unit 104 has a microphone. A detailed configuration example of the audio input unit 104 will be described later.

[0026] The volatile memory 105 temporarily stores image data generated by the imaging unit 101 and audio data generated by the audio input unit 104. The volatile memory 105 is also used as a temporary recording area for image data displayed on the display unit 107 and as a work area for the control unit 111.

[0027] The display control unit 106 controls the display unit 107 to display image data output from the imaging unit 101, interactive operation characters, and menu screens. Furthermore, during still image and video recording, the display control unit 106 controls the display unit 107 to sequentially display the digital data output from the imaging unit 101, thereby enabling the display unit 107 to function as an electronic viewfinder. For example, the display unit 107 is a liquid crystal display or an organic EL display. The display control unit 106 can also control the display of image data and video data output from the imaging unit 101, interactive operation characters, and menu screens on an external display via an external output unit 115, which will be described later.

[0028] The encoding processing unit 108 can encode image data and audio data temporarily recorded in the volatile memory 105, respectively. For example, the encoding processing unit 108 can generate video data encoded and compressed according to the JPEG standard or RAW image format for image data. For example, the encoding processing unit 108 can generate video data encoded and compressed according to the MPEG2 standard or H.264 / MPEG4-AVC standard for video data. Also, for example, the encoding processing unit 108 can generate audio data encoded and compressed according to the AC3AAC standard, ATRAC standard, or ADPCM method for audio data. Furthermore, the encoding processing unit 108 may encode audio data without data compression, for example, according to the linear PCM method.

[0029] The recording control unit 109 can record data onto the recording medium 110 and read data from the recording medium 110. For example, the recording control unit 109 can record still image data, video data, and audio data generated by the encoding processing unit 108 onto the recording medium 110 and read data from the recording medium 110. The recording medium 110 is, for example, an SD card, a CF card, an XQD memory card, an HDD (magnetic disk), an optical disk, and a semiconductor memory. The recording medium 110 may be configured to be detachable from the imaging device 100, or it may be built into the imaging device 100. In other words, the recording control unit 109 only needs to have means to access the recording medium 110.

[0030] The control unit 111 controls each component of the imaging device 100 via the data bus 116 according to the input signals and the program described later. The control unit 111 has a CPU, ROM, and RAM for performing various controls. Alternatively, instead of the control unit 111 controlling the entire imaging device 100, multiple hardware components may share the task of controlling the entire imaging device. The ROM of the control unit 111 stores programs for controlling each component. The RAM of the control unit 111 is a volatile memory used for calculations and other processing.

[0031] The control unit 112 is a user interface for receiving instructions from the user for the imaging device 100. The control unit 112 includes, for example, a power switch 72 for turning the imaging device 100 on or off, a release switch 61 for instructing shooting, a playback button for instructing playback of image data or video data, and a mode switching switch 60.

[0032] The operation unit 112 outputs control signals to the control unit 111 in response to user operations. A touch panel formed on the display unit 107 can also be included in the operation unit 112. The release switch 61 has SW1 and SW2. When the release switch 61 is half-pressed, SW1 turns on. This allows the system to receive preparation instructions for image capture preparation operations such as AF (autofocus), AE (automatic exposure), AWB (auto white balance), and EF (flash pre-flash). When the release switch 61 is fully pressed, SW2 turns on. This allows the system to receive imaging instructions for image capture operations. The operation unit 112 also includes an operating element (e.g., a button) that can adjust the volume of audio data played back from the speaker 114, which will be described later.

[0033] The audio output unit 113 can output audio data to the speaker 114 and the external output unit 115. The audio data input to the audio output unit 113 consists of audio data read from the recording medium 110 by the recording control unit 109, audio data output from the non-volatile memory 117, and audio data output from the encoding processing unit. The speaker 114 is an electroacoustic converter capable of reproducing audio data.

[0034] The external output unit 115 can output image data, video data, and audio data to external devices. The external output unit 115 is composed of, for example, a video terminal, a microphone terminal, and a headphone terminal.

[0035] The data bus 116 is a data bus for transmitting various types of data, such as audio data, video data, and image data, as well as various control signals, to each block of the imaging device 100.

[0036] The non-volatile memory 117 is a non-volatile memory that stores programs and other data executed by the control unit 111, as described later. The non-volatile memory 117 also stores audio data. This audio data includes, for example, the focus sound output when the subject is in focus, the electronic shutter sound output when photography is instructed, and the operation sound output when the imaging device 100 is operated.

[0037] <Operation of imaging device 100> Next, I will explain the operation of the imaging device 100 in this embodiment.

[0038] In this embodiment, the imaging device 100 receives power from a power source (not shown) to each component of the imaging device when the user turns on the power switch 72. For example, the power source is a battery such as a lithium-ion battery or an alkaline manganese dry cell.

[0039] The control unit 111 determines, for example, whether to operate in shooting mode or playback mode, based on the state of the mode selector switch 60 in response to the supply of power. In video recording mode, the control unit 111 records the video data output from the imaging unit 101 and the audio data output from the audio input unit 104 as a single video data with audio. In playback mode, the control unit 111 controls the recording control unit 109 to read the image data or video data recorded on the recording medium 110 and display it on the display unit 107.

[0040] First, let's explain the video recording mode. In video recording mode, the control unit 111 first sends control signals to each component of the imaging device 100 to switch the imaging device 100 into a shooting standby state. For example, the control unit 111 controls the imaging unit 101 and the audio input unit 104 to perform the following operations.

[0041] The imaging unit 101 converts the optical image of the subject, which is formed on the imaging surface via the optical lens 300, into an electrical signal, and generates video data from the electrical signal generated by the image sensor. The imaging unit 101 then transmits the video data to the display control unit 106, which displays it on the display unit 107. The user can prepare to shoot while viewing the video data displayed on the display unit 107.

[0042] The audio input unit 104 converts the analog audio signals input from multiple microphones using A / D conversion to generate multiple digital audio signals. The audio input unit 104 then generates audio data for multiple channels from these multiple digital audio signals. The audio input unit 104 transmits the generated audio data to the audio output unit 113, which plays the audio data from the speaker 114. The user can adjust the volume of the audio data recorded in the audio-enabled video data using the control unit 112 while listening to the audio data played from the speaker 114.

[0043] Next, in response to the user pressing the LV button 76, the control unit 111 sends an instruction signal to start shooting to each component of the imaging device 100. For example, the control unit 111 controls the imaging unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.

[0044] The imaging unit 101 converts the optical image of the subject formed on the imaging surface via the optical lens 300 into an electrical signal, and generates video data from the electrical signal generated by the image sensor. The imaging unit 101 then transmits the video data to the display control unit 106, which displays it on the display unit 107. The imaging unit 101 also transmits the generated video data to the volatile memory 105.

[0045] The audio input unit 104 converts the analog audio signals input from multiple microphones into digital signals, generating multiple digital audio signals. The audio input unit 104 then generates multi-channel audio data from these multiple digital audio signals. Finally, the audio input unit 104 transmits the generated audio data to the volatile memory 105.

[0046] The encoding processing unit 108 reads the video data and audio data temporarily recorded in the volatile memory 105 and encodes them, respectively. The control unit 111 generates a data stream from the video data and audio data encoded by the encoding processing unit 108 and outputs it to the recording control unit 109. The recording control unit 109 records the input data stream as video data with audio on the recording medium 110 according to a file system such as UDF or FAT.

[0047] Each component of the imaging device 100 continues the above operations during video recording.

[0048] Then, in response to the user pressing the LV button 76, the control unit 111 sends an instruction signal to each component of the imaging device 100 to complete the shooting. For example, the control unit 111 controls the imaging unit 101, the audio input unit 104, the encoding processing unit 108, and the recording control unit 109 to perform the following operations.

[0049] The imaging unit 101 stops generating video data. The audio input unit 104 stops generating audio data.

[0050] The encoding processing unit 108 reads the remaining video and audio data recorded in the volatile memory 105 and encodes it. The control unit 111 generates a data stream from the video and audio data encoded by the encoding processing unit 108 and outputs it to the recording control unit 109.

[0051] The recording control unit 109 records the data stream as a video data file with sound on the recording medium 110 according to a file system such as UDF or FAT. Then, the recording control unit 109 completes the video data with sound in response to the cessation of data stream input. Upon completion of the video data with sound, the recording operation of the imaging device 100 stops.

[0052] The control unit 111 transmits control signals to each component of the imaging device 100 to transition to the shooting standby state when the recording operation stops. This allows the control unit 111 to control the imaging device 100 to return to the shooting standby state.

[0053] Next, the playback mode will be described. In playback mode, the control unit 111 transmits control signals to each component of the imaging device 100 to transition to the playback state. For example, the control unit 111 controls the encoding processing unit 108, the recording control unit 109, the display control unit 106, and the audio output unit 113 to perform the following operations.

[0054] The recording control unit 109 reads the video data with audio recorded on the recording medium 110 and transmits the read video data with audio to the encoding processing unit 108.

[0055] The encoding processing unit 108 decodes image data and audio data from the video data with audio. The encoding processing unit 108 transmits the decoded video data to the display control unit 106 and the decoded audio data to the audio output unit 113.

[0056] The display control unit 106 displays the decoded image data using the display unit 107. The audio output unit 113 plays the decoded audio data using the speaker 114.

[0057] As described above, the imaging device 100 of this embodiment can record and reproduce image data and audio data.

[0058] In this embodiment, the audio input unit 104 performs audio processing, such as adjusting the level of the audio signal input from the microphone. In this embodiment, the audio input unit 104 performs this audio processing in response to the start of video recording. This audio processing may also be performed after the power of the imaging device 100 is turned on. Furthermore, this audio processing may be performed in response to the selection of a shooting mode. Furthermore, this audio processing may be performed in response to the selection of a mode related to audio recording, such as a video recording mode or an audio memo function. Furthermore, this audio processing may be performed in response to the start of recording of an audio signal.

[0059] <Configuration of the voice input unit 104> Figure 3 is a block diagram showing an example of a detailed configuration of the audio input unit 104 in this embodiment.

[0060] In this embodiment, the audio input unit 104 has three microphones: an L microphone 201a, an R microphone 201b, and a noise microphone 201c. The L microphone 201a and the R microphone 201b are examples of first microphones. In this embodiment, the imaging device 100 picks up ambient sounds using the L microphone 201a and the R microphone 201b, and records the audio signals input from the L microphone 201a and the R microphone 201b in stereo. Ambient sounds include sounds generated outside the housing of the imaging device 100 and the housing of the optical lens 300, such as the user's voice, animal sounds, rain sounds, and music.

[0061] Furthermore, the noise microphone 201c is an example of a second microphone. The noise microphone 201c is a microphone for acquiring noise such as drive sounds from predetermined noise sources that occur inside the housing of the imaging device 100 and the housing of the optical lens 300. The noise sources are, for example, motors such as an ultrasonic motor (USM) and a stepper motor (STM). The noise is, for example, vibration sound generated by the driving of motors such as USM and STM. For example, the motor is driven in the AF processing to focus on the subject. The imaging device 100 acquires noise such as drive sounds generated inside the housing of the imaging device 100 and the housing of the optical lens 300 using the noise microphone 201c, and generates noise parameters, which will be described later, using the acquired noise audio data. In this embodiment, the L microphone 201a, the R microphone 201b, and the noise microphone 201c are omnidirectional microphones. An example of the arrangement of the L microphone 201a, R microphone 201b, and noise microphone 201c in this embodiment will be described later with reference to Figure 4.

[0062] The left microphone 201a, the right microphone 201b, and the noise microphone 201c each generate an analog audio signal from the acquired audio and input it to the A / D conversion unit 202. Here, the audio signal input from the left microphone 201a is denoted as Lch, the audio signal input from the right microphone 201b is denoted as Rch, and the audio signal input from the noise microphone 201c is denoted as Nch.

[0063] The A / D conversion unit 202 converts the analog audio signals input from the left microphone 201a, the right microphone 201b, and the noise microphone 201c into digital audio signals. The A / D conversion unit 202 outputs the converted digital audio signals to the FFT unit 203. In this embodiment, the A / D conversion unit 202 converts the analog audio signals into digital audio signals by performing sampling processing with a sampling frequency of 48kHz and a bit depth of 16 bits.

[0064] The FFT unit 203 performs a Fast Fourier Transform (FFT) on the time-domain digital audio signal input from the A / D conversion unit 202, converting it into a frequency-domain digital audio signal. In this embodiment, the frequency-domain digital audio signal has a frequency spectrum of 1024 points in the frequency band from 0 Hz to 48 kHz. Furthermore, the frequency-domain digital audio signal has a frequency spectrum of 513 points in the frequency band from 0 Hz to the Nyquist frequency of 24 kHz. In this embodiment, the imaging device 100 uses the 513-point frequency spectrum from 0 Hz to 24 kHz of the audio data output from the FFT unit 203 to perform noise reduction processing.

[0065] Here, the frequency spectrum of the Lch channel, after being Fast Fourier Transformed, is represented by a 513-point array of data, Lch_Before[0] to Lch_Before

[0512] . When referring to these array data collectively, we will use the term Lch_Before. Similarly, the frequency spectrum of the Rch channel, after being Fast Fourier Transformed, is represented by a 513-point array of data, Rch_Before[0] to Rch_Before

[0512] . When referring to these array data collectively, we will use the term Rch_Before. Note that Lch_Before and Rch_Before are examples of the first frequency spectrum data, respectively.

[0066] Furthermore, the Fast Fourier Transformed frequency spectrum of Nch is represented by a 513-point array of data, Nch_Before[0] to Nch_Before

[0512] . When referring to these array data collectively, it is written as Nch_Before. Note that Nch_Before is an example of the second frequency spectrum data.

[0067] The noise data generation unit 204 generates data to reduce the noise contained in Lch_Before and Rch_Before based on Nch_Before. In this embodiment, the noise data generation unit 204 generates array data NL[0] to NL

[0512] to reduce the noise contained in Lch_Before[0] to Lch_Before

[0512] using noise parameters. The noise data generation unit 204 also generates array data NR[0] to NR

[0512] to reduce the noise contained in Rch_Before[0] to Rch_Before

[0512] . The frequency points in the array data NL[0] to NL

[0512] are the same as the frequency points in the array data Lch_Before[0] to Lch_Before

[0512] . Furthermore, the frequency points in the array data NR[0]~NR

[0512] are the same as the frequency points in the array data Rch_Before[0]~Rch_Before

[0512] .

[0068] Note that when referring to the sequence data from NL[0] to NL

[0512] collectively, it will be written as NL. Similarly, when referring to NR[0] to NR

[0512] collectively, it will be written as NR. NL and NR are examples of third frequency spectrum data, respectively.

[0069] The noise parameter recording unit 205 records noise parameters used by the noise data generation unit 204 to generate NL and NR from Nch_Before. The noise parameter recording unit 205 records multiple types of noise parameters depending on the type of noise. When referring to the noise parameters used to generate NL from Nch_Before collectively, it is written as PLx. When referring to the noise parameters used to generate NR from Nch_Before collectively, it is written as PRx.

[0070] PLx and PRx have the same number of elements as NL and NR, respectively. For example, PL1 is the element data from PL1[0] to PL1

[0512] . Also, the frequency points of PL1 are the same as the frequency points of Lch_Before. Similarly, PR1 is the element data from PR1[0] to PR1

[0512] . The frequency points of PR1 are the same as the frequency points of Rch_Before. The noise parameters will be described later using Figure 5.

[0071] The noise parameter selection unit 206 determines the noise parameters to be used in the noise data generation unit 204 from the noise parameters recorded in the noise parameter recording unit 205. The noise parameter selection unit 206 determines the noise parameters to be used in the noise data generation unit 204 based on Lch_Before, Rch_Before, Nch_Before, and data received from the lens control unit 102. The operation of the noise parameter selection unit 206 will be described in detail later with reference to Figure 8.

[0072] In this embodiment, the noise parameter recording unit 205 records all coefficients for each of the 513 frequency spectrum points as noise parameters. However, the noise parameter recording unit 205 only needs to record the coefficients for at least the frequency points necessary to reduce noise, rather than the coefficients for all 513 frequencies. For example, the noise parameter recording unit 205 may record the coefficients for each of the frequency spectrum points from 20 Hz to 20 kHz, which are considered typical audible frequencies, as noise parameters, and may not need to record the coefficients for other frequency spectrum points. Also, for example, coefficients for frequency spectrum points with a coefficient value of zero do not need to be recorded in the noise parameter recording unit 205 as noise parameters.

[0073] The subtraction processing unit 207 subtracts NL and NR from Lch_Before and Rch_Before, respectively. For example, the subtraction processing unit 207 has an L subtractor 207a that subtracts NL from Lch_Before, and an R subtractor 207b that subtracts NR from Rch_Before. The L subtractor 207a subtracts NL from Lch_Before and outputs 513 points of sequence data from Lch_After[0] to Lch_After

[0512] . The R subtractor 207b subtracts NR from Rch_Before and outputs 513 points of sequence data from Rch_After[0] to Rch_After

[0512] . In this embodiment, the subtraction processing unit 207 performs the subtraction process using the spectral subtraction method.

[0074] The iFFT unit 208 converts the frequency-domain digital audio signal input from the subtraction unit 207 into a time-domain digital audio signal by performing an inverse fast Fourier transform (inverse Fourier transform).

[0075] The audio processing unit 209 performs time-domain audio processing on the digital audio signal, such as equalization, auto-level control, and stereo enhancement processing. The audio processing unit 209 outputs the processed audio data to the volatile memory 105.

[0076] In this embodiment, the imaging device 100 has two microphones as the first microphones, but the imaging device 100 may have one microphone or three or more microphones as the first microphones. For example, if the imaging device 100 has one microphone as the first microphone in the audio input unit 104, it will record the audio data picked up by the one microphone in a monaural manner. Alternatively, if the imaging device 100 has three or more microphones as the first microphones in the audio input unit 104, it will record the audio data picked up by the three or more microphones in a surround sound manner.

[0077] In this embodiment, the L microphone 201a, R microphone 201b, and noise microphone 201c are omnidirectional microphones, but these microphones may also be directional microphones.

[0078] <Positioning of microphones in the audio input unit 104> Here, we will describe an example of the arrangement of microphones in the audio input unit 104 of this embodiment. Figure 4 shows an example of the arrangement of the left microphone 201a, the right microphone 201b, and the noise microphone 201c.

[0079] Figure 4 is an example of a cross-sectional view of the imaging device 100 to which the L microphone 201a, R microphone 201b, and noise microphone 201c are attached. This part of the imaging device 100 consists of an outer casing 302, a microphone bush 303, and a fixing part 304.

[0080] The exterior part 302 has a hole (hereinafter referred to as the microphone hole) for inputting ambient sound into the microphone. In this embodiment, the microphone hole is formed above the L microphone 201a and the R microphone 201b. On the other hand, the noise microphone 201c is provided to acquire driving noise generated inside the housing of the imaging device 100 and the housing of the optical lens 300, and does not need to acquire ambient sound. Therefore, in this embodiment, the exterior part 302 does not have a microphone hole above the noise microphone 201c.

[0081] The drive noise generated inside the housing of the imaging device 100 and the housing of the optical lens 300 is acquired by the L microphone 201a and the R microphone 201b via microphone holes. When ambient noise is low and drive noise is generated inside the housings of the imaging device 100 and the optical lens 300, the sound acquired by each microphone will mainly be this drive noise. Therefore, the sound level from the noise microphone 201c is higher than the sound levels from the L microphone 201a and the R microphone 201b. In other words, in this case, the relationship between the levels of the audio signals output from each microphone is as follows. Lch ≈ Rch <Nch Furthermore, when ambient noise increases, the audio level of ambient noise from microphones L 201a and R 201b becomes higher than the audio level of the drive noise generated by the imaging device 100 or optical lens 300 from microphone 201c. Therefore, in this case, the relationship between the levels of the audio signals output from each microphone is as follows. Lch ≈ Rch > Nch In this embodiment, the shape of the microphone hole formed in the outer casing 302 is elliptical, but it may be circular, rectangular, or any other shape. Also, the shape of the microphone hole on microphone 201a and the shape of the microphone hole on microphone 201b may be different from each other.

[0082] In this embodiment, the noise microphone 201c is positioned close to the L microphone 201a and the R microphone 201b. Also, in this embodiment, the noise microphone 201c is positioned between the L microphone 201a and the R microphone 201b. As a result, the audio signal generated by the noise microphone 201c from the driving noise etc. generated inside the housing of the imaging device 100 and the housing of the optical lens 300 will be similar to the audio signal generated by the L microphone 201a and the R microphone 201b from the same driving noise etc.

[0083] The microphone bush 303 is a component for fixing the left microphone 201a, the right microphone 201b, and the noise microphone 201c. The fixing part 304 is a component for fixing the microphone bush 303 to the outer casing 302.

[0084] In this embodiment, the exterior part 302 and the fixing part 304 are made of molded material such as PC material. Alternatively, the exterior part 302 and the fixing part 304 may be made of metal material such as aluminum or stainless steel. In this embodiment, the microphone bush 303 is made of rubber material such as ethylene propylene diene rubber.

[0085] <Noise parameters> Figure 5 shows an example of noise parameters recorded in the noise parameter recording unit 205. Noise parameters are parameters for correcting the audio signal generated by the noise microphone 201c acquiring the drive sound generated inside the housing of the imaging device 100 and the housing of the optical lens 300. As shown in Figure 5, in this embodiment, PLx and PRx are recorded in the noise parameter recording unit 205. In this embodiment, the source of the drive sound is assumed to be inside the housing of the optical lens 300. The drive sound generated inside the housing of the optical lens 300 is transmitted to the housing of the imaging device 100 via the lens mount 301 and acquired by the L microphone 201a, the R microphone 201b, and the noise microphone 201c.

[0086] The frequency of the drive noise differs depending on the type of drive noise. Therefore, in this embodiment, the imaging device 100 records multiple noise parameters corresponding to the type of drive noise (noise). Then, noise data is generated using one of these multiple noise parameters. In this embodiment, the imaging device 100 records noise parameters for white noise as a constant noise. The imaging device 100 also records noise parameters for short-term noise generated, for example, by the meshing of gears in the optical lens 300. Furthermore, the imaging device 100 records noise parameters for long-term noise, for example, sliding noise within the housing of the lens 300. In addition, the imaging device 100 may record noise parameters for each type of optical lens 300, as well as for the temperature inside the housing of the imaging device 100 and the tilt of the imaging device 100 as detected by the information acquisition unit 103.

[0087] <Method for generating noise data> The noise data generation process in the noise data generation unit 204 will be explained using Figures 6 and 7. While this explanation focuses on the noise data generation process for the Lch, the method for generating noise data for the Rch is similar.

[0088] First, we will explain the process of generating noise parameters in a situation where ambient noise can be considered absent. Figure 6(a) is an example of the frequency spectrum of Lch_Before when driving noise occurs inside the housing of the optical lens 300 in a situation where ambient noise can be considered absent. Figure 6(b) is an example of the frequency spectrum of Nch_Before when driving noise occurs inside the housing of the optical lens 300 in a situation where ambient noise can be considered absent. The horizontal axis represents the frequency from point 0 to point 512, and the vertical axis represents the amplitude of the frequency spectrum.

[0089] Since the situation can be considered as having no ambient noise, the amplitude of the frequency spectrum in the same frequency band is larger in Lch_Before and Nch_Before. Also, because driving noise is generated inside the housing of the optical lens 300, the amplitude of each frequency spectrum for the same driving noise tends to be larger in Nch_Before than in Lch_Before.

[0090] Figure 6(c) shows an example of PLx in this embodiment. In this embodiment, PLx is a coefficient of each frequency spectrum calculated by dividing the amplitude of each frequency spectrum of Lch_Before by the amplitude of each frequency spectrum of Nch_Before. The result of this division is denoted as Lch_Before / Nch_Before. That is, PLx is the ratio of the amplitudes of Lch_Before and Nch_Before. The noise parameter recording unit 205 records the value of Lch_Before / Nch_Before as the noise parameter PLx. As mentioned above, the amplitude of the frequency spectrum for the same drive sound tends to be larger for Nch_Before than for Lch_Before, so the values ​​of each coefficient of the noise parameter PLx tend to be less than 1. However, if the value of Nch_Before[n] is smaller than a predetermined threshold, the noise parameter recording unit 205 records the noise parameter PLx as PLx[n]=0.

[0091] Next, we will explain the process of applying the generated noise parameters to Nch_Before. Figure 7(a) is an example of the frequency spectrum of Lch_Before when driving noise is generated inside the housing of the optical lens 300 in the presence of ambient noise. Figure 7(b) is an example of the frequency spectrum of Nch_Before when driving noise is generated inside the housing of the optical lens 300 in the presence of ambient noise. The horizontal axis represents the frequency from point 0 to point 512, and the vertical axis represents the amplitude of the frequency spectrum.

[0092] Figure 7(c) shows an example of NL when driving noise is generated inside the housing of the optical lens 300 in the presence of ambient noise. The noise data generation unit 204 generates NL by multiplying each frequency spectrum of Nch_Before by each coefficient of PLx. NL is the frequency spectrum thus generated.

[0093] Figure 7(d) shows an example of Lch_After when driving noise is generated inside the housing of the optical lens 300 in the presence of ambient noise. The subtraction processing unit 207 subtracts NL from Lch_Before to generate Lch_After. Lch_After is the frequency spectrum generated in this way.

[0094] As a result, the imaging device 100 can reduce noise caused by the drive sound inside the housing of the optical lens 300 and record ambient sounds with less noise.

[0095] <Explanation of Noise Parameter Selection Unit 206> Figure 8 is a block diagram showing an example of a detailed configuration of the noise parameter selection unit 206.

[0096] The noise parameter selection unit 206 receives Lch_Before, Rch_Before, Nch_Before, and lens control signals as inputs.

[0097] The Nch noise detection unit 2061 detects noise caused by driving sounds generated within the housing of the optical lens 300 from Nch_Before. Based on the noise detection result, the Nch noise detection unit 2061 outputs data related to the noise detection result to the noise determination unit 2063. In this embodiment, the Nch noise detection unit 2061 uses 513 points of data from Nch_Before to detect noise.

[0098] The ambient sound detection unit 2062 detects the ambient sound level from Lch_Before and Rch_Before. Based on the detection result of the ambient sound level, the ambient sound detection unit 2062 outputs data related to the detection result of the ambient sound level to the noise determination unit 2063.

[0099] The noise determination unit 2063 determines the noise parameters to be used by the noise data generation unit 204 based on the lens control signal input from the lens control unit 102, the data input from the Nch noise detection unit 2061, and the data input from the ambient sound detection unit 2062. The noise determination unit 2063 outputs data indicating the type of noise parameter determined to the noise data generation unit 204.

[0100] The Nch differentiation unit 2064 performs differentiation on Nch_Before. The Nch differentiation unit 2064 outputs data showing the result of differentiating Nch_Before to the short-term noise detection unit 2065. The short-term noise detection unit 2065 detects whether or not short-term noise is occurring based on the data input from the Nch differentiation unit 2064. The short-term noise detection unit 2065 outputs data indicating whether or not short-term noise is occurring to the noise determination unit 2063. Note that the Nch differentiation unit 2064 and the short-term noise detection unit 2065 are included in the Nch noise detection unit 2061.

[0101] The Nch integration unit 2066 performs integration on Nch_Before. The Nch integration unit 2066 outputs data showing the result of differentiating Nch_Before to the long-term noise detection unit 2067. The long-term noise detection unit 2067 detects whether or not long-term noise is occurring based on the data input from the Nch integration unit 2066. The long-term noise detection unit 2067 outputs data indicating whether or not long-term noise is occurring to the noise determination unit 2063. Note that the Nch integration unit 2066 and the long-term noise detection unit 2067 are included in the Nch noise detection unit 2061.

[0102] The ambient sound extraction unit 2068 extracts ambient sounds. In this embodiment, the ambient sound extraction unit 2068 extracts data at frequencies where the noise has little effect, based on a noise parameter. For example, the ambient sound extraction unit 2068 extracts data at frequencies where the noise parameter is below a predetermined value. Then, based on the extracted frequency data, the ambient sound extraction unit 2068 outputs data indicating the magnitude of the ambient sound. The ambient sound extraction unit 2068 is included in the ambient sound detection unit 2062.

[0103] The ambient sound determination unit 2069 determines the magnitude of ambient sound. The ambient sound determination unit 2069 inputs data indicating the magnitude of the determined ambient sound to the Nch noise detection unit 2061 and the noise determination unit 2063. The Nch noise detection unit 2061 changes the first threshold and the second threshold, described later, based on the data indicating the magnitude of ambient sound input from the ambient sound determination unit 2069. The ambient sound determination unit 2069 is included in the ambient sound detection unit 2062.

[0104] <Timing chart for noise reduction processing> The noise reduction process in this embodiment will be explained with reference to Figure 9.

[0105] Figures 9(a) to (i) are examples of timing charts for audio processing in the noise data generation unit 204, the noise parameter selection unit 206, and the subtraction processing unit 207. In this embodiment, for the sake of simplicity, the audio processing of the Lch will be described, but the audio processing of the Rch is similar. The horizontal axis of the graphs in Figures 9(a) to (i) is the time axis.

[0106] Figure 9(a) shows an example of a lens control signal. The lens control signal is a signal that instructs the lens control unit 102 to drive the optical lens 300. In this embodiment, the level of the lens control signal is represented by two values: High and Low. When the level of the lens control signal is High, the lens control unit 102 is instructing the optical lens 300 to drive. When the level of the lens control signal is Low, the lens control unit 102 is not instructing the optical lens 300 to drive.

[0107] Figure 9(b) is a graph showing an example of the value of Lch_Before[n]. The vertical axis represents the value of Lch_Before[n]. In this embodiment, Lch_Before[n] is the signal at the nth frequency point among the Lch_Before signals output from the FFT unit 203, in which the signal indicating the driving sound of the optical lens 300 is characteristically present. In this embodiment, the signal at the nth frequency point is described, but audio processing is performed similarly for other frequencies. Also, the signals X and Y are signals that contain noise. In this embodiment, signal X represents a signal that contains short-term noise. Signal Y represents a noise signal that contains long-term noise.

[0108] Figure 9(c) is a graph showing an example of the magnitude of ambient sound extracted by the ambient sound extraction unit 2068. The vertical axis shows the level of the audio signal generated from the acquired ambient sound. Thresholds Th1 and Th2 are two thresholds used in the ambient sound determination unit 2069.

[0109] Figure 9(d) is a graph showing an example of the value of Nch_Before[n]. Nch_Before[n] is the signal at the nth frequency point in the Nch_Before output from the FFT unit 203 where the signal indicating the driving sound of the optical lens 300 is characteristically present. The vertical axis represents the value of Nch_Before[n]. The noise signals shown by signals X and Y in Figure 9(b) are more characteristically present in Nch_Before[n] than in Lch_Before.

[0110] Figure 9(e) is a graph showing an example of the value of Ndiff[n]. Ndiff[n] represents the signal value at the nth frequency point of the Ndiff output from the Nch differential unit 2064. The vertical axis represents the value of Ndiff[n]. When the change in the value of Nch_Before[n] per predetermined time is large, the value of Ndiff[n] becomes large. The short-term noise detection unit 2065 has a first threshold, threshold Th_Ndiff[n], to detect short-term noise. The threshold Th_Ndiff[n] changes between levels 1 and 3 based on data indicating the magnitude of ambient noise input from the ambient sound determination unit 2069 and the lens control signal. The initial value of threshold Th_Ndiff[n] is set to level 2. The level of threshold Th_Ndiff[n] is represented by the horizontal dashed line.

[0111] Figure 9(f) is a graph showing an example of the value of Nint[n]. In this embodiment, Nint[n] represents the value of the signal at the nth frequency point among the Nint output from the Nch integration unit 2066. The vertical axis represents the value of Nint[n]. If Nch_Before[n] is continuously large, the value of Nint[n] will increase. The long-term noise detection unit 2067 has a second threshold, Th_Nint[n], to detect long-term noise. The threshold Th_Nint[n] changes between levels 1 and 3 based on data indicating the magnitude of ambient noise input from the ambient sound determination unit 2069 and the lens control signal. The initial value of the threshold Th_Nint[n] is set to level 2. The level of the threshold Th_Nint[n] is represented by the horizontal dashed line.

[0112] Figure 9(g) shows an example of noise parameters selected by the noise parameter selection unit 206. In this embodiment, the solid color area indicates that only the PL1 noise parameter has been selected. The shaded area indicates that both the PL1 and PL2 noise parameters have been selected. The checkered area indicates that both the PL1 and PL3 noise parameters have been selected.

[0113] Figure 9(h) is a graph showing an example of the value of NL[n]. In this embodiment, NL[n] represents the signal value at the nth frequency point among the NL generated by the noise data generation unit 204. The vertical axis represents the value of NL[n].

[0114] Figure 9(i) is a graph showing an example of the value of Lch_After[n]. In this embodiment, Lch_After[n] represents the value of the signal at the nth frequency point among the Lch_After output from the subtraction processing unit 207. The vertical axis represents the value of Lch_After[n].

[0115] Next, the timing of each operation will be explained using the time intervals t701 to t709.

[0116] At time t701, the lens control unit 102 outputs a High signal as a lens control signal to the optical lens 300 and the noise parameter selection unit 206 (Figure 9(a)). At time t701, there is a high possibility that driving noise will occur inside the housing of the optical lens 300, so the short-term noise detection unit 2065 lowers the threshold Th_Ndiff[n] to level 1 (Figure 9(e)). Also at time t701, there is a high possibility that driving noise will occur inside the housing of the optical lens 300, so the long-term noise detection unit 2067 lowers the threshold Th_Nint[n] to level 1 (Figure 9(f)).

[0117] At time t702, the optical lens 300 is driven, generating short-term driving noises such as the sound of gears meshing. The noise microphone 201c picks up these short-term driving noises, causing the value of Ndiff[n] to exceed the threshold Th_Ndiff[n] (Figure 9(e)). Accordingly, the noise parameter selection unit 206 selects noise parameters PL1 and PL2 (Figure 9(g)). The noise data generation unit 204 generates NL[n] based on Nch_Before[n] and noise parameters PL1 and PL2 (Figure 9(h)). The subtraction processing unit 207 subtracts NL[n] from Lch_Before[n] and outputs Lch_After[n] (Figure 9(i)). In this case, Lch_After[n] becomes an audio signal with constant noise and short-term noise reduced.

[0118] At time t703, the optical lens 300 begins continuous operation, generating long-term driving noises such as sliding noises within the housing of the optical lens 300. The noise microphone 201c picks up these long-term driving noises, causing the value of Nint[n] to exceed the threshold Th_Nint[n] (Figure 9(f)). Accordingly, the noise parameter selection unit 206 selects noise parameters PL1 and PL3 (Figure 9(g)). The noise data generation unit 204 generates NL[n] based on Nch_Before[n] and noise parameters PL1 and PL3 (Figure 9(h)). The subtraction processing unit 207 subtracts NL[n] from Lch_Before[n] and outputs Lch_After[n] (Figure 9(i)). In this case, Lch_After[n] becomes an audio signal with constant noise and long-term noise reduced.

[0119] At time t704, the optical lens 300 stops continuous operation. Since the noise microphone 201c no longer picks up its long-term operation noise, the value of Nint[n] becomes less than or equal to the threshold Th_Nint[n] (Figure 9(f)). Accordingly, the noise parameter selection unit 206 selects the noise parameter PL1 (Figure 9(g)). The noise data generation unit 204 generates NL[n] based on Nch_Before[n] and the noise parameter PL1 (Figure 9(h)). The subtraction processing unit 207 subtracts NL[n] from Lch_Before[n] and outputs Lch_After[n] (Figure 9(i)). In this case, Lch_After[n] becomes an audio signal with the constant noise reduced.

[0120] At time t705, the lens control unit 102 outputs a Low signal as a lens control signal to the optical lens 300 and the noise parameter selection unit 206 (Figure 9(a)). In this case, the possibility of drive noise being generated inside the housing of the optical lens 300 is reduced, so the short-term noise detection unit 2065 raises the threshold Th_Ndiff[n] to level 2 (Figure 9(e)). Also in this case, the possibility of drive noise being generated inside the housing of the optical lens 300 is reduced, so the long-term noise detection unit 2067 raises the threshold Th_Nint[n] to level 2 (Figure 9(f)).

[0121] At time t706, the magnitude of the ambient noise extracted by the ambient noise extraction unit 2068 exceeds the threshold Th1. When the ambient noise is loud, the user is less likely to perceive the noise contained in the audio signal, so the short-term noise detection unit 2065 raises the threshold Th_Ndiff[n] to level 3 (Figure 9(e)). Also, when the ambient noise is loud, the user is less likely to perceive the noise contained in the audio signal, so the long-term noise detection unit 2067 raises the threshold Th_Nint[n] to level 3 (Figure 9(f)).

[0122] At time t707, the lens control unit 102 outputs a High signal as a lens control signal to the optical lens 300 and the noise parameter selection unit 206 (Figure 9(a)). In this case, there is a high possibility that driving noise will be generated inside the housing of the optical lens 300, so the short-term noise detection unit 2065 lowers the threshold Th_Ndiff[n] to level 2 (Figure 9(e)). Also in this case, there is a high possibility that driving noise will be generated inside the housing of the optical lens 300, so the long-term noise detection unit 2067 lowers the threshold Th_Nint[n] to level 2 (Figure 9(f)).

[0123] At time t708, the magnitude of the ambient sound extracted by the ambient sound extraction unit 2068 exceeds the threshold Th2. Here, the noise parameter selection unit 206 selects only the noise parameter PL1 regardless of the data input from the Nch noise detection unit 2061. In this way, when the ambient sound is very loud, the noise contained in the audio signal is difficult for the user to perceive. Therefore, the imaging device 100 reduces only the constant noise, resulting in a more natural recording of ambient sound than if processing to further reduce short-term and long-term noise were performed.

[0124] As described above, the imaging device 100 can record ambient sound with reduced noise by performing noise reduction processing using the second microphone, the noise microphone 201c.

[0125] Furthermore, the imaging device 100 detects the presence of noise using the output signal from the noise microphone 201c and sets the noise parameters in accordance with the timing of noise detection. Therefore, the imaging device 100 can appropriately set the noise parameters in synchronization with the occurrence of noise and appropriately reduce the noise.

[0126] Furthermore, when the ambient noise level is below the threshold Th2, the imaging device 100 performs noise reduction processing according to the noise detected by the Nch noise detection unit 2061, and when the ambient noise level is above the threshold Th2, it reduces only the constant noise. As a result, the imaging device 100 can record ambient noise with reduced noise to minimize discomfort for the user, depending on the ambient noise level.

[0127] In this embodiment, the imaging device 100 reduces the drive noise generated within the housing of the optical lens 300, but the drive noise generated within the imaging device 100 may also be reduced. Examples of drive noise generated within the imaging device 100 include substrate noise and radio wave noise. Substrate noise is, for example, the sound generated by the creaking of the substrate when a voltage is applied to a capacitor on the substrate.

[0128] The thresholds Th1 and Th2 of the ambient sound determination unit 2069, the threshold Th_Ndiff[n] of the short-term noise detection unit 2065, and the threshold Th_Nint[n] of the long-term noise detection unit 2067 are determined based on the generated driving noise and ambient noise. Therefore, the imaging device 100 may change these thresholds depending on the type of optical lens 300 and the tilt of the imaging device 100.

[0129] [Second example] Here, Figure 10 is a block diagram showing an example configuration of the audio input unit 104 in the second embodiment. The parts that differ from the configuration of the audio input unit 104 shown in Figure 3 are the subtraction processing unit 207 and the iFFT unit 208. Here, a description of the processing unit, which is the same as in Figure 3, is omitted.

[0130] The iFFT unit 208a performs inverse fast Fourier transforms on Lch_Before and Rch_Before, input from the FFT unit 203, respectively, to convert the frequency-domain digital audio signals into time-domain digital audio signals. The iFFT unit 208b also performs inverse fast Fourier transforms on NL and NR, respectively, to convert the frequency-domain digital audio signals into time-domain digital audio signals.

[0131] The subtraction processing unit 207 subtracts the digital audio signal input from iFFT unit 208b from the digital audio signal input from iFFT unit 208a. The calculation in the subtraction processing unit 207 is a waveform subtraction method that subtracts the digital audio signal in the time domain.

[0132] Furthermore, when performing waveform subtraction, the imaging device 100 may also record parameters related to the phase of the digital audio signal as noise parameters.

[0133] The configuration and operation of the other imaging device 100 are the same as in the first embodiment.

[0134] [Third example] In the third embodiment, a configuration in which the imaging device 100 has two subtraction processing units will be described.

[0135] Figure 11 is a block diagram showing an example of the configuration of the audio input unit 104 in the third embodiment.

[0136] Here, the microphone 201, A / D converter 202, FFT unit 203, iFFT unit 208, and audio processing unit 209 shown in Figure 11 are the same as those shown in Figure 3, so their explanation is omitted.

[0137] The switching unit 210 switches paths based on control information from the lens control unit 102. In this embodiment, when the optical lens 300 is driven, the switching unit 210 switches paths so that noise reduction processing is performed by the arithmetic processing unit A217, which will be described later. When the optical lens 300 is not driven, the switching unit 210 switches paths so that noise reduction processing is not performed by the arithmetic processing unit A217.

[0138] The noise data generation unit A214 generates data to reduce noise related to lens drive included in Lch_Before and Rch_Before based on Nch_Before. Noise related to lens drive included in the audio signals input from the L microphone and R microphone is an example of the first type of noise. In this embodiment, the noise data generation unit A214 generates array data NLA[0] to NLA

[0512] using noise parameters to reduce the noise included in Lch_Before[0] to Lch_Before

[0512] , respectively. The noise data generation unit A214 also generates array data NRA[0] to NRA

[0512] to reduce the noise included in Rch_Before[0] to Rch_Before

[0512] , respectively.

[0139] Note that the frequency points in the sequence data NLA[0]~NLA

[0512] are the same as the frequency points in the sequence data Lch_Before[0]~Lch_Before

[0512] . Also, the frequency points in the sequence data NRA[0]~NRA

[0512] are the same as the frequency points in the sequence data Rch_Before[0]~Rch_Before

[0512] .

[0140] Note that when referring collectively to the sequence data from NLA[0] to NLA

[0512] , it will be written as NLA. Similarly, when referring collectively to NRA[0] to NRA

[0512] , it will be written as NRA. NLA and NRA are examples of third frequency spectrum data, respectively.

[0141] The noise parameter recording unit 205 stores noise parameters for the noise data generation unit A214 to generate NLA and NRA from Nch_Before. In this embodiment, the noise parameter recording unit 205 stores noise parameters related to lens drive for each lens type, which are used by the noise data generation unit A214. In this embodiment, the noise data generation unit A214 does not switch noise parameters while recording audio data.

[0142] Here, when referring to the noise parameters used to generate NLA from Nch_Before, we will use the descriptive term PLxA. When referring to the noise parameters used to generate NRA from Nch_Before, we will use the descriptive term PRxA.

[0143] PLxA and PRxA have the same number of elements as NLA and NRA, respectively. For example, PL1A is the element data from PL1A[0] to PL1A

[0512] . Also, the frequency points of PL1A are the same as the frequency points of Lch_Before. Similarly, PR1A is the element data from PR1A[0] to PR1A

[0512] . The frequency points of PR1A are the same as the frequency points of Rch_Before. The noise parameters will be described later using Figure 12.

[0144] In this embodiment, the noise parameter recording unit 205 records all coefficients for each of the 513 frequency spectrum points as noise parameters. However, the noise parameter recording unit 205 only needs to record the coefficients for at least the frequency points necessary to reduce noise, rather than the coefficients for all 513 frequencies. For example, the noise parameter recording unit 205 may record the coefficients for each of the frequency spectrum points from 20 Hz to 20 kHz, which are considered typical audible frequencies, as noise parameters, and may not need to record the coefficients for other frequency spectrum points. Also, for example, coefficients for frequency spectrum points with a coefficient value of zero do not need to be recorded in the noise parameter recording unit 205 as noise parameters.

[0145] The subtraction processing unit A217 subtracts NLA and NRA from Lch_Before and Rch_Before, respectively. For example, the subtraction processing unit A217 has an L subtractor A217a that subtracts NLA from Lch_Before, and an R subtractor A217b that subtracts NRA from Rch_Before. The L subtractor A217a subtracts NLA from Lch_Before and outputs 513-point sequence data from Lch_A_After[0] to Lch_A_After

[0512] . The R subtractor A217b subtracts NRA from Rch_Before and outputs 513-point sequence data from Rch_A_After[0] to Rch_A_After

[0512] . In this embodiment, the subtraction processing unit A217 performs subtraction processing using the spectral subtraction method, and in particular subtracts noise related to lens driving, which has a large amount of noise.

[0146] The noise data generation unit B224 generates data to reduce a second noise related to lens drive included in Lch_A_After and Rch_A_After, based on Nch_Before.

[0147] In this embodiment, the noise data generation unit B224 generates array data of NLB[0] to NLB

[0512] to reduce the noise contained in Lch_A_After[0] to Lch_A_After

[0512] using noise parameters. The noise data generation unit B224 also generates array data of NRB[0] to NRB

[0512] to reduce the noise contained in Rch_A_After[0] to Rch_A_After

[0512] using noise parameters.

[0148] The frequency points in the sequence data NLB[0]~NLB

[0512] are the same as the frequency points in the sequence data Lch_A_After[0]~Lch_A_After

[0512] . Also, the frequency points in the sequence data NRB[0]~NRB

[0512] are the same as the frequency points in the sequence data Rch_A_After[0]~Rch_A_After

[0512] .

[0149] Note that when referring collectively to the sequence data from NLB[0] to NLB

[0512] , it will be written as NLB. Similarly, when referring collectively to NRB[0] to NRB

[0512] , it will be written as NRB. NLB and NRB are examples of (the fourth frequency spectrum data, respectively).

[0150] The noise parameter recording unit 205 records the noise parameters that the noise data generation unit B224 uses to generate NLB and NRB from Nch_Before.

[0151] In this embodiment, the noise parameter recording unit 205 records noise parameters used in the noise data generation unit B224, such as microphone floor noise and electrical noise. In this embodiment, the noise data generation unit A214 does not switch noise parameters while recording audio data.

[0152] Here, when referring to the noise parameters for generating NLB from Nch_Before, we will use the descriptive term PLxB. When referring to the noise parameters for generating NRB from Nch_Before, we will use the descriptive term PRxB.

[0153] PLxB and PRxB have the same number of array elements as NLB and NRB, respectively. For example, PL1B is the array data from PL1B[0] to PL1B

[0512] . Also, the frequency points of PL1B are the same as the frequency points of Lch_Before. Similarly, PR1B is the array data from PR1B[0] to PR1B

[0512] . The frequency points of PR1B are the same as the frequency points of Rch_Before. The noise parameters will be described later using Figure 12.

[0154] In this embodiment, the noise parameter recording unit 205 records all coefficients for each of the 513 frequency spectrum points as noise parameters. However, the noise parameter recording unit 205 only needs to record the coefficients for at least the frequency points necessary to reduce noise, rather than recording the coefficients for all 513 frequencies. For example, the noise parameter recording unit 205 may record the coefficients for each of the frequency spectrum points from 20 Hz to 20 kHz, which are considered typical audible frequencies, as noise parameters, and may not need to record the coefficients for other frequency spectrum points. Also, for example, coefficients for frequency spectrum points with a coefficient value of zero do not need to be recorded in the noise parameter recording unit 205 as noise parameters.

[0155] The subtraction processing unit B227 subtracts NLB and NRB from Lch_A_After and Rch_A_After, respectively. For example, the subtraction processing unit B227 has an L subtractor B227a that subtracts NLB from Lch_A_AFTER, and an R subtractor B227b that subtracts NRB from Rch_Before. The L subtractor B227a subtracts NLB from Lch_Before and outputs 513-point sequence data from Lch_After[0] to Lch_After

[0512] . The R subtractor B227b subtracts NRA from Rch_Before and outputs 513-point sequence data from Rch_After[0] to Rch_After

[0512] . In this embodiment, the subtraction processing unit B227 performs subtraction processing using the spectral subtraction method, and in particular subtracts noise related to lens driving, which has a large amount of noise.

[0156] In this embodiment, the subtraction processing unit B227 subtracts constantly occurring noises other than noise generated by lens drive, such as microphone floor noise and electrical noise. In this embodiment, the noise data generation unit B224 generates NLB and NRB based on Nch_Before, but other methods may be used. For example, NLB and NRB may be recorded in the noise parameter recording unit 205, and the subtraction processing unit B227 may read NLB and NRB directly from the noise parameter recording unit 205 without going through the noise data generation unit B224. This is because there is little need to refer to noises included in Nch_Before, such as microphone floor noise and electrical noise, which are constantly occurring.

[0157] In this embodiment, the noise reduction process is described as being performed in the order of subtraction processing unit A217 to subtraction processing unit B227, but the noise reduction process may also be performed in the reverse order, from subtraction processing unit B227 to subtraction processing unit A217.

[0158] The configuration and operation of the other imaging device 100 are the same as in the first embodiment.

[0159] <Noise parameters in the third embodiment> Figure 12 shows an example of noise parameters recorded in the noise parameter recording unit 205 in the third embodiment. Noise parameters are parameters for correcting the audio signal generated by the noise microphone 201c acquiring the drive sound generated inside the housing of the imaging device 100 and the housing of the optical lens 300. As shown in Figure 12, in this embodiment, PLxA, PRxA, PLxB, and PRxB are recorded in the noise parameter recording unit 205. In this embodiment, the source of the drive sound, as PLxA and PRxA, is described as being inside the housing of the optical lens 300. The drive sound generated inside the housing of the optical lens 300 is transmitted to the housing of the imaging device 100 via the lens mount 301 and acquired by the L microphone 201a, the R microphone 201b, and the noise microphone 201c.

[0160] In this embodiment, multiple noise parameters corresponding to the type of optical lens 300 are recorded in the noise parameter recording unit 205. This is because the frequency of the drive sound differs depending on the type of optical lens 300. The imaging device 100 generates noise data using the noise parameter corresponding to the type of optical lens 300 from among these multiple noise parameters.

[0161] In this embodiment, the imaging device 100 records noise parameters for persistent noise as PLxB and PRxB for each video mode. Persistent noise includes, for example, white noise, microphone floor noise, and electrical noise. In addition, the imaging device 100 may record noise parameters for each type of optical lens 300, as well as for the temperature inside the housing of the imaging device 100 and the tilt of the imaging device 100, as detected by the information acquisition unit 103.

[0162] Furthermore, the average values ​​of the PLxA and PRxA coefficients are greater than the average values ​​of the PLxB and PRxB coefficients. This is because the noise reduced by PLxA and PRxA is louder and more jarring than the noise reduced by PLxB and PRxB.

[0163] <Timing chart of noise reduction processing in the third embodiment> The noise reduction process in this embodiment will be explained with reference to Figure 13.

[0164] Figures 13(a),(b),(d),(g),(hA),(hB), and(i) are examples of timing charts for audio processing in the noise data generation unit 204, the noise parameter selection unit 206, and the subtraction processing unit 207. In this embodiment, for the sake of simplicity, the audio processing of the Lch is described, but the audio processing of the Rch is similar. In Figures 9(a),(b),(d),(g),(hA),(hB), and(i), the horizontal axis of the graphs is the time axis.

[0165] Figure 13(a) shows an example of a lens control signal. The lens control signal is a signal that instructs the lens control unit 102 to drive the optical lens 300. In this embodiment, the level of the lens control signal is represented by two values: High and Low. When the level of the lens control signal is High, the lens control unit 102 is instructing the optical lens 300 to drive. That is, when the level of the lens control signal from the control unit 111 is High, it can be determined that noise is being generated from the optical lens 300. When the level of the lens control signal is Low, the lens control unit 102 is not instructing the optical lens 300 to drive.

[0166] Figure 13(b) is a graph showing an example of the value of Lch_Before[n]. The vertical axis represents the value of Lch_Before[n]. In this embodiment, Lch_Before[n] is the signal at the nth frequency point among the Lch_Before signals output from the FFT unit 203, in which the signal indicating the driving sound of the optical lens 300 is characteristically present. In this embodiment, the signal at the nth frequency point is described, but audio processing is performed similarly for other frequencies. In addition, the signals V and W are signals that contain noise. In this embodiment, signal V represents a signal that contains noise associated with lens driving. Signal W represents a noise signal that contains constant noise such as microphone floor noise and electrical noise.

[0167] Figure 13(d) is a graph showing an example of the value of Nch_Before[n]. Nch_Before[n] is the signal at the nth frequency point in the Nch_Before output from the FFT unit 203 where the signal indicating the driving sound of the optical lens 300 is characteristically present. The vertical axis represents the value of Nch_Before[n]. The noise signals shown by signals V and W in Figure 9(b) are more characteristically present in Nch_Before[n] than in Lch_Before.

[0168] Figure 13(g) shows an example of the operation state of subtraction processing units A217 and B227 selected by the switching unit 210. In this embodiment, the plain areas indicate that noise reduction processing is performed only by subtraction processing unit B227. The checkered areas indicate that noise reduction processing is performed by both subtraction processing units A217 and B227.

[0169] Figure 13(hA) is a graph showing an example of the NLA[n] value. In this embodiment, NLA[n] represents the signal value at the nth frequency point of the NLA generated by the noise data generation unit A214. The vertical axis represents the value of NLA[n].

[0170] Figure 13(hB) is a graph showing an example of the NLB[n] value. In this embodiment, NLB[n] represents the signal value at the nth frequency point of the NLB generated by the noise data generation unit B224. The vertical axis represents the value of NLB[n].

[0171] Figure 13(i) is a graph showing an example of the value of Lch_After[n]. In this embodiment, Lch_After[n] represents the signal value at the nth frequency point among the Lch_After output from the subtraction processing unit 207. The vertical axis represents the value of Lch_After[n].

[0172] Next, the timing of each operation will be explained using the time intervals t1301 to t1302.

[0173] At time t1301, the lens control unit 102 outputs a High signal as a lens control signal to the optical lens 300 and the noise parameter selection unit 206 (Figure 13(a)).

[0174] At time t1301, the switching unit 210 switches from noise reduction processing performed only by subtraction processing unit B227 to noise reduction processing performed by both subtraction processing unit A217 and subtraction processing unit B227 (Figure 13(g)). From time t1301, the noise data generation unit A214 generates NLA[n] based on Nch_Before[n] and the noise parameter PLxA[n] (Figure 13(hA)). Alternatively, the noise data generation unit A214 may always generate NLA[n], and the subtraction processing unit A217 may start subtraction when the lens control signal becomes high. The noise data generation unit B224 generates NLB[n] based on Nch_Before[n] and the noise parameter PLxB[n] (Figure 13(hB)).

[0175] From time t1301, subtraction processing units A217 and B227 subtract NLA[n] and NLB[n] from Lch_Before[n] and output Lch_After[n] (Figure 13(i)).

[0176] At time t1302, the lens control unit 102 determines that the driving of the optical lens 300 has finished and outputs a Low signal as a lens control signal to the optical lens 300 and the noise parameter selection unit 206 (Figure 13(a)).

[0177] At time t1302, the switching unit 210 switches from noise reduction processing by subtraction processing units A217 and B227 to noise reduction processing performed only by subtraction processing unit B227 (Figure 13(g)). Since subtraction processing unit A217 is no longer used from time t1302 onwards, NLA[n] is not used (shaded area in Figure 13(hA)). On the other hand, NLB[n] generated by noise data generation unit B224 continues to be used by subtraction processing unit B227 (Figure 13(hB)).

[0178] At time t1302, the subtraction processing unit B227 subtracts NLB[n] from Lch_Before[n] and outputs Lch_After[n] (Figure 13(i)).

[0179] Subsequently, the audio input unit 104 performs noise reduction processing as described above based on the signal output from the lens control unit 102.

[0180] In this way, the imaging device 100 can reduce power consumption by switching the noise reduction processing, which is performed only while the optical lens 300 is in motion, based on the lens control signal.

[0181] [Other examples] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or recording medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0182] It should be noted that the present invention is not limited to the above embodiments, and the components can be modified and implemented in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the above embodiments. For example, some components may be deleted from all the components shown in the embodiments. Moreover, components from different embodiments may be appropriately combined.

Claims

1. a first microphone for capturing environmental sounds; a second microphone for capturing sound from the noise source; a first conversion means for Fourier-transforming the audio signal from the first microphone to generate a first audio signal; a second conversion means for Fourier-transforming the audio signal from the second microphone to generate a second audio signal; a generating means for generating noise data using the second audio signal and parameters related to the noise of the noise source; a subtraction means for subtracting the noise data from the first audio signal; a third transform means for performing an inverse Fourier transform on the audio signal from the subtraction means; 10. A voice processing device comprising:

2. 2. The audio processing device according to claim 1, wherein the generating means generates the noise data using at least one of a plurality of parameters including a first parameter corresponding to a first type of noise and a second parameter corresponding to a second type of noise and the second audio signal.

3. 3. The voice processing apparatus according to claim 2, further comprising a recording means for recording information on the plurality of parameters.

4. the first microphone is composed of a plurality of microphones, The recording means records the parameters for each microphone constituting the first microphone.

4. The audio processing device according to claim 3.

5. The generating means generates the noise data using a parameter corresponding to a type of noise included in the second audio signal and the second audio signal, among the plurality of parameters.

5. The audio processing device according to claim 2, wherein the audio processing device is a voice processing device.

6. a driving means as the noise source, 6. The audio processing device according to claim 2, wherein the generating means selects noise parameters according to the driving state of the driving means, and generates the noise data using the selected parameters and the second audio signal.

7. 7. The audio processing device according to claim 2, wherein the generating means selects one of the plurality of parameters based on the second audio signal.

8. The noise includes at least one of constant noise, short-term noise, and long-term noise.

8. The audio processing device according to claim 1, wherein the audio processing device is a voice processing device.

9. In the sound processing device, a hole for inputting environmental sound is formed above the first microphone, and a hole for inputting environmental sound is not formed above the second microphone.

9. The audio processing device according to claim 1, wherein the audio processing device is a voice processing device.

10. 10. The audio processing device according to claim 1, wherein the parameter is a ratio between the amplitudes of the first audio signal and the second audio signal.

11. further comprising an imaging means; The noise source is a member that is driven during imaging by the imaging means.

11. The audio processing device according to claim 1.

12. 12. The audio processing device according to claim 1, further comprising a switching unit that switches the noise data generated by the generating unit so as to be different depending on whether or not noise is generated from the noise source.

13. 13. The audio processing device according to claim 12, wherein the subtraction means comprises first subtraction means for reducing constant noise and second subtraction means for reducing noise other than constant noise.

14. 14. The audio processing device according to claim 13, wherein the generating means comprises: first generating means for generating a noise parameter to be used in the first subtracting means; and second generating means for generating a noise parameter to be used in the second subtracting means.

15. 14. The audio processing device according to claim 13, wherein the generating means generates noise parameters to be used in the first subtracting means, and the noise parameters to be used in the second subtracting means are noise parameters recorded in advance in a recording means.

16. a first microphone for capturing environmental sounds; a second microphone for acquiring sound from a noise source, generating a first audio signal by Fourier transforming the audio signal from the first microphone; performing a Fourier transform on the audio signal from the second microphone to generate a second audio signal; generating noise data using the second audio signal and parameters related to noise from the noise source; a subtraction step of subtracting the noise data from the first audio signal; performing an inverse Fourier transform on the audio signal produced by the subtraction step; A control method comprising:

17. A computer-readable program for causing a computer to function as each of the means of the voice processing device according to any one of claims 1 to 15.

18. a first microphone for capturing environmental sounds; a second microphone for capturing sound from the noise source; a first conversion means for Fourier-transforming the audio signal from the first microphone to generate a first audio signal; a second conversion means for Fourier-transforming the audio signal from the second microphone to generate a second audio signal; a generating means for generating noise data using the second audio signal and parameters related to the noise of the noise source; a third transforming means for performing an inverse Fourier transform on the first audio signal, the second audio signal, and the noise data; a subtraction means for subtracting the inverse Fourier transformed noise data from each of the inverse Fourier transformed first audio signal and the inverse Fourier transformed second audio signal; 10. A voice processing device comprising:

19. 19. The audio processing device according to claim 18, wherein the generating means generates the noise data using at least one of a plurality of parameters including a first parameter corresponding to a first type of noise and a second parameter corresponding to a second type of noise and the second audio signal.

20. a driving means as the noise source, 20. The audio processing device according to claim 19, wherein the generating means selects a noise parameter according to a driving state of the driving means, and generates the noise data using the selected parameter and the second audio signal.

21. further comprising an imaging means; The noise source is a member that is driven during imaging by the imaging means.

21. The audio processing device according to claim 18, wherein:

22. a first microphone for capturing environmental sounds; a second microphone for acquiring sound from a noise source, generating a first audio signal by Fourier transforming the audio signal from the first microphone; performing a Fourier transform on the audio signal from the second microphone to generate a second audio signal; generating noise data using the second audio signal and parameters related to noise from the noise source; performing an inverse Fourier transform on the first audio signal, the second audio signal, and the noise data; subtracting the inverse Fourier transformed noise data from each of the inverse Fourier transformed first audio signal and the inverse Fourier transformed second audio signal; A control method comprising:

23. A computer-readable program for causing a computer to function as each of the means of the voice processing device according to any one of claims 18 to 21.