Imaging apparatus, control method, and program
The imaging device addresses the challenge of adjusting audio levels for each channel by detecting the subject's position in the image and adjusting the recording levels accordingly, enhancing the playback experience by synchronizing audio and video.
Patent Information
- Application Number
- JP2023203121
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-11
AI Technical Summary
Existing methods for recording monaural audio with cameras using wireless microphones struggle to accurately adjust audio levels for each channel based on the position of the sound source in the image, leading to discomfort during playback.
An imaging device with an imaging unit, a sound collection unit, a synthesis unit, a detection unit, a setting unit, and determination units that detect a subject in the image, associate it with a sound collection device, and adjust the recording level for each channel of audio data based on the subject's position.
The solution allows for precise adjustment of audio levels for each channel in association with a subject in the image, reducing user discomfort caused by mismatched audio and video during playback.
Smart Images

Figure 2025088427000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to recording / recording processing by an imaging device.
Background Art
[0002] Conventionally, when recording monaural audio with a camera using a wireless microphone or the like, audio data mixed in each channel (for example, the L channel and the R channel) is recorded, but the audio data is evenly distributed and recorded with respect to the sound source position in the image. Therefore, when playing back a moving image including audio, the sound source position in the image does not match the audio level (volume) for each channel, which may give the user a sense of discomfort.
[0003] In response to such a background, Patent Document 1 describes a method of selecting audio data corresponding to a subject identified as a sound source from monaural audio data and adjusting the volume for each channel according to the position of the identified subject in the image. Patent Document 2 describes a method of separating image data with audio data into audio data and image data, separating the sound source of a specific subject in the image from the separated audio data, and rearranging the separated sound source in a sound field space suitable for the image.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in Patent Documents 1 and 2, it is necessary to extract only the audio data corresponding to the subject from the audio data with high accuracy, adjust the volume for each channel, and recombine it with the data other than the audio data of the subject. Therefore, if the audio data of the subject cannot be appropriately extracted, unnatural audio data may be generated.
[0006] The present invention has been made in view of the above problems, and an object thereof is to realize a technique capable of adjusting the level of audio data recorded for each channel in association with a subject in an image.
Means for Solving the Problems
[0007] In order to solve the above problems and achieve the object, an imaging device of the present invention includes an imaging unit, a connection unit connected to a sound collection device, a synthesis unit that synthesizes audio data acquired from the sound collection device and moving image data acquired by the imaging unit, a detection unit that detects a subject included in a screen of a moving image captured by the imaging unit, a setting unit that associates the subject with a sound collection device connected to the imaging device, a determination unit that determines a position of the subject detected by the detection unit within a screen of a moving image captured by the imaging unit, and a determination unit that determines a position of the subject detected by the detection unit within a screen of a moving image captured by the imaging unit. A determination means for determining the recording level for each channel of the audio data acquired from the sound collection device associated with the subject based on the position of the subject determined by the determination means.
Effects of the Invention
[0008] According to the present invention, the level of audio data recorded for each channel can be adjusted in association with a subject in an image.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant explanations are omitted.
[0011] [Embodiment 1] FIG. 1 is a diagram illustrating a system including the imaging device and the sound collection device according to Embodiment 1.
[0012] The system of the present embodiment includes an imaging device 100 and a sound collection device 400. In the present embodiment, as the imaging device 100 and the sound collection device 400, an example of a camera and a wireless microphone (hereinafter, wireless mic) that can be connected by a wireless communication method will be described.
[0013] The wireless microphone 400 is wirelessly communicably connected to the camera 100 via a receiver 300 connected to the camera 100. A plurality of wireless microphones 400 can be connected to the receiver 300.
[0014] The wireless microphone 400 picks up the sound around the wireless microphone 400 to generate an analog audio signal, performs various signal processes on the analog audio signal, converts it into a digital signal, and generates audio data. The wireless microphone 400 transmits the audio signal to the camera 100 via the receiver 300 by wireless communication. The camera 100 receives the audio signal transmitted from the wireless microphone 400 via the receiver 300.
[0015] In the present embodiment, an example will be described in which the wireless microphone 400 is attached to two subjects that are the shooting targets of the camera 100, and the wireless microphone 400 picks up the sound of the subject to which it is attached. However, the number of wireless microphones and subjects may be one or three or more. Also, an example will be described in which the channels capable of recording audio data in the camera 100 are two channels, an L channel corresponding to the left direction of the camera 100 and an R channel corresponding to the right direction of the camera 100. However, the number of channels is not limited to two types and may be three or more types. Further, in the present embodiment, an example of an interchangeable-lens camera 100 in which the lens unit 200 is detachable will be described. However, the present invention is not limited to an interchangeable-lens camera, and a camera with a built-in lens may be used. Further, the camera 100 may be a smartphone with a camera function or a head-mounted display with a built-in camera that can be connected to the wireless microphone 400 by a wireless communication method.
[0016] In FIG. 1, a lens unit 200 and a receiver 300 can be attached to the camera 100. The lens unit 200 is attached to the front surface portion of the camera 100. The camera 100 and the lens unit 200 are mechanically and electrically connected via a lens connection portion 103. The receiver 300 is attached to the upper surface portion of the camera 100. The camera 100 and the receiver 300 are mechanically and electrically connected via an accessory connection portion 109.
[0017] The camera control unit 101 is a microcomputer that controls each component of the camera 100. The camera control unit 101 includes a processor such as a CPU that performs arithmetic processing and control processing related to the camera 100, a non-volatile memory that stores programs executed by the processor and data referred to by the programs, and a volatile memory into which programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM. Note that the program of this embodiment includes a program for executing the flowcharts described later in FIGS. 2, 7, and 9.
[0018] The imaging unit 102 includes an image sensor composed of a CCD, a CMOS, or the like that converts the optical image of a subject incident through the lens 202 into an electrical signal, and an A / D converter that converts the analog image signal output from the image sensor into a digital signal. The imaging unit 102 converts the subject image formed by the lens 202 into an electrical signal by the image sensor based on the control information of the camera control unit 101, performs noise reduction processing, etc., and outputs image data composed of a digital signal.
[0019] The focal plane shutter (hereinafter referred to as the shutter) 104 is interposed between the imaging unit 102 and the lens 202 based on the control information of the camera control unit 101, and can freely control the exposure time in the imaging unit 102. The shutter 104 includes a front curtain and a rear curtain. In the shutter 104, the exposure of the imaging unit 102 is started when the front curtain travels and the shutter 104 opens, and the exposure of the imaging unit 102 ends when the rear curtain travels and the shutter 104 closes.
[0020] The camera operation unit 105 includes operation members such as push buttons, slide switches, rotary dials, and a touch panel integrally formed with the camera display unit 106 described later that receive user operations. The camera operation unit 105 sends state information corresponding to the operation state of the camera operation unit 105 to the camera control unit 101. The operation members included in the camera operation unit 105 include a power button for turning on and off the power of the camera 100 and a mode dial for setting the operation mode of the camera 100. Further, the operation members included in the camera operation unit 105 include a video shooting button for instructing the start or end of a video shooting operation and a still image shooting button for instructing the shooting of a still image.
[0021] In the still image shooting mode, when the still image shooting button is half-pressed, the camera operation unit 105 sends a shooting preparation signal (SW1 signal) to the camera control unit 101. Further, when the still image shooting button is further pressed from the half-pressed state to the fully pressed state, the camera operation unit 105 sends a shooting signal (SW2 signal) to the camera control unit 101. Also, in the video recording mode, when the video shooting button is first pressed, the camera operation unit 105 sends a recording start signal (REC signal) to the camera control unit 101. Further, when the video shooting button is pressed again, the camera operation unit 105 sends a recording stop signal (REC signal) to the camera control unit 101.
[0022] The camera control unit 101 controls each component of the camera 100 based on the state information of the camera operation unit 105. When the camera control unit 101 receives the SW1 signal from the camera operation unit 105, it calculates focus information such as defocus information indicating the in-focus state of the subject image based on the image (captured image) captured by the imaging unit 102. Further, the camera control unit 101 detects the subject from the captured image and executes photometric control (AE operation) for measuring the luminance of the subject, and determines exposure control information such as the shutter speed, aperture value, and ISO sensitivity at the time of shooting from the photometric result. The determined exposure control information is displayed on the camera display unit 106.
[0023] The camera display unit 106 includes a display device such as a liquid crystal or an organic EL for displaying a live view image, a playback image, the setting information and operating state of the camera 100, and a GUI, and a light source such as an LED.
[0024] When the camera control unit 101 receives the SW2 signal from the camera operation unit 105, it drives the aperture 203 of the lens unit 200, sets the sensitivity (ISO sensitivity) of the imaging unit 102, and controls the shutter 104 to expose the imaging unit 102. When the camera control unit 101 receives the REC signal from the camera operation unit 105, it sets the sensitivity (ISO sensitivity) and frame rate of the imaging unit 102, calculates focus information based on the captured image of the imaging unit 102, and exposes the imaging unit 102 while repeating photometric control (AE operation). The lens control unit 201, which will be described later, drives a focus lens (not shown) for focus adjustment included in the lens 202 based on the control information of the camera control unit 101 to repeat autofocus (AF operation). The camera control unit 101 sequentially displays the image data output from the imaging unit 102 as a live view image on the camera display unit 106. Also, in the video shooting mode, the camera control unit 101 stores a video file including the captured video data or a video file including video data and audio data in the storage unit 107.
[0025] The storage unit 107 is an auxiliary storage device such as a magnetic disk, an optical disk, or a semiconductor memory that can store video files.
[0026] The camera communication unit 108 is an interface that enables wireless communication connection between the camera 100 and an external device. The camera communication unit 108 transmits and receives, between the camera 100 and the external device, for example, image data, audio data, composite data of image data and audio data, device information of the camera 100 and the external device, operation information, control information, and other information. The camera communication unit 108 is, for example, an infrared communication module, a Bluetooth (registered trademark) module, a wireless LAN (Local Area Network) module, a wireless communication module such as WirelessUSB. In the present embodiment, the configuration in which the camera 100 receives audio data transmitted from the wireless microphone 400 via the receiver 300 is illustrated, but it is not limited thereto. For example, by adding the function of the receiver 300 to the camera communication unit 108, the camera 100 may be configured to be able to communicate directly with the wireless microphone 400.
[0027] The camera audio input unit 110 is a sound collection unit such as a microphone built in the camera 100 or an external microphone connected via an audio input terminal. The camera audio input unit 110 collects the sound around the camera 100 to generate an analog audio signal, performs various signal processes on the analog audio signal, converts it into a digital signal, and generates audio data. The camera control unit 101 performs processes related to audio, such as normalization processing of the signal level of the audio data, reduction processing of a specific frequency, and audio detection processing, on the audio data generated by the camera audio input unit 110. Then, the camera control unit 101 performs a process of synthesizing the audio data subjected to audio processing with the image data acquired by the imaging unit 102 or the audio data acquired from the wireless microphone 400, and stores the encoded image file in the storage unit 107. Note that the microphone is a monaural microphone capable of collecting a one-channel audio signal or a stereo microphone capable of collecting a two-channel audio signal.
[0028] Next, the configuration of the lens unit 200 will be described.
[0029] The lens control unit 201 is a microcomputer that controls each component of the lens unit 200. The lens control unit 201 includes a processor and a built-in memory. The processor is a CPU that performs arithmetic processing and control processing related to the lens unit 200. The built-in memory includes a non-volatile memory that stores programs executed by the processor and data referred to by the programs, and a volatile memory into which programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM.
[0030] The lens 202 includes a plurality of lenses and forms a subject image on the imaging unit 102. The lens 202 includes a diaphragm 203 for adjusting the amount of light transmitted through the lens 202 and a focus lens (not shown) for focus adjustment. The lens control unit 201 adjusts the amount of light transmitted through the lens 202 and the focal length based on the control information received from the camera control unit 101 via the lens connection unit 103, and transmits the adjustment result to the camera control unit 101.
[0031] Next, the configuration of the receiver 300 will be described.
[0032] The receiver control unit 301 is a microcomputer that controls each component of the receiver 300. The receiver control unit 301 includes a processor and a built-in memory. The processor is a CPU that performs arithmetic processing and control processing related to the receiver 300. The built-in memory includes a non-volatile memory that stores programs executed by the processor and data referred to by the programs, and a volatile memory into which programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM.
[0033] The receiver control unit 301 is communicably connected to the camera control unit 101 via the accessory connection unit 109. The receiver control unit 301 communicates with the camera control unit 101 to transmit and receive, for example, audio data generated by the receiver audio input unit 305 and the microphone audio input unit 405, information on the receiver 300 and the wireless microphone 400, control information of the camera 100, and other information.
[0034] The receiver communication unit 302 is an interface that wirelessly communicably connects the receiver 300 and an external device (in this embodiment, the wireless microphone 400). The receiver communication unit 302 transmits and receives, for example, audio data generated by the wireless microphone 400, information on the receiver 300 and the wireless microphone 400, control information of the camera 100, and other information between the receiver 300 and the wireless microphone 400. The receiver communication unit 302 is, for example, a wireless communication module such as an infrared communication module, a Bluetooth (registered trademark) module, a wireless LAN module, or a WirelessUSB.
[0035] The receiver operation unit 303 includes operation members such as push buttons, slide switches, and rotary dials that receive user operations. The receiver operation unit 303 sends state information corresponding to the operation state of the receiver operation unit 303 to the receiver control unit 301.
[0036] The receiver display unit 304 includes a display device such as a liquid crystal or an organic EL that displays the setting information, operation state, and GUI of the receiver 300, and a light source such as an LED.
[0037] The receiver voice input unit 305 is a sound collection unit such as a microphone built into the receiver 300 or an external microphone connected via a voice input terminal. The receiver voice input unit 305 collects the voice around the receiver 300 to generate an analog voice signal, performs various signal processes on the analog voice signal, converts it into a digital signal, and generates voice data. The receiver control unit 301 performs voice-related processes such as signal level optimization processing, specific frequency reduction processing, and voice detection processing on the voice data generated by the receiver voice input unit 305. Note that the microphone is a monaural microphone capable of collecting a one-channel voice signal or a stereo microphone capable of collecting a two-channel voice signal.
[0038] Next, the configuration of the wireless microphone 400 will be described.
[0039] The microphone control unit 401 is a microcomputer that controls each component of the wireless microphone 400. The microphone control unit 401 includes a processor and a built-in memory. The processor is a CPU that performs arithmetic processing and control processing related to the wireless microphone 400. The built-in memory includes a non-volatile memory that stores programs executed by the processor and data referred to by the programs, and a volatile memory into which programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM.
[0040] The microphone control unit 401 is wirelessly communicably connected to the receiver control unit 301 via the microphone communication unit 402 and the receiver communication unit 302. The microphone control unit 401 transmits and receives voice data, device information and operation information of the receiver 300 and the wireless microphone 400, control information of the camera 100, and other information to and from the receiver control unit 301.
[0041] The microphone communication unit 402 is an interface that wirelessly connects the wireless microphone 400 and an external device (in this embodiment, the receiver 300). The microphone communication unit 402 transmits and receives, between the wireless microphone 400 and the receiver 300, for example, audio data generated by the wireless microphone 400, device information and operation information of the wireless microphone 400, control information of the camera 100, and other information. The microphone communication unit 402 is, for example, a wireless communication module such as an infrared communication module, a Bluetooth (registered trademark) module, a wireless LAN module, or a WirelessUSB.
[0042] The microphone operation unit 403 includes operation members such as push buttons, slide switches, and rotary dials that receive user operations. The microphone operation unit 403 sends state information corresponding to the operation state of the microphone operation unit 403 to the microphone control unit 401.
[0043] The microphone display unit 404 includes a display device such as a liquid crystal or an organic EL that displays the setting information, operation state, and GUI of the wireless microphone 400, and a light source such as an LED.
[0044] The microphone audio input unit 405 is a sound collection unit such as a microphone built in the wireless microphone 400 or an external microphone connected via an audio input terminal. The microphone audio input unit 405 collects the sound around the wireless microphone 400 to generate an analog audio signal, performs various signal processes on the analog audio signal, converts it into a digital signal, and generates audio data. The microphone control unit 401 performs processes related to audio, such as optimization processing of the signal level of the audio data, reduction processing of a specific frequency, and audio detection processing, on the audio data generated by the microphone audio input unit 405. Note that the microphone is a monaural microphone capable of collecting a one-channel audio signal or a stereo microphone capable of collecting a two-channel audio signal.
[0045] [Embodiment 1] Next, with reference to FIGS. 2 to 6, the recording / recording process of the camera 100 according to Embodiment 1 will be described.
[0046] FIG. 2 is a flowchart illustrating the recording / recording process of the camera 100 according to Embodiment 1.
[0047] The process of FIG. 2 is realized by the camera control unit 101 executing a program stored in the built-in memory. The same applies to FIGS. 7 and 9 described later.
[0048] When the power button included in the camera operation unit 105 is turned on, the camera 100 is supplied with power from a power supply unit (not shown) to each component of the camera 100, and each component of the camera 100 becomes operable. When the power is turned on, the camera control unit 101 controls each part of the camera 100, controls the imaging unit 102 to capture a live view image, and displays the live view image on the camera display unit 106. When the power button included in the receiver operation unit 303 of the receiver 300 is turned on, the receiver 300 is supplied with power from the camera 100 connected via the accessory connection unit 109, and each component of the receiver 300 becomes operable. Then, the receiver 300 is in a state where it can be wirelessly connected to the wireless microphone 400 via Bluetooth (registered trademark). When the power button included in the microphone operation unit 403 of the wireless microphone 400 is turned on, the wireless microphone 400 is supplied with power from a power supply unit (not shown) of the wireless microphone 400 to each component of the wireless microphone 400, and each component of the wireless microphone 400 becomes operable. Then, the wireless microphone 400 is in a state where it can be wirelessly connected to the receiver 300 via Bluetooth (registered trademark). The same applies to FIGS. 7 and 9 described later.
[0049] In this embodiment, an example in which the receiver 300 and the wireless microphone 400 are wirelessly connected via Bluetooth (registered trademark) will be described. However, the connection is not limited to Bluetooth (registered trademark), and they may be wirelessly connected via a wireless LAN. In this case, the receiver 300 functions as a master unit that forms a wireless LAN access point. The wireless microphone 400 functions as a slave unit that connects to the network formed by the receiver 300 through the wireless LAN access point.
[0050] In step S200, in response to the user performing an operation to connect the wireless microphone 400 and the receiver 300, the camera control unit 101 pairs the wireless microphone 400 and the receiver 300 and connects them so that wireless communication is possible. If the wireless microphone 400 and the receiver 300 are already connected, the processing of step S200 may be omitted. Here, when a plurality of wireless microphones 400 are arranged around the camera 100, the receiver 300 communicates with each of the plurality of wireless microphones 400 and performs control to connect to each wireless microphone 400.
[0051] In step S201, the camera control unit 101 determines whether the receiver control unit 301 can communicate with the microphone control unit 401. If the camera control unit 101 determines that the receiver control unit 301 can communicate with the microphone control unit 401, the process proceeds to step S203. If the camera control unit 101 determines that the receiver control unit 301 cannot communicate wirelessly with the microphone control unit 401, the camera control unit 101 stores error information indicating that there is an abnormality in the communication state in its built-in memory and proceeds to step S202. Also, in the microphone control unit 401, if it is determined that communication with the receiver control unit 301 is not possible, error information indicating that there is an abnormality in the wireless communication state is stored in the built-in memory of the microphone control unit 401.
[0052] In step S202, based on the error information stored in the built-in memory, the camera control unit 101 displays a warning on the camera display unit 106, or the receiver control unit 301 displays a warning on the receiver display unit 304 to notify the user. Also, in the microphone control unit 401, based on the error information stored in the built-in memory, a warning is displayed on the microphone display unit 404 to notify the user. After displaying the warning, the camera control unit 101 returns the process to step S200.
[0053] In step S203, the camera control unit 101 acquires information on the wireless microphone 400 from the microphone control unit 401 by the receiver control unit 301 and stores it in the built-in memory of the camera control unit 101. Also, when a plurality of wireless microphones 400 are arranged around the camera 100, the camera control unit 101 stores the information received from each wireless microphone in the built-in memory. The information of the wireless microphone 400 is, for example, device information, operation information, and setting information. The device information is unique identification information of the wireless microphone 400, etc., the operation information is battery remaining amount information, etc., and the setting information is sound collection information such as mono, stereo, and mute (silence). Further, the camera control unit 101 performs status display on the camera display unit 106 or on the receiver display unit 304 by the receiver control unit 301 based on the information acquired from the wireless microphone 400. Also, the microphone control unit 401 performs status display based on the information of the wireless microphone 400 so that the user can check the operation state of the wireless microphone 400 and the connection state with the receiver 300, etc.
[0054] In step S204, the camera control unit 101 performs recording settings for synthesizing the audio data received from the wireless microphone 400 with the image data acquired by the imaging unit 102 according to the setting operation performed by the user using the camera operation unit 105. If the recording settings have already been performed, the process of step S204 may be omitted or updated to new recording settings. When the recording settings in step S204 are not set, the default settings stored in the built-in memory of the camera control unit 101 are applied. The recording settings include settings for the channel for recording the audio data received from the wireless microphone 400, the recording hold timer when subject detection becomes impossible during recording, settings for the processing after timeout, etc.
[0055] Here, with reference to FIG. 3, a recording setting method for setting the channels for recording the audio data of each wireless microphone 400 when two wireless microphones 400 are connected to the receiver 300 will be described.
[0056] FIG. 3 illustrates a UI screen of the camera 100 for performing recording settings displayed on the camera display unit 106. FIG. 3 illustrates a state where "MIX" is selected in the recording settings for channel L (left).
[0057] Hereinafter, two wireless microphones 400 connected to the receiver 300 will be described by rephrasing them as microphone_ID1 and microphone_ID2 using the identification information of each wireless microphone.
[0058] On the screen of FIG. 3, it is possible to set the "channel" for recording the audio data of the wireless microphone 400. In FIG. 3, the L channel and the R channel are displayed as the settable channel items. The user selects one of the channels. Also, when a channel is selected, the setting items for the audio data to be assigned to each selected channel are displayed. In FIG. 3, as the audio data to be assigned to the L channel, "MIX", "microphone_ID1", and "microphone_ID2" are displayed by a pull-down menu. As shown in FIG. 3, when "MIX" is selected, the audio data of microphone_ID1 and the audio data of microphone_ID2 are synthesized (mixed) and assigned to channel L for recording. Also, in this case, the other channel R is also automatically in the state where "MIX" is selected. When "MIX" is not selected, "microphone_ID1" or "microphone_ID2" can be selected as the audio data to be recorded for each channel. In this case, the respective audio data of microphone_ID1 and microphone_ID2 are assigned to the set channel for recording.
[0059] Also, on the screen of FIG. 3, it is possible to set the "recording hold timer". In the present embodiment, during recording, when subject detection becomes impossible after performing the association setting between the wireless microphone and the subject described later, the time for continuing the recording of the audio data of the wireless microphone associated with the subject being set can be set. The time in this case can be set using the "recording hold timer". In FIG. 3, the set time of the "recording hold timer" is set to 10 seconds, but it can be set to an arbitrary time.
[0060] Furthermore, on the screen of FIG. 3, it is possible to set "processing after timeout". In this embodiment, it is possible to set the processing after the set time of the "recording hold timer" has elapsed. On the screen of FIG. 3, "mute" and "maintain recording" can be selected from a pull-down menu as the "processing after timeout" when the subject is not detected again within the set time of the "recording hold timer". When the setting by the screen of FIG. 3 is completed, the user operates the camera operation unit 105 to instruct the end of the setting.
[0061] Hereinafter, as an example of the recording setting in FIG. 3, an example will be described in which the audio data recorded in "channel" is "MIX", the "recording hold timer" is "10 seconds", and the "processing after timeout" is set to "mute".
[0062] When the setting by the screen of FIG. 3 is completed, in step S205, the camera control unit 101 executes an autofocus (AF operation) to focus on a specific subject in the captured image. When the user turns the camera 100 toward the subject and the subject is displayed in the live view image, the camera control unit 101 calculates focus information such as defocus information indicating the focus state of the subject image based on the image captured by the imaging unit 102. The camera control unit 101 calculates control information for driving the lens 202 of the lens unit 200 to a focused state based on the focus information, and transmits it to the lens control unit 201. The lens control unit 201 drives the focus lens included in the lens 202 of the lens unit 200 based on the control information received from the camera control unit 101 via the lens connection unit 103.
[0063] In step S206, the camera control unit 101 executes subject detection processing to detect a subject from the image captured by the imaging unit 102. The subject detection method is executed by pattern matching using parts such as a person's face, eyes, and mouth, and the size of the detected subject (for example, the face) can be compared. In the following description, an example of detecting a person's face as the subject detection processing will be described, but it is not limited to this, and parts other than a person's face or a person other than a person may be detected.
[0064] In step S207, the camera control unit 101 performs an association setting for associating the wireless microphone 400 connected in step S200 with the subject detected in step S206 according to the setting operation performed by the user using the camera operation unit 105. The control unit 101 displays a screen for the association setting shown in FIG. 4 on the camera display unit 106. Here, for example, among the subjects detected in step S206, the subject wearing the wireless microphone 400 connected to the receiver 300 is associated with the wireless microphone 400 worn as the subject to be recorded. If the association setting has already been made, the processing of step S207 may be omitted. If the association setting is not set, the determination of the recording level based on the subject position and the recording gain based on the subject size, which will be described later, is not performed. Based on the recording setting set in step S204, the audio data received from each wireless microphone 400 is recorded on the channel set for recording.
[0065] Here, with reference to FIG. 4, a method of associating and setting a subject in the captured image with the wireless microphone 400 will be described.
[0066] FIG. 4 illustrates a display screen (UI screen) displayed on the camera display unit 106 of the camera 100 when performing the association setting. FIG. 4(a) illustrates a state in which the microphone_ID1 is associated with the central subject among the three subjects in the captured image. FIG. 4(b) illustrates a state in which the microphone_ID2 is associated with the left subject among the three subjects in the captured image.
[0067] On the screen of FIG. 4, the user can move the AF frame with the camera operation unit 105 to select any one of the three subjects on the screen, and can further select either "Microphone_ID1" or "Microphone_ID2" at the bottom of the screen.
[0068] The example of FIG. 4(a) is the case where the central subject is equipped with Microphone_ID1. When the user selects the central subject, the AF frame is displayed as a solid line, and the left and right subjects are displayed with the AF frame as a dotted line and can be selected. Also, in the example of FIG. 4(a), the "Microphone_ID1" selected by the user is displayed in the selected state at the bottom of the screen.
[0069] The example of FIG. 4(b) is the case where the left subject is equipped with Microphone_ID2. When the user selects the left subject, the AF frame is displayed as a solid line. Also, in the example of FIG. 4(b), the "Microphone_ID2" selected by the user is displayed in the selected state at the bottom of the screen. Also, when making the association setting of FIG. 4(b), if the association setting of FIG. 4(a) is not completed, the central subject and the right subject are displayed with the AF frame as a dotted line and can be selected. On the other hand, when making the association setting of FIG. 4(b), if the association setting of FIG. 4(a) is completed, the central subject has already completed the association setting and thus cannot be selected, and the dotted line of the AF frame of the central subject is grayed out. When the association setting is completed, the user operates the camera operation unit 105 to instruct the completion of the association setting. When the association setting is completed, the camera control unit 101 switches the image displayed on the camera display unit 106 to a normal live view image.
[0070] When there is an instruction to complete the association setting, the camera control unit 101 determines, in step S208, whether the association setting has been performed for the information of the wireless microphone 400 acquired in step S203 in step S207. If the camera control unit 101 determines that the association setting has been performed for all the connected wireless microphones 400, the process proceeds to step S210. If the camera control unit 101 determines that the association setting has not been performed for all the connected wireless microphones 400, the camera control unit 101 stores error information indicating that there is a wireless microphone 400 for which the association setting has not been performed in its built-in memory.
[0071] In step S209, based on the error information stored in the built-in memory, the camera control unit 101 gives a warning display on the camera display unit 106, or the receiver control unit 301 gives a warning display on the receiver display unit 304 to notify the user. After giving the warning display, the camera control unit 101 returns the process to step S207.
[0072] In step S210, the camera control unit 101 determines whether the video shooting button included in the camera operation unit 105 has been pressed and the REC signal has been received. If the camera control unit 101 determines that the REC signal has been received from the camera operation unit 105, the process proceeds to step S211. If the camera control unit 101 determines that the REC signal has not been received from the camera operation unit 105, the process returns to step S200, and the processes from step S200 to S208 are repeated until the REC signal is received.
[0073] In step S211, the camera control unit 101 controls each component of the camera 100 to start the video shooting process. The camera control unit 101 stores the image data generated by the imaging unit 102 in the built-in memory at a predetermined frame rate. Also, the camera control unit 101 sends an instruction to start video shooting and an instruction to start recording to the receiver control unit 301 and the microphone control unit 401. Further, the camera control unit 101 displays on the camera display unit 106 that video shooting is in progress.
[0074] In step S212, the camera control unit 101 transmits control information for starting the sound collection process to the wireless microphone 400 via the receiver 300. In accordance with the start instruction for sound collection received from the camera control unit 101, the wireless microphone 400 starts generating audio data by the microphone voice input unit 405 under the control of the microphone control unit 401. The microphone control unit 401 transmits the audio data to the camera control unit 101 via the receiver control unit 301, and the camera control unit 101 stores the audio data received from the wireless microphone 400 in the built-in memory of the camera control unit 101.
[0075] In step S213, the camera control unit 101 performs a subject detection process on each frame of the video captured by the imaging unit 102, and stores subject position information indicating the position of the detected subject within each frame in the built-in memory. If the channel setting in the recording setting of step S204 is one channel or the association setting of step S207 is not set, the process of step S213 may be omitted.
[0076] In step S214, based on the subject position information stored in the built-in memory, the camera control unit 101 determines the recording level of the audio data of each of the microphone_ID1 and microphone_ID21 when recording the audio data on the set channels, and stores it in the built-in memory. The recording level indicates the ratio or proportion of the audio data when recording the audio data of each of the microphone_ID1 and microphone_ID21 on the set channels. The recording level is a value representing the magnitude of the signal of the audio data recorded for each channel, and corresponds to the volume according to the position of the sound source. If the channel setting in the recording setting of step S204 is a single channel or the association setting of step S207 is not set, the process of step S214 may be omitted. When the channel setting of step S204 or the association setting of step S207 is a single channel, the recording level for each channel is set to a fixed value so that the audio data is recorded only on one channel. Also, when the channel setting of step S204 or the association setting of step S207 is not set, the recording level for each channel is set to a fixed value so that the audio data is evenly recorded on each channel.
[0077] Here, with reference to FIG. 5, a method for determining the recording level of the audio data to be recorded for each channel based on the subject position will be described.
[0078] FIG. 5(a) illustrates a live view screen displayed on the camera display unit 106 during video shooting of a first subject associated with microphone_ID1 and a second subject associated with microphone_ID2 in the association setting of step S207. Also, FIG. 5(a) schematically shows the recording level when the recording setting of step S204 is "MIX" and the audio data received from microphone_ID1 and microphone_ID2 are recorded on channel L and channel R. FIG. 5(b) illustrates a state where, from the state of FIG. 5(a), the first subject and the second subject have moved to the right side of the angle of view, and the first subject associated with microphone_ID1 is no longer included in the shooting range (angle of view).
[0079] In the example of Fig. 5(a), the first subject associated with Mic_ID1 is located near the center of the shooting range, and the second subject associated with Mic_ID2 is located on the left side of the shooting range. In this case, in the subject position determination of step S213, the center of the detection range (circumscribed rectangle) of the face of each subject becomes the subject position (the position of the dashed line in Fig. 5(a)). In the example of Fig. 5(a), the position of the first subject associated with Mic_ID1 is at the position of W1 from the left end of the screen (W2 from the right end of the screen), and the position of the second subject associated with Mic_ID2 is at the position of W3 from the left end of the screen (W4 from the right end of the screen). The directions of arrows 1 to 4 indicate channel L or channel R in which the voice data of the subject acquired by each of the microphone voice input units 405 of Mic_ID1 and Mic_ID2 is recorded. The thicknesses of arrows 1 to 4 indicate the recording levels of the voice data acquired by each of the microphone voice input units 405 of Mic_ID1 and Mic_ID2. In the example of Fig. 5(a), the relationship of the thicknesses of arrows 1 to 4 is "arrow 3 > arrow 2 ≒ arrow 1 > arrow 4", which is represented by the following formula 1. (Formula 1) TIFF2025088427000002.tif8982In the example of Fig. 5(b), the first subject associated with Mic_ID1 is located on the right side of the shooting range, and the second subject associated with Mic_ID2 is located on the right side of the shooting range. In this case, the recording level of arrow 4 changes to the relationship of "arrow 4 > arrow 3" according to the third and fourth formulas of formula 1. For arrows 1 and 2, by substituting "W2 = 0" into formula 1, arrow 1 disappears and only arrow 2 remains. Therefore, the recording levels of arrows 1 to 4 are in the relationship of "arrow 2 > arrow 4 > arrow 3 > arrow 1 = 0". However, the first subject associated with Mic_ID1 cannot be detected from the shooting range. In this case, the camera control unit 101 transmits control information to the microphone control unit 401 via the receiver control unit 301 so as to shift to the "mute" state according to the setting of the post-timeout process after 10 seconds set by the recording hold timer in step S204.
[0080] In step S215, the camera control unit 101 executes a process of detecting the size of a subject detected from the image captured by the imaging unit 102, and stores the subject size information in the built-in memory.
[0081] In step S216, the camera control unit 101 determines a gain (recording gain) for correcting the recording level when recording each audio data of the microphone_ID1 and the microphone_ID2 on the set channels based on the subject size stored in the built-in memory, and stores it in the built-in memory.
[0082] Here, with reference to FIG. 6, a method of determining the recording gain based on the subject size will be described.
[0083] FIG. 6 illustrates a live view screen displayed on the camera display unit 106 during video shooting of a first subject associated with the microphone_ID1 and a second subject associated with the microphone_ID2 in the association setting of step S207, similar to FIG. 5. Further, FIG. 6 schematically shows the recording level when the recording setting in step S204 is "MIX" and the audio data received from the microphone_ID1 and the microphone_ID2 are recorded on the channel L and the channel R. The first subject associated with the microphone_ID1 exists on the left back of the road extending towards the left back of the shooting range, and the second subject associated with the microphone_ID2 exists in front of the right side of the same road. In this case, in the subject size determination in step S215, each face detection range (circumscribed rectangle) is the subject size. Let the area of the face detection range of the first subject associated with the microphone_ID1 be κ, and the area of the face detection range of the second subject associated with the microphone_ID2 be λ. In this case, considering the recording level of the channels for recording the audio data in Equation 1, the recording levels of arrows 1 to 4 are represented by the following Equation 2 based on the recording gain determined in step S216. (Equation 2) TIFF2025088427000003.tif89116I1 indicates the signal level of the audio data acquired by microphone_ID1. I2 indicates the signal level of the audio data acquired by microphone_ID2. m and n are offset adjustment amounts, which vary for each wireless microphone. When the distance in the depth direction of the subject (distance to camera 100) is the same (κ = λ), they are automatically adjusted so that "I1 + m = I2 + n". In the example of Figure 6, the recording levels of arrow 1 to arrow 4 have the relationship of "arrow 4 > arrow 3 > arrow 1 > arrow 2".
[0084] In step S217, the camera control unit 101 corrects the recording level determined in step S214 for the audio data stored in the built-in memory based on the recording gain determined in S216, synthesizes it with the image data stored in the built-in memory, and stores it in the built-in memory. In this case, the camera control unit 101 synthesizes so that the recording start time of the audio data coincides with the recording start time of the image data, and also adds device information and setting information of the camera 100 and the wireless microphone 400 to the audio data or the image data.
[0085] In step S218, the camera control unit 101 encodes the audio data and video data synthesized in step S217 stored in the built-in memory, stores them in a video file, and saves them in the storage unit 107.
[0086] In step S219, the camera control unit 101 determines whether the video shooting button included in the camera operation unit 105 has been pressed and the REC signal has been received. If the camera control unit 101 determines that the REC signal has been received from the camera operation unit 105, the process proceeds to step S220. If the camera control unit 101 determines that the REC signal has not been received from the camera operation unit 105, the process returns to step S213, and the processes from step S213 to S218 are repeated until the REC signal is received.
[0087] In step S220, after the camera control unit 101 stores all the audio data and image data stored in the built-in memory in the storage unit 107, the camera control unit 101 controls each component of the camera 100 to end the video shooting process. Further, the camera control unit 101 transmits an instruction to end the video shooting and an instruction to end the recording to the receiver control unit 301 and the microphone control unit 401. Furthermore, the camera control unit 101 turns off the display during video shooting on the camera display unit 106.
[0088] According to Embodiment 1, the recording level for each channel when recording the audio data acquired by the wireless microphone 400 associated with the subject on the set channel is adjusted based on the position and size of the subject. Thereby, it is possible to reduce the sense of discomfort given to the user due to the mismatch between the image and the audio.
[0089] [Embodiment 2] Next, with reference to FIGS. 7 and 8, the recording / recording process of the camera 100 according to Embodiment 2 will be described.
[0090] Embodiment 2 picks up the ambient sound around the camera 100 by the receiver audio input unit 305 and synthesizes it with the audio data acquired by the microphone audio input unit 405. In Embodiment 2, an example of picking up the ambient sound by the receiver audio input unit 305 will be described, but it may be picked up by the camera audio input unit 110. Further, by the user operating the camera operation unit 105, the audio input unit for picking up the ambient sound may be switchable between the camera audio input unit 110 and the receiver audio input unit 305. Alternatively, also, by the user operating the camera operation unit 105, the recording of the ambient sound of the camera audio input unit 110 or the receiver audio input unit 305 may be switchable. Also, when the connection of the receiver 300 or the wireless microphone 400 is unexpectedly disconnected, the audio input unit may be switched to the camera audio input unit 110 so that recording can be continued.
[0091] FIG. 7 is a flowchart illustrating the recording / recording process of the camera 100 according to Embodiment 2.
[0092] Steps S700 to S711 are the same as the processes of steps S200 to S211 in FIG. 2.
[0093] In step S712, the camera control unit 101 transmits control information for starting the sound collection process to the receiver 300 and the wireless microphone 400 via the receiver 300. The microphone control unit 401 starts generating voice data by the microphone voice input unit 405 according to the start instruction of the sound collection process received from the camera control unit 101. The microphone control unit 401 transmits the voice data to the camera control unit 101 via the receiver control unit 301, and the camera control unit 101 stores the voice data received from the wireless microphone 400 in the built-in memory of the camera control unit 101. Also, the receiver control unit 301 starts generating voice data by the receiver voice input unit 305 according to the start instruction of the sound collection process received from the camera control unit 101. The receiver control unit 301 transmits the voice data to the camera control unit 101, and the camera control unit 101 stores the voice data received from the receiver 300 in the built-in memory of the camera control unit 101.
[0094] Steps S713 to S715 are the same as the processes of steps S213 to S215 in FIG. 2.
[0095] In step S716, the camera control unit 101 determines a recording gain for correcting the recording level when recording each voice data of the wireless microphone 400 and the receiver 300 on the set channel based on the subject size stored in the built-in memory, and stores it in the built-in memory.
[0096] Here, with reference to FIG. 8, the determination method of the subject size in step S715 and the determination method of the recording gain in step S716 will be described.
[0097] FIG. 8, similar to FIG. 6, illustrates a live view screen displayed on the camera display unit 106 during video shooting of a first subject associated with microphone_ID1 and a second subject associated with microphone_ID2 in the association setting of step S707. Further, FIG. 8 schematically shows the recording levels when the recording setting in step S704 is "MIX" and the audio data received from microphone_ID1 and microphone_ID2 are recorded on channel L and channel R. In FIG. 8, different from FIG. 6, the recording levels when the audio data of the receiver audio input unit 305 are recorded on channel L and channel R are shown. In this case, similar to FIG. 6, the recording levels of arrow 1 to arrow 6 are represented by the following formula 3 based on the recording gain determined in step S716. (Formula 3) TIFF2025088427000004.tif115133Arrow 5 and arrow 6 indicate the recording levels of the ambient sound acquired by the receiver audio input unit 305. I3 indicates the signal level of the audio data acquired by the receiver audio input unit 305. γ indicates the recording gain of the audio data acquired by the receiver audio input unit 305. l indicates the offset adjustment amount of the audio signal acquired by the receiver audio input unit 305.
[0098] The offset adjustment amount l is automatically set so that when the distances in the depth direction between the subjects are the same (κ = λ), the signal levels of the audio data acquired by each audio input unit are "I1 + m = I2 + n > I3 + l". Thereby, it is possible to prevent a state where the ambient sound is so loud that the conversation between two subjects cannot be heard when the two subjects are talking side by side.
[0099] Also, the recording gain γ of the audio data acquired by the receiver audio input unit 305 is set such that the recording level of the audio data acquired by the receiver audio input unit 305 is lower than that of the audio data of a subject near the camera 100 (receiver 300). In other words, the recording level of the audio data of a subject near the camera 100 (receiver 300) is set to be higher than the recording level of the audio data acquired by the receiver audio input unit 305.
[0100] Thereby, the recording levels of the audio data of a subject near the camera 100 (receiver 300) and the audio data acquired by the receiver audio input unit 305 are set so as not to be reversed (when λ ≧ κ, λ / κ > γ).
[0101] Similarly, the recording gain γ of the audio data acquired by the receiver audio input unit 305 is set such that the recording level of the audio data acquired by the receiver audio input unit 305 is higher than that of the audio data of a subject far from the camera 100 (receiver 300). In other words, the recording level of the audio data of a subject far from the camera 100 (receiver 300) is set to be lower than the recording level of the audio data acquired by the receiver audio input unit 305.
[0102] Thereby, the recording levels of the audio data of a subject far from the camera 100 (receiver 300) and the audio data acquired by the receiver audio input unit 305 are set so as not to be reversed.
[0103] Note that a reference subject size may be set in advance, and the recording levels of the audio data of the subject and the audio data acquired by the receiver audio input unit 305 may be adjusted based on the comparison with the reference subject size. In this case, when the recording gain in the case where the reference subject size is the area ε is γ as in the ambient sound, the relationship between the recording levels of the wireless microphone 400 and the receiver 300 is represented by the following formula 4. (Formula 4) Considering a relationship such as 115147λ≧ε≧κ, when the subject size of the recording target is equal to the reference subject size, the recording gains of arrows 1 to 6 are γ. The larger the area λ of the subject size closer to the camera 100 than the area ε of the reference subject size, the larger the recording gain of the audio signal of the microphone_ID2. The smaller the area κ of the subject size farther from the camera 100 than the area ε of the reference subject size, the smaller the recording gain of the audio signal of the microphone_ID1.
[0104] In the example of FIG. 8, the relationship is "arrow 4 > arrow 3 > arrow 5 = arrow 6 > arrow 1 > arrow 2". The receiver voice input unit 305 may be either a monaural microphone or a stereo microphone. From Equation 3 or Equation 4, even if it is a stereo microphone, the recording gain γ of the audio data of the receiver voice input unit 305 is common for channel L and channel R.
[0105] However, since the signal level I3 of the audio data of the receiver voice input unit 305 changes, for example, when a large ambient sound is input from the right side, the signal levels of the audio data in channel L and channel R are different, so arrows 5 and 6 also have different recording levels. This is the same when recording the audio data of the camera voice input unit 110.
[0106] In step S717, the camera control unit 101 performs a voice detection process of detecting the voice data of the first subject associated with the microphone_ID1 or the voice data of the second subject associated with the microphone_ID2 from the voice data of the receiver voice input unit 305 stored in the built-in memory. The voice detection process is performed by comparing the voice data of the receiver voice input unit 305 with the voice data of the first subject associated with the microphone_ID1 or the voice data of the second subject associated with the microphone_ID2. Since the voice detection processing method is a known technique such as voice feature amount analysis, the description thereof is omitted.
[0107] When voice data of a first subject associated with microphone_ID1 or voice data of a second subject associated with microphone_ID2 is detected from the voice data of the receiver voice input unit 305, the camera control unit 101 calculates a delay amount for synchronizing the voice data of the receiver voice input unit 305 of the detected subject with the voice data of microphone_ID1 or microphone_ID2, and stores it in the built-in memory. Thereby, when the distance between the camera 100 and the subject associated with microphone_ID1 or microphone_ID2 is physically close, it is possible to prevent a state in which the voice of the subject for which the voice synthesis process described later is performed is heard twice.
[0108] Note that the method for preventing a state in which the voice of the subject is heard twice is not limited to the method of synchronizing voice data. For example, when voice data of a subject associated with microphone_ID1 or microphone_ID2 is detected from the voice data of the receiver voice input unit 305, the recording gain γ of the voice data of the receiver voice input unit 305 may be reduced to lower the recording level of the ambient sound.
[0109] In step S718, the camera control unit 101 adjusts the voice data acquired by the wireless microphone 400 stored in the built-in memory based on the recording level determined in step S714 and the recording gain determined in S716. Then, the camera control unit 101 synthesizes the image data captured by the imaging unit 102 stored in the built-in memory and the voice data synchronized based on the delay amount calculated in step S717, and stores it in the built-in memory. In this case, device information, setting information, etc. of the camera 100 are also added to the synthesized image data. When performing the process of reducing the recording gain γ of the ambient sound in step S717, the camera control unit 101 synthesizes the voice data of microphone_ID1 or microphone_ID2 and the voice data of the receiver voice input unit 305 so that the elapsed time from the start of recording matches, and stores it in the built-in memory.
[0110] Steps S719 to S721 are the same as the processes of steps S218 to S220 in FIG. 2.
[0111] According to Embodiment 2, when synthesizing and recording the audio data acquired by the wireless microphone 400 associated with the subject and the audio data including the ambient sound around the camera 100, the recording level for each channel is adjusted based on the position and size of the subject. This can reduce the sense of discomfort given to the user due to the mismatch between the image and the audio.
[0112] [Embodiment 3] Next, with reference to FIGS. 9 and 10, the recording / recording process of the camera 100 according to Embodiment 3 will be described.
[0113] In Embodiment 3, the subject distance is detected from the defocus information of the imaging unit 102, the focal length information of the lens unit 200, etc., and the recording gain is determined based on the subject distance instead of the subject size in Embodiment 2. Further, by selecting a subject for which the audio is to be emphasized during video shooting, a subject emphasis process is performed to increase the recording level of the subject selected according to the user's intention.
[0114] FIG. 9 is a flowchart illustrating the recording / recording process of the camera 100 according to Embodiment 3.
[0115] Steps S900 to S914 in FIG. 9 are the same as the processes of steps S700 to S714 in FIG. 7.
[0116] In step S915, the camera control unit 101 detects the subject distance from the defocus information of the imaging unit 102 and stores it in the built-in memory. Note that the focal length information of the lens unit 200 may be acquired from the lens control unit 201 and used instead of the subject distance information.
[0117] In step S916, the camera control unit 101 determines a gain (recording gain) for correcting the recording level when recording each audio data of the microphone_ID1 and the microphone_ID2 on the set channel based on the subject distance information stored in the built-in memory, and stores it in the built-in memory.
[0118] Here, with reference to FIG. 8 of Embodiment 2, the method for determining the subject distance in step S915 and the method for determining the recording gain in step S916 will be described.
[0119] Let the distance from the camera 100 to the first subject associated with the microphone_ID1 be x, and the distance from the camera 100 to the second subject associated with the microphone_ID2 be y. In this case, the recording levels of arrows 1 to 6 are represented by the following Equation 5 based on the recording gain γ of the receiver audio input unit 305 of Equation 4. (Equation 5) Considering a relationship such as TIFF2025088427000006.tif115122x > y, the recording gain of the wireless microphone 400 decreases as the distance from the camera 100 increases. In the example of FIG. 8, the recording gain of the audio data of microphone_ID1, which is farther from the camera 100, is smaller than the recording gain of the audio data of microphone_ID2. Conversely, the recording gain of the wireless microphone 400 increases as it approaches the camera 100. In the example of FIG. 8, the recording gain of the audio data of microphone_ID2, which is closer to the camera 100, is larger than the recording gain of the audio data of microphone_ID1. However, the balance between the subject voice and the ambient sound is automatically set so that the signal levels of the audio data acquired by each audio input unit are "I1 + m = I2 + n > I3 + l" when the distances in the far - near direction between the subjects are the same (x = y), as described in Embodiment 2. Therefore, it is possible to prevent a state where the ambient sound is too loud and the voice of the subject cannot be heard.
[0120] In step S917, the camera control unit 101 selects a first subject associated with the microphone_ID1 or a second subject associated with the microphone_ID2 when the user operates the camera operation unit 105 during video shooting, and performs subject emphasis processing setting. If the subject emphasis processing setting has already been completed or if no emphasis processing is to be performed, the processing in step S917 may be omitted. When the subject emphasis processing setting is performed, the camera control unit 101 displays on the camera display unit 106 that the selected subject is being emphasized. Further, the camera control unit 101 stores in the built-in memory an offset value in the subject emphasis processing described later in accordance with the recording time of the image and the object to be emphasized.
[0121] FIG. 10 illustrates a subject emphasis processing setting screen.
[0122] FIG. 10 illustrates a state in which a first subject associated with the microphone_ID1 is selected in the live view screen of FIG. 5(a) and the first subject is being emphasized, and a state in which the ambient sound of the receiver voice input unit 305 is being recorded. In the example of FIG. 10, a double frame is displayed for the first subject associated with the microphone_ID1 to be emphasized, and a gray frame is displayed for the subject associated with the microphone_ID2 that is not emphasized. The method for setting the subject emphasis processing is, for example, to touch, operate, and select a subject associated with a desired wireless microphone 400 in the live view screen of FIG. 5(a) using a joystick, dial, button, etc. Further, a subject that the photographer has gazed at for a predetermined time by gaze input may be set as the object to be emphasized.
[0123] In the subject emphasis processing, the recording level of the audio data of the wireless microphone associated with the subject to be processed is increased, and the recording levels of other audio data are decreased. In the example of FIG. 5(a), the relationship is "arrow 3 > arrow 2 ≒ arrow 1 > arrow 4", but in the example of FIG. 10, by performing the emphasis processing including the recording level of the receiver voice input unit 305, the relationship becomes "arrow 2 ≒ arrow 1 > arrow 3 > arrow 5 = arrow 6 > arrow 4". In this case, the recording levels of arrows 1 to 6 are represented by the following formula 6. (Formula 6) TIFF2025088427000007.tif115141 ρ, σ, and τ are offset values in the subject enhancement process of the audio data of the wireless microphone 400 and the receiver 300. In the example of FIG. 10, since the first subject associated with the microphone_ID1 is enhanced, predetermined values where ρ > 0, σ < 0, and τ < 0 are substituted. In Equation 6, although the offset values in the subject enhancement process of the audio data of the wireless microphone 400 and the receiver 300 are different, the recording level may be decreased at the same level for audio data other than the audio data to be processed (σ = τ). By performing the subject enhancement process setting in this way, it becomes possible to adjust the recording level of the audio data recorded for each channel while adding an audio production effect during recording / recording. When the first subject associated with the microphone_ID1 is selected again, the currently executing subject enhancement process ends, and the display on the camera display unit 106 also returns to the state before the start of the subject enhancement process.
[0124] Steps S918 to S922 are the same as the processes of steps S717 to S721 in FIG. 7.
[0125] According to Embodiment 3, an audio production effect can be added by the subject enhancement process during recording / recording. Also, when adding an audio production effect, the recording level for each channel when synthesizing and recording the audio data acquired by the wireless microphone 400 associated with the subject and the audio data including the ambient sound around the camera 100 is adjusted based on the position and size of the subject. Thereby, it is possible to reduce the discomfort given to the user due to the mismatch between the image and the audio.
[0126] Each process in the flowchart of each of the above-described embodiments is an example, and as long as the present embodiment can be realized, the order of the processes in each flowchart may be appropriately changed and executed.
[0127] [Other Embodiments] The present invention can also be implemented by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and causing one or more processors in a computer of the system or device to read and execute the program. It can also be implemented by a circuit (for example, ASIC) that realizes one or more functions.
[0128] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the claims are attached to disclose the scope of the invention.
[0129] The disclosure of this specification includes the following imaging device, control method, and program. [Configuration 1] An imaging device, imaging means, connection means for connecting to a sound collection device, synthesis means for synthesizing the audio data acquired from the sound collection device and the video data acquired by the imaging means, detection means for detecting a subject included in a screen of a video imaged by the imaging means, setting means for associating the subject with a sound collection device connected to the imaging device, determination means for determining a position of the subject detected by the detection means within a screen of the video imaged by the imaging means, determination means for determining, based on the position of the subject determined by the determination means, a recording level for each channel of the audio data acquired from the sound collection device associated with the subject, and an imaging device characterized by comprising the same. [Configuration 2] The recording level is a ratio of audio data recorded for each channel, the determination means determines the ratio based on the position of the subject determined by the determination means, the synthesis means synthesizes the audio data acquired from the sound collection device with the imaged video data based on the ratio determined by the determination means, and the imaging device according to Configuration 1, characterized by the same. [Configuration 3] The detection means detects the size of a subject included in the screen of the video imaged by the imaging means, Based on the size of the subject detected by the detection means, the determination means determines a gain for correcting the recording level, and the imaging apparatus according to Configuration 1 or 2. [Configuration 4] The imaging apparatus according to Configuration 3, wherein the determination means determines the gain based on the ratio between the size of a first subject and the size of a second subject detected by the detection means. [Configuration 5] When the subject associated with the sound collection device is no longer detected by the detection means, the setting means can set to mute the sound collection device associated with the subject after a predetermined time has elapsed, or can set to continue recording the audio data of the channel in a direction outside the shooting range. The imaging apparatus according to any one of Configurations 1 to 4. [Configuration 6] The imaging apparatus according to any one of Configurations 1 to 5, wherein it is possible to set to mix and record the audio data acquired from a plurality of sound collection devices connected to the imaging apparatus for each channel, or to record without mixing. [Configuration 7] When the channel for recording the audio data acquired from the sound collection device is one channel, the determination means sets the recording level so that the audio data is recorded only on that channel. The imaging apparatus according to any one of Configurations 1 to 5. [Configuration 8] When the association by the setting means is not set, the determination means sets the recording level so that the audio data is evenly recorded for each channel. The imaging apparatus according to any one of Configurations 1 to 5. [Configuration 9] It has a sound collection unit, and a receiving device that receives audio data from the sound collection device can be connected, Of the subjects detected by the detection means, the recording level of the audio data acquired by the sound collection unit is made smaller than the recording level of the audio data acquired by the sound collection device associated with the subject close to the imaging device, and the recording level of the audio data acquired by the sound collection unit is made larger than the recording level of the audio data acquired by the sound collection device associated with the subject far from the imaging device. The imaging device according to any one of Configurations 1 to 8, characterized in that. [Configuration 10] It has a sound collection unit, and a receiving device for receiving audio data from the sound collection device is connectable, Of the subjects detected by the detection means, the recording level of the audio data acquired by the sound collection device associated with the subject close to the imaging device is made larger than the recording level of the audio data acquired by the sound collection unit, and the recording level of the audio data acquired by the sound collection device associated with the subject far from the imaging device is made smaller than the recording level of the audio data acquired by the sound collection unit. The imaging device according to any one of Configurations 1 to 8, characterized in that. [Configuration 11] It has a sound collection unit, and a receiving device for receiving audio data from the sound collection device is connectable, When the size of the subject detected by the detection means is larger than a predetermined size, the recording level of the audio data acquired by the sound collection device associated with the subject is made larger than the recording level of the audio data acquired by the sound collection unit, and when the size of the subject detected by the detection means is smaller than the predetermined size, the recording level of the audio data acquired by the sound collection device associated with the subject is made smaller than the recording level of the audio data acquired by the sound collection unit. The imaging device according to any one of Configurations 1 to 8, characterized in that. [Configuration 12] It has a sound collection unit, and a receiving device for receiving audio data from the sound collection device is connectable, When the size of the subject detected by the detection means is larger than a predetermined size, the recording level of the audio data acquired by the sound collection unit is made smaller than the recording level of the audio data acquired by the sound collection device associated with the subject. The imaging device according to any one of configurations 1 to 8, wherein when the size of the subject detected by the detection means is smaller than the predetermined size, the recording level of the audio data acquired by the sound collection unit is made larger than the recording level of the audio data acquired by the sound collection device associated with the subject. [Configuration 13] The imaging device according to any one of configurations 9 to 12, wherein when the audio data acquired by the sound collection unit and the audio data acquired by the sound collection device include audio of the same sound source, the audio data of the same sound source is synchronized and synthesized with the moving image data by the synthesizing means. [Configuration 14] The imaging device according to any one of configurations 9 to 12, wherein when the audio data acquired by the sound collection unit and the audio data acquired by the sound collection device include audio of the same sound source, the recording level of the audio data of the sound collection unit is reduced and synthesized with the moving image data by the synthesizing means. [Configuration 15] The detection means detects the distance between the imaging device and the subject included in the image captured by the imaging means. The imaging device according to configuration 1 or 2, wherein based on the distance between the imaging device and the subject associated with the sound collection device, the determination means determines a gain for correcting the recording level. [Configuration 16] The imaging device according to configuration 15, wherein the determination means reduces the gain of the recording level of the audio data acquired by the sound collection device associated with the subject as the subject associated with the imaging device and the sound collection device moves away, and increases the gain of the recording level of the audio data acquired by the sound collection device associated with the subject as the subject associated with the imaging device and the sound collection device approaches. [Configuration 17] The setting means can select a specific subject from the subjects associated with the sound collection device, The imaging device according to any one of configurations 1 to 16, wherein the determination means makes the recording level of the audio data acquired by the sound collection device associated with the subject selected by the setting means higher than the recording levels of other audio data. [Configuration 18] The imaging device according to any one of configurations 1 to 17, wherein the channel includes a channel corresponding to the left side of the imaging device and a channel corresponding to the right side of the imaging device. [Configuration 19] The imaging device according to any one of configurations 1 to 18, further comprising notification means for notifying that none of the sound collection devices connected to the imaging device are associated with the subject included in the captured image. [Configuration 20] A control method for an imaging device, wherein the imaging device includes imaging means, connection means for connecting to a sound collection device, synthesis means for synthesizing the audio data acquired from the sound collection device and the video data acquired by the imaging means, and detection means for detecting a subject included in the screen of the video captured by the imaging means, and the control method includes a step of associating the subject with the sound collection device connected to the imaging device, a step of determining the position of the subject detected by the detection means within the screen of the video captured by the imaging means, and a step of determining the recording level for each channel of the audio data acquired from the sound collection device associated with the subject based on the determined position of the subject. [Configuration 21] A program for causing a computer to function as the imaging device according to any one of configurations 1 to 19.
Description of Reference Numerals
[0130] 100… Camera, 101… Camera control unit, 300… Receiver, 301… Receiver control unit, 305… Receiver voice input unit, 400… Wireless microphone, 401… Microphone control unit, 405… Microphone voice input unit
Claims
1. An imaging device, comprising: imaging means; connection means for connecting to a sound collection device; synthesis means for synthesizing audio data acquired from the sound collection device and video data acquired by the imaging means; detection means for detecting a subject included in a screen of a video imaged by the imaging means; setting means for associating the subject with a sound collection device connected to the imaging device; determination means for determining a position of the subject detected by the detection means within a screen of a video imaged by the imaging means; determination means for determining a recording level for each channel of audio data acquired from a sound collection device associated with the subject based on the position of the subject determined by the determination means. The imaging device is characterized by comprising the above components.
2. The recording level is a ratio of audio data recorded for each channel, the determination means determines the ratio based on the position of the subject determined by the determination means, and the synthesis means synthesizes the audio data acquired from the sound collection device with the imaged video data based on the ratio determined by the determination means. The imaging device according to claim 1 is characterized by the above features.
3. The detection means detects a size of a subject included in a screen of a video imaged by the imaging means, and based on the size of the subject detected by the detection means, the determination means determines a gain for correcting the recording level. The imaging device according to claim 1 is characterized by the above features.
4. The determination means determines the gain based on a ratio of a size of a first subject detected by the detection means to a size of a second subject. The imaging device according to claim 3 is characterized by the above features.
5. When a subject associated with the sound collection device is no longer detected by the detection means, the setting means can set to mute the sound collection device associated with the subject after a predetermined time has elapsed, or can set to continue recording audio data of a channel in a direction outside the shooting range. The imaging device according to claim 1 is characterized by the above features.
6. The imaging device according to claim 1 is characterized in that it is possible to set to mix or not mix audio data acquired from a plurality of sound collection devices connected to the imaging device for recording for each channel.
7. When the channel for recording the audio data acquired from the sound collection device is one channel, the determination means sets the recording level so that the audio data is recorded only on that channel. The imaging device according to claim 1, characterized in that.
8. When the association by the setting means is not set, the determination means sets the recording level so that the audio data is evenly recorded for each channel. The imaging device according to claim 1, characterized in that.
9. It has a sound collection unit, and a receiving device that receives audio data from the sound collection device can be connected. Among the subjects detected by the detection means, the determination means makes the recording level of the audio data acquired by the sound collection device associated with the subject close to the imaging device smaller than the recording level of the audio data acquired by the sound collection unit. The imaging device according to claim 1, characterized in that the recording level of the audio data acquired by the sound collection unit is made larger than the recording level of the audio data acquired by the sound collection device associated with the subject far from the imaging device.
10. It has a sound collection unit, and a receiving device that receives audio data from the sound collection device can be connected. Among the subjects detected by the detection means, the determination means makes the recording level of the audio data acquired by the sound collection device associated with the subject close to the imaging device larger than the recording level of the audio data acquired by the sound collection unit. The imaging device according to claim 1, characterized in that the recording level of the audio data acquired by the sound collection device associated with the subject far from the imaging device is made smaller than the recording level of the audio data acquired by the sound collection unit.
11. It has a sound collection unit, and a receiving device that receives audio data from the sound collection device can be connected. When the size of the subject detected by the detection means is larger than a predetermined size, the recording level of the audio data acquired by the sound collection device associated with the subject is made larger than the recording level of the audio data acquired by the sound collection unit. The imaging device according to claim 1, characterized in that when the size of the subject detected by the detection means is smaller than the predetermined size, the recording level of the audio data acquired by the sound collection device associated with the subject is made smaller than the recording level of the audio data acquired by the sound collection unit.
12. It has a sound collection unit and is connectable to a receiving device that receives voice data from the sound collection device. When the size of the subject detected by the detection means is larger than a predetermined size, the recording level of the voice data acquired by the sound collection unit is made smaller than the recording level of the voice data acquired by the sound collection device associated with the subject. The imaging device according to claim 1, wherein when the size of the subject detected by the detection means is smaller than the predetermined size, the recording level of the voice data acquired by the sound collection unit is made larger than the recording level of the voice data acquired by the sound collection device associated with the subject.
13. The imaging device according to any one of claims 9 to 12, wherein when the voice data acquired by the sound collection unit and the voice data acquired by the sound collection device include voices of the same sound source, the voice data of the same sound source is synchronized and synthesized with the moving image data by the synthesizing means.
14. The imaging device according to any one of claims 9 to 12, wherein when the voice data acquired by the sound collection unit and the voice data acquired by the sound collection device include voices of the same sound source, the recording level of the voice data of the sound collection unit is reduced and synthesized with the moving image data by the synthesizing means.
15. The detection means detects the distance between the imaging device and the subject included in the image captured by the imaging means. The imaging device according to claim 1, wherein based on the distance between the imaging device and the subject associated with the sound collection device, the determination means determines a gain for correcting the recording level.
16. The imaging device according to claim 15, wherein the determination means makes the gain of the recording level of the voice data acquired by the sound collection device associated with the subject smaller as the subject associated with the imaging device and the sound collection device moves away, and makes the gain of the recording level of the voice data acquired by the sound collection device associated with the subject larger as the subject associated with the imaging device and the sound collection device approaches.
17. The setting means can select a specific subject from the subject associated with the sound collection device. The imaging device according to claim 1, wherein the determining means makes the recording level of the audio data acquired by the sound recording device associated with the subject selected by the setting means higher than the recording level of other audio data.
18. The imaging device according to claim 1, wherein the channels include a channel corresponding to the left side of the imaging device and a channel corresponding to the right side of the imaging device.
19. The imaging device according to claim 1, further comprising a notifying means for notifying that all the sound recording devices connected to the imaging device are not associated with the subject included in the captured image.
20. A control method for an imaging device, wherein the imaging device comprises: imaging means; connection means for connecting to a sound recording device; synthesizing means for synthesizing the audio data acquired from the sound recording device and the video data acquired by the imaging means; detecting means for detecting a subject included in the screen of the video captured by the imaging means, and wherein the control method comprises: associating the subject with the sound recording device connected to the imaging device; determining the position of the subject detected by the detecting means within the screen of the video captured by the imaging means; and determining the recording level for each channel of the audio data acquired from the sound recording device associated with the subject based on the determined position of the subject.
21. A program for causing a computer to function as the imaging device according to any one of claims 1 to 12 and 15 to 19.
Citation Information
Patent Citations
Method and system for sound source localization using image information and storage medium storing program to realize the method
JP2000295700A
Imaging apparatus and program
JP2012151544A