Video generation method
The method allows for user-controlled emphasis or reduction of secondary audio in videos by separately recording main and secondary audio, enhancing audio quality and aligning with user intent, addressing the limitations of existing technologies in managing background noises.
Patent Information
- Application Number
- JP2024003291
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-26
- Filing Date
- 2024-01-12
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2040-07-20
AI Technical Summary
Existing video recording technologies fail to effectively separate and manage background noises such as wind noise, environmental sounds, and operation sounds from the main audio, limiting user control over audio emphasis or reduction in recorded videos.
A method and device that records both main and secondary audio separately, allowing for the generation of videos with audio where the secondary audio can be emphasized or reduced based on user input, using distinct microphones for each audio type and processing conditions, and incorporating movement detection for adaptive audio processing.
Enables user-controlled emphasis or reduction of secondary audio in videos, improving audio quality and user intent alignment by separating and editing specific sounds, even in dynamic recording conditions.
Smart Images

Figure 0007720428000001 
Figure 0007720428000002 
Figure 0007720428000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation method, an image generation device, and an image generation program. [Background technology]
[0002] When recording video and audio, sounds other than the main audio (wind noise, environmental sounds, operation sounds, speaking voices, etc.) may be recorded together.
[0003] Patent document 1 describes a video camera equipped with a means for displaying the presence or absence or strength of wind noise, a means for manually selecting the presence or strength of wind noise countermeasures, and a means for selecting the presence or absence of wind noise countermeasures even after recording.
[0004] Furthermore, Patent Document 2 describes a video camera that has a function for automatically reducing wind noise and allows the operation of this function to be set arbitrarily. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-4339 [Patent Document 2] Japanese Patent Application Laid-Open No. 2009-124414 Summary of the Invention
[0006] One embodiment of the technique of the present disclosure provides an image generation method, an image generation device, and an image generation program that can generate an image with audio in which a specific audio is emphasized or reduced. [Means for solving the problem]
[0007] (1) A video generation method comprising: a video recording step of recording a first video captured by an imaging unit; a first audio recording step of recording a first audio in synchronization with the first video; a second audio recording step of recording a second audio different from the first audio; an audio generation step of processing the first audio using the second audio to generate a third audio including the second audio that has been emphasized or reduced; and a video generation step of generating the second video by associating the first video with the third audio.
[0008] (2) The video generation method of (1), wherein the sound generation step synthesizes the emphasized or reduced second sound with the first sound to generate a third sound.
[0009] (3) A video generation method according to (2), further comprising an intensity setting step of setting the intensity of the second audio before the audio generation step, wherein the audio generation step synthesizes the second audio with the first audio at the intensity set in the intensity setting step.
[0010] (4) A video generation method according to (1), in which the first audio includes a common component that is an audio component common to the second audio, and the audio generation step uses the second audio to perform processing on the first audio to emphasize or reduce the common component, thereby generating a third audio.
[0011] (5) A video generation method according to (4), further comprising a processing condition setting step for setting processing conditions for the common component before the audio generation step, wherein the audio generation step performs processing on the first audio to emphasize or reduce the common component in accordance with the processing conditions set in the processing condition setting step.
[0012] (6) A video generation method according to any one of (1) to (5), further comprising a detection step of detecting movement of the imaging device body including the imaging unit, and wherein the audio generation step, when movement is detected in the detection step, performs a predetermined process on the first audio or the second audio to generate a third audio.
[0013] (7) The image generating method according to any one of (1) to (6), further comprising a first information acquisition step of acquiring imaging information of a first image by an imaging unit, and a first display step of displaying the imaging information.
[0014] (8) The image generating method of (7), wherein the imaging information includes at least one of information on the movement of the imaging device body including the imaging unit and information on the focal length.
[0015] (9) Any one of the video generation methods (1) to (8) further comprising a second information acquisition step of acquiring information about a sound collection section that collects the first sound and the second sound, and a second display step of displaying the information about the sound collection section.
[0016] (10) The video generating method according to any one of (1) to (9), wherein the second audio recording step records the second audio in synchronization with the first video.
[0017] (11) The video generation method of (10) further comprising a second audio detection step of detecting the timing at which the second audio was recorded, and an association step of associating the information detected in the second audio detection step with the first video.
[0018] (12) The video generation method according to any one of (1) to (11), wherein the second audio recording step records the second audio before the video recording step.
[0019] (13) A video generation method according to any one of (1) to (12), wherein the first audio recording step records the first audio via a first audio collection unit, and the second audio recording step records the second audio via a second audio collection unit different from the first audio collection unit.
[0020] (14) The image generation method of (13), wherein the second sound collection unit has directional sound collection characteristics, and the first sound collection unit has less directional sound collection characteristics than the second sound collection unit.
[0021] (15) A video generation method according to (13) or (14), wherein the second sound collection unit has directional sound collection characteristics, and the second sound generation process detects the position of the sound source of the second sound and directs the second sound collection unit in the direction of the detected sound source. [Brief explanation of the drawings]
[0022] [Figure 1]FIG. 1 is a block diagram showing the schematic configuration of an imaging device having a function for generating an image using the image generation method according to the present invention. [Figure 2] Block diagram of the main functions realized by the CPU when recording video and audio [Figure 3] Block diagram of the main functions realized by the CPU when playing back recorded video [Figure 4] Block diagram of the main functions realized by the CPU when generating video with audio [Figure 5] Block diagram of the functions of the third sound generation unit [Figure 6] Block diagram of the main functions realized by the CPU when generating video with audio [Figure 7] Block diagram of the functions of the third sound generation unit [Figure 8] FIG. 10 is a block diagram showing a schematic configuration of an imaging device according to a third embodiment. [Figure 9] Block diagram of the main functions realized by the CPU when recording video and audio [Figure 10] 10 is a block diagram of functions of a third voice generation unit according to a third embodiment; [Figure 11] FIG. 13 is a diagram illustrating a modification of the third voice generation unit of the third embodiment. [Figure 12] Block diagram of functions realized by the CPU when acquiring and recording image information and when displaying image information [Figure 13] Block diagram of the functions implemented by the CPU when acquiring and recording microphone information and when displaying microphone information [Figure 14] Block diagram of functions implemented by the CPU when detecting and recording the timing at which the second audio is recorded and when displaying the recorded information. DETAILED DESCRIPTION OF THE INVENTION
[0023] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0024] [First embodiment] FIG. 1 is a block diagram showing a schematic configuration of an imaging device having a function for generating an image using the image generation method according to the present invention.
[0025] The imaging device 1 of this embodiment records the first audio and the second audio in synchronization with imaging. After imaging, the first audio is processed using the second audio to generate a third audio containing the second audio at a predetermined intensity (audio level). The generated third audio is then associated with the video obtained by imaging (first video) to generate a video with audio (second video).
[0026] As shown in FIG. 1, the imaging device 1 includes an imaging unit 10, a first audio input unit 12, a second audio input unit 14, a display unit 16, a storage unit 18, an audio output unit 20, an operation unit 22, a central processing unit (CPU) 24, a read-only memory (ROM) 26, and a random access memory (RAM) 28. The imaging unit 10 captures video. The imaging unit 10 includes an imaging optical system 10A, an imaging element 10B, and an image signal processing unit 10C. The imaging optical system 10A forms an image of a subject on the light receiving surface of the imaging element 10B. The imaging element 10B converts the image of the subject formed on the light receiving surface by the imaging optical system 10A into an electrical signal. The image signal processing unit 10C performs predetermined signal processing on the signal output from the imaging element 10B to generate a video signal.
[0027] The first audio input unit 12 is an input unit for the main audio (first audio). The first audio input unit 12 includes a first microphone 12A and a first audio signal processing unit 12B. The first microphone 12A collects the first audio as the main audio. This first audio is audio that does not include the second audio (including cases where the second audio is slightly included). The first microphone 12A is an example of a first audio collection unit. The first audio signal processing unit 12B performs predetermined signal processing on the signal from the first microphone 12A to generate an audio signal of the first audio.
[0028] The second audio input unit 14 is an input unit for a specific audio (second audio) to be synthesized with the main audio. The second audio input unit 14 includes a second microphone 14A and a second audio signal processing unit 14B. The second microphone 14A collects the second audio, which is a specific audio. This second audio is audio that does not include the first audio (including cases where it is deemed not to substantially include the first audio). The second microphone 14A is an example of a second audio collection unit. The second audio signal processing unit 14B performs predetermined signal processing on the signal from the second microphone 14A to generate an audio signal of the second audio.
[0029] The display unit 16 displays in real time the video being captured by the imaging unit 10. The display unit 16 also displays the played back video. The display unit 16 also displays an operation screen, a menu screen, messages, etc. as necessary. The display unit 16 is configured to include, for example, a display device such as an LCD (Liquid Crystal Display), a drive circuit for the display device, etc.
[0030] The storage unit 18 mainly stores captured images and collected sounds. The storage unit 18 includes, for example, a storage medium such as a nonvolatile memory, and a control circuit therefor.
[0031] The audio output unit 20 outputs the reproduced audio. The audio output unit 20 also outputs warning sounds and the like as necessary. The audio output unit 20 includes a speaker, a signal processing circuit that processes the audio signal to be output from the speaker, and the like.
[0032] The operation unit 22 receives operation inputs from the user and includes various operation buttons such as a record button, and a detection circuit for detecting the operations.
[0033] The CPU 24 executes a predetermined control program to function as a control unit for the entire device. The CPU 24 controls the operation of each component based on user operations and performs overall control of the operation of the entire device. The CPU 24 also executes a predetermined program to function as a video generation device that generates video with audio using recorded video and audio. The CPU 24, functioning as a video generation device, processes recorded video and audio based on user operations to generate video with audio. The ROM 26 stores various programs executed by the CPU 24, as well as data necessary for control. The RAM 28 provides the CPU 24 with working memory space.
[0034] 2 is a block diagram of the main functions realized by the CPU when recording video and audio. As shown in the figure, the CPU 24 functions as an imaging control unit 101, a video output unit 102, a first video recording unit 103, a first audio recording unit 104, and a second audio recording unit 105.
[0035] The imaging control unit 101 controls imaging by the imaging unit 10. The imaging control unit 101 controls the imaging unit 10 so that an image is captured with proper exposure, based on a video signal obtained from the imaging unit 10. The imaging control unit 101 also controls the imaging unit 10 so that a main subject is in focus, based on the video signal obtained from the imaging unit 10.
[0036] The video output unit 102 outputs the video captured by the imaging unit 10 to the display unit 16 in real time, whereby a live view is displayed on the display unit 16.
[0037] The first video recording unit 103 records the video (first video) captured by the imaging unit 10 in the storage unit 18. The first video recording unit 103 starts recording the video in response to an instruction from the user. Also, it stops recording the video in response to an instruction from the user. The user issues instructions to start and stop recording via the operation unit 22. The video (first video) is recorded in the storage unit 18 in association with the first audio and second audio collected in synchronization with the capture.
[0038] The first audio recording unit 104 records the first audio (main audio) input from the first audio input unit 12 in the storage unit 18 in synchronization with the capture of the first video. The first audio is recorded in the storage unit 18 in association with the first video.
[0039] The second audio recording unit 105 records the second audio (specific audio) input from the second audio input unit 14 in the storage unit 18 in synchronization with the capture of the first video. The second audio is recorded in the storage unit 18 in association with the first video.
[0040] 3 is a block diagram of the main functions realized by the CPU when playing back recorded video. As shown in the figure, the CPU 24 functions as a video playback unit 111, an audio playback unit 112, and the like.
[0041] In response to a playback instruction from the user, the video playback unit 111 plays back the video recorded in the storage unit 18 on the display unit 16. The user selects the video to be played back and issues a playback instruction using the display unit 16 and the operation unit 22. The video playback unit 111 reads out the selected video from the storage unit 18 and plays it back.
[0042] When audio is associated with a video, the audio playback unit 112 plays the audio in synchronization with the video. When a first audio and a second audio are associated with the video, the audio playback unit 112 synthesizes and plays the first audio and the second audio. The played audio is output from the audio output unit 20.
[0043] 4 is a block diagram of the main functions realized by the CPU when generating video with audio. As shown in the figure, the CPU 24 functions as a first video acquisition unit 121, a first audio acquisition unit 122, a second audio acquisition unit 123, a third audio generation unit 124, an intensity setting unit 125, a video generation unit 126, a second video recording unit 127, etc.
[0044] The first video acquisition unit 121 reads and acquires the video (first video) selected by the user as the processing target from the storage unit 18. The user selects the video to be processed using the display unit 16 and the operation unit 22. The acquired video data is added to the video generation unit 126.
[0045] The first audio acquisition unit 122 reads and acquires data of the first audio (main audio) associated with the video selected as the processing target from the storage unit 18. The acquired first audio data is added to the third audio generation unit 124.
[0046] The second audio acquisition unit 123 reads and acquires the second audio (specific audio) associated with the video selected as the processing target from the storage unit 18. The acquired second audio data is added to the third audio generation unit 124. Note that the first audio acquisition unit 122 and the second audio acquisition unit 123 may acquire the corresponding audio data directly from the first audio input unit 12 and the second audio input unit 14 without going through the storage unit 18. Furthermore, the imaging device 1 may record audio data in an external storage unit instead of the internal storage unit 18 of the device. In this case, the first audio acquisition unit 122 and the second audio acquisition unit 123 may acquire the audio data from the external storage unit.
[0047] The third audio generation unit 124 processes the first audio using the second audio to generate the third audio. The third audio is generated as audio in which the second audio is included in the first audio at a predetermined intensity (audio level). The predetermined intensity is set by the user. FIG. 5 is a block diagram of the functions of the third audio generation unit. As shown in the figure, the third audio generation unit 124 has the functions of an intensity adjustment unit 124A and a synthesis unit 124B. The intensity adjustment unit 124A adjusts the intensity of the second audio according to the setting of the intensity setting unit 125. The synthesis unit 124B synthesizes the second audio, after the intensity adjustment, with the first audio to generate the third audio. This generates audio (third audio) in which the second audio is included in the first audio at a predetermined intensity. Note that, as described above, the first audio is recorded in synchronization with the video, so the generated third audio is also synchronized with the video. The generated third audio data is supplied to the video generation unit 126.
[0048] The intensity setting unit 125 sets the intensity (audio level) of the second audio when it is synthesized with the first audio. The intensity setting unit 125 sets the intensity based on an operation input from the operation unit 22. By setting the intensity of the second audio via the operation unit 22, the user can emphasize or attenuate the second audio relative to the first audio.
[0049] The video generation unit 126 associates the video (first video) acquired by the first video acquisition unit 121 with the third audio generated by the third audio generation unit 124 to generate a video with audio (second video). For example, the video generation unit 126 containerizes the video file and the audio file to generate a video file in a predetermined video format. For example, the generated file may be AVI (Audio Video Interleave), MP4 (MPEG-4 Part 14 (ISO / IEC 14496-14:2003, ISO / IEC JTC 1)), or the like.
[0050] The second video recording unit 127 stores the video with audio (second video) generated by the video generating unit 126 in the storage unit 18.
[0051] Next, a procedure (video generation method) for generating video with audio using the imaging device 1 configured as described above will be described.
[0052] First, an image is captured, and the image, first audio, and second audio are recorded. Specifically, the image captured by the imaging unit 10 (first image) is recorded in the storage unit 18 (image recording step). Synchronously with the image capture, the first audio and second audio are collected and recorded in the storage unit 18 (first audio recording step and second audio recording step). Here, the first audio is recorded as a main audio. On the other hand, the second audio is recorded as a specific audio. The "specific audio" here refers to audio that is different from the main audio and is to be included in the main audio. For example, when capturing an image of a person talking in a windy environment, the person's voice can be recorded as the main audio, and the wind noise (the sound caused by wind hitting a microphone) can be recorded as the specific audio. Alternatively, when capturing an image of a person talking on the beach, the person's voice can be recorded as the main audio, and the sound of waves can be recorded as the specific audio.
[0053] Next, the intensity (audio level) of the second audio when the second audio is synthesized with the first audio is set (intensity setting step). The user sets the intensity via the operation unit 22. This setting allows the user to arbitrarily emphasize or attenuate the second audio relative to the first audio.
[0054] Next, the first sound is processed using the second sound to generate a third sound in which the second sound is included in the first sound at a predetermined intensity (sound generation process). The predetermined intensity is the intensity set in the intensity setting process. In this process, the second sound is first adjusted to the intensity set by the user. As a result, a second sound that is emphasized or reduced relative to the first sound is generated. Then, the second sound after the intensity adjustment is synthesized with the first sound to generate a third sound. As a result, a sound (third sound) in which the second sound is included in the first sound at a predetermined intensity is generated. For example, if wind noise is recorded as the second sound, a sound (third sound) in which the wind noise is included in the main sound at a predetermined intensity is generated. Furthermore, for example, if the sound of waves is recorded as the second sound, a sound (third sound) in which the sound of waves is included in the main sound at a predetermined intensity is generated.
[0055] Next, the sound (third sound) generated in the sound generating step is associated with the image (first image) obtained by imaging, and a video with sound (second image) is generated (image generating step).
[0056] Through the above series of steps, a video with audio is generated. The generated video with audio is recorded in the storage unit 18.
[0057] According to the video generation method of this embodiment, by recording a specific audio (second audio) separately from the main audio (first audio), the specific audio can be separated and edited, thereby generating a video with audio according to the user's intention.
[0058] [Modification of the first embodiment] (1) Modifications for synthesis of second voice The second audio can be configured to be synthesized only in a specific section of the video (a section on the time axis). In this case, the section to be synthesized is specified, and the second audio is synthesized with the first audio. The section to be synthesized is specified, for example, while the first video and the first audio are being played back.
[0059] Additionally, the intensity of the secondary audio can be adjusted along the time axis during synthesis, allowing you to create video with audio in which the intensity of a specific audio varies depending on the scene, for example.
[0060] (2) Modification of intensity setting In the above embodiment, the user can arbitrarily set the intensity of the second voice when synthesizing it with the first voice, but it is also possible to configure it so that it is synthesized at an intensity selected from a plurality of predetermined intensity settings (for example, strong reduction, weak reduction, strong emphasis, weak emphasis).
[0061] (3) Modifications to the recording of the second audio The second audio does not necessarily have to be recorded in synchronization with the capture of the first video. For example, when the second audio is to be synthesized only in a specific section as described above, the second audio may be recorded before or after the fact.
[0062] (4) Modifications of the first voice input unit and the second voice input unit It is preferable that the first audio input unit 12 and the second audio input unit 14 perform filtering processing on the audio signal of the collected audio as needed. For example, it is preferable that the first audio input unit 12 performs filtering processing so that the main audio is recorded clearly. Similarly, it is preferable that the second audio input unit 14 performs filtering processing so that a specific audio is recorded clearly.
[0063] It is also preferable to use microphones suited to the purpose for the first audio input unit 12 and the second audio input unit 14. For example, when collecting a wide-area sound as the main sound, a microphone having a less omnidirectional (preferably omnidirectional) sound collection characteristic than the second microphone 14A is used as the first microphone 12A. Furthermore, a microphone having a directional sound collection characteristic (for example, a gun microphone) is used as the second microphone 14A, which collects a specific sound. This allows the first sound and the second sound to be recorded with high accuracy.
[0064] The first microphone 12A and the second microphone 14A may be built into the main body of the imaging device 1, or may be attached externally.
[0065] [Second embodiment] As in the first embodiment, an example in which an image is generated using an imaging device will be described.
[0066] In this embodiment, audio including a second audio is recorded as a first audio. Then, the second audio recorded separately from the first audio is used to perform processing on the first audio to emphasize or reduce the second audio, thereby generating a third audio. The basic configuration of the imaging device is the same as that of the first embodiment, but the functions realized by the CPU 24 are different.
[0067] 6 is a block diagram of the main functions realized by the CPU when generating video with audio. As shown in the figure, the CPU 24 functions as a first video acquisition unit 121, a first audio acquisition unit 122, a second audio acquisition unit 123, a third audio generation unit 124, a processing condition setting unit 128, a video generation unit 126, and a second video recording unit 127. The functions of each unit except for the third audio generation unit 124 and the processing condition setting unit 128 are substantially the same as those in the first embodiment. Therefore, only the functions of the third audio generation unit 124 and the processing condition setting unit 128 will be described here.
[0068] FIG. 7 is a block diagram of functions of the third sound generating unit.
[0069] As described above, in this embodiment, a sound including the second sound is recorded as the first sound. The third sound generation unit 124 generates the third sound by performing processing on the first sound to emphasize or reduce sound components (common components) that are common to the second sound. Specifically, sound components with the same frequency as the second sound are set as common components, and the sound components with the same frequency as the second sound are processed under processing conditions set by the user to generate the third sound. For this reason, the third sound generation unit 124 has the functions of a frequency detection unit 124C and a sound processing unit 124D.
[0070] The frequency detection unit 124C analyzes the data of the second sound to detect the frequency of the second sound. The second sound is a specific sound in the first sound, and is a sound that the user wishes to emphasize or reduce. As in the first embodiment, the second sound is collected by the second microphone 14A. Information detected by the frequency detection unit 124C is sent to the sound processing unit 124D.
[0071] The audio processing unit 124D acquires information about the frequency of the second audio detected by the frequency detection unit 124C, and generates a third audio by processing the first audio under the processing conditions set by the processing condition setting unit 128. That is, the audio components (common components) having the same frequency as the second audio are processed under the processing conditions set by the user to generate the third audio.
[0072] The processing condition setting unit 128 sets processing conditions for processing the first voice. Specifically, it sets processing conditions (sound emphasis or reduction processing) for common components, which are voice components that are common to the second voice. The processing condition setting unit 128 sets the processing conditions based on operation input from the operation unit 22. By setting processing conditions for processing the first voice via the operation unit 22, the user can emphasize, reduce, or cancel the second voice included in the first voice.
[0073] Next, a procedure (video generation method) for generating video with audio using the imaging device configured as described above will be described.
[0074] First, an image is captured, and the image, first audio, and second audio are recorded. Specifically, the image captured by the imaging unit 10 (first image) is recorded in the storage unit 18 (image recording process). Synchronously with the image capture, the first audio and second audio are collected and recorded in the storage unit 18 (first audio recording process and second audio recording process). As described above, audio including the second audio is recorded as the first audio. That is, audio including a common audio component that is common to the second audio is recorded. Meanwhile, a specific audio component in the first audio is recorded as the second audio. Here, the "specific audio component" refers to audio included in the first audio that the user wishes to emphasize or reduce. For example, when capturing an image of a person talking in a windy environment, wind noise can be recorded as the second audio. Or, when capturing an image of a person talking on the beach, the sound of waves can be recorded as the second audio.
[0075] Next, processing conditions for the common components of the first voice when processing the first voice are set (processing condition setting step). The user sets the processing conditions via the operation unit 22. By setting these processing conditions, the user can arbitrarily emphasize, reduce, or cancel the second voice contained in the first voice.
[0076] Next, the first sound is processed using the second sound to emphasize or reduce the common components, thereby generating a third sound (sound generation process). In this process, the frequency of the second sound is first detected. Then, the first sound is processed to emphasize or reduce the sound components of that frequency according to the processing conditions set in the processing condition setting process, thereby generating a third sound. This generates a sound (third sound) that includes the second sound contained in the first sound at the intensity intended by the user, or that has been canceled. For example, if the first sound is recorded in a windy environment and wind noise is recorded as the second sound, it is possible to generate a sound in which the wind noise has been reduced or canceled. It is also possible to generate a sound in which the wind noise has been emphasized as necessary.
[0077] Next, the sound (third sound) generated in the sound generating step is associated with the image (first image) obtained by imaging, and a video with sound (second image) is generated (image generating step).
[0078] Through the above series of steps, a video with audio is generated. The generated video with audio is recorded in the storage unit 18.
[0079] According to the video generation method of this embodiment, by recording a specific audio (second audio) included in the main audio separately from the main audio (first audio), the specific audio can be separated and edited, thereby generating a video with audio according to the user's intention.
[0080] [Modification of the second embodiment] (1) Modifications for generating the third voice The process of enhancing, reducing, or canceling the common component can be performed only in a specific section of the video. In this case, the process is performed by specifying the section.
[0081] Furthermore, the processing conditions for the common components can be partially varied along the time axis, which makes it possible to generate video with audio in which the intensity of a specific audio is changed depending on the scene, for example.
[0082] (2) Modifications to the recording of the second audio The second audio does not necessarily have to be recorded in synchronization with the capture of the first video. The second audio may be recorded in advance or afterward. For example, the audio of the environment to be edited (wind noise, waterfall noise, construction noise, etc.) may be recorded in advance as the second audio. Alternatively, the audio of the environment to be edited may be recorded in advance as sample audio, and the second audio may be created from the common components of the recorded audio and the sample audio during the video recording process.
[0083] Furthermore, representative second audio can be stored in advance as preset data in the imaging device. This allows, for example, when audio including audio stored as preset data is recorded together with video, audio can be generated with the audio emphasized or reduced. For example, if wind noise is stored as preset data, video with audio in which the wind noise of the first audio is emphasized or reduced can be generated using the wind noise frequency data contained in the preset data. The preset data is stored, for example, in ROM 26 or storage unit 18. When generating video, the user selects the audio data to be edited.
[0084] (3) Modifications of the first voice input unit and the second voice input unit As in the first embodiment, the first voice input unit 12 and the second voice input unit 14 preferably perform filtering processing on the voice signal of the collected voice as needed. Also, the first voice input unit 12 and the second voice input unit 14 preferably use microphones according to the purpose.
[0085] Except for the case where the second audio is recorded in synchronization with the image capture, the second audio can also be collected by the first microphone 12A. That is, when the second audio is recorded in advance or afterward, the second audio can be collected and recorded using the first audio input unit 12. Therefore, when the second audio is recorded in advance or afterward, the second audio input unit 14 is not required in the device main body.
[0086] [Third embodiment] In this embodiment, the movement of the imaging device body is detected while recording video and audio, and the third audio is generated taking into account information about that movement.
[0087] 8 is a block diagram showing a schematic configuration of an imaging device according to this embodiment. As shown in the figure, the imaging device 1 according to this embodiment differs from the imaging devices according to the first and second embodiments in that it further includes a motion detection unit 30.
[0088] The motion detection unit 30 detects the motion of the imaging device body including the imaging unit 10. The motion detection unit 30 detects the motion of the imaging device body in synchronization with the imaging by the imaging unit 10. That is, it starts detecting motion simultaneously with the start of imaging and ends detection simultaneously with the end of imaging (motion detection process). The motion detection unit 30 is configured, for example, with an acceleration sensor or the like. Note that if the imaging device body has an image shake correction function or the like, the sensor used for shake detection or the like can be used as the sensor for motion detection.
[0089] 9 is a block diagram of the main functions realized by the CPU when recording video and audio. As shown in the figure, the CPU 24 functions as an imaging control unit 101, a video output unit 102, a first video recording unit 103, a first audio recording unit 104, a second audio recording unit 105, a motion recording unit 106, etc. The functions of each unit except for the motion recording unit 106 are substantially the same as those in the first embodiment.
[0090] The motion recording unit 106 records information about the motion of the imaging device body detected by the motion detection unit 30 in the storage unit 18 in synchronization with the imaging of the first video. The motion information is associated with the first video and recorded in the storage unit 18. When generating a video with audio (second video), the motion information stored in the storage unit 18 is used to generate audio (third audio) to be associated with the video.
[0091] The third sound generation unit 124 generates the third sound by taking into account the movement information. Here, an example will be described in which the third sound is generated by synthesizing the second sound with the first sound. As described in the first embodiment above, the first sound is a sound that does not substantially include the second sound, and the second sound is a sound that does not substantially include the first sound.
[0092] 10 is a block diagram of the functions of the third sound generation unit of this embodiment. As shown in the figure, the third sound generation unit 124 of this embodiment has the functions of a second sound processing unit 124E, an intensity adjustment unit 124A, and a synthesis unit 124B. The functions of each unit except for the second sound processing unit 124E are substantially the same as those of the first embodiment.
[0093] The second audio processor 124E acquires information about the movement of the imaging device body when recording video and audio, and processes the second audio based on the movement information. Specifically, the second audio is processed according to predetermined processing conditions in response to the movement of the imaging device body. As an example, a case will be described in which audio (third audio) is generated that includes wind noise (second audio) in addition to the main audio (first audio). Assume that the second microphone 14A is composed of a pair of left and right microphones and is integrally mounted on the imaging device body. Therefore, in this case, the second microphone 14A moves integrally with the imaging device body. When panning of the imaging device body is detected, the second audio processor 124E performs processing to change the intensity of the left and right audio in response to the movement of the imaging device body. Specifically, the audio of the moving microphone is weakened. This allows appropriate processing of wind noise that changes between the left and right microphones in response to the movement of the imaging device body. In other words, since wind noise is stronger in the moving microphone, weakening the wind noise in response to the movement allows for balanced left and right audio to be synthesized.
[0094] The intensity adjustment unit 124A adjusts the intensity of the processed second voice in accordance with the setting of the intensity setting unit 125. The synthesis unit 124B synthesizes the second voice, whose intensity has been adjusted, with the first voice to generate a third voice.
[0095] In this way, in this embodiment, the second sound is automatically processed in accordance with the movement of the imaging device body to generate the third sound, thereby automatically eliminating the influence of the movement of the imaging device body.
[0096] [Modification of the third embodiment] In the following, we will explain the case where a second audio is used to perform processing on a first audio to emphasize or reduce common components, thereby generating a third audio, and where information about the movement of the imaging device body is added to generate the third audio.
[0097] 11 is a block diagram of the functions of the third sound generation unit 124 of this example. As shown in the figure, the third sound generation unit 124 of this example differs from the third sound generation unit 124 of the second embodiment in that a sound processing unit 124D processes the first sound based on information about the movement of the imaging device body.
[0098] The audio processor 124D processes the first audio in accordance with predetermined processing conditions in response to the movement of the imaging device body. As an example, a case will be described in which an audio (first audio) including wind noise (second audio) in the main audio is recorded in synchronization with imaging, and an audio (third audio) is generated in which the wind noise is emphasized or reduced. In this case, the first audio includes the main audio as well as the wind noise (second audio). The second audio includes the wind noise.
[0099] The frequency detection unit 124C analyzes the data of the second sound and detects the frequency of the second sound.
[0100] The user then selects one intensity setting from multiple predetermined audio intensity settings. The audio processing unit 124D acquires information on the frequency of the second audio detected by the frequency detection unit 124C and information on the movement of the image capture device body, and processes the first audio by combining the processing conditions set by the processing condition setting unit 128 with the movement information to generate the third audio. For example, consider a case where video and audio are recorded while moving in an environment where wind noise (second audio) is recorded along with the main audio (first audio). The intensity of the wind noise (second audio) changes depending on the moving speed. Therefore, the audio processing unit 124D corrects the preset intensity setting or changes the frequency to be processed (the frequency of the common component) depending on the moving speed (movement) of the image capture device body. This allows the target audio (second audio) to be appropriately processed even when video and audio are recorded while moving. For example, if the user's setting slightly reduces wind noise (second audio), a scene in which the moving speed of the image capture device body is determined to be fast is corrected to significantly reduce wind noise compared to other scenes, and the third audio is generated. This prevents certain sounds (wind, waves, etc.) from becoming too loud in only certain scenes of the third audio.
[0101] [Fourth embodiment] In this embodiment, when a video (first video) is captured, imaging information of the video is acquired and recorded in storage unit 18. The recorded imaging information is used when generating a video with audio (second video). Specifically, when generating a video with audio, the imaging information is displayed on display unit 16. The user uses the information displayed on display unit 16 to identify the section (scene) for which audio editing is to be performed. Here, imaging information is information related to capturing the video. For example, it includes information on the focal length when the video was captured, information on the subject distance, information on exposure, etc. Furthermore, if the imaging device has a function for detecting the movement of the imaging device itself, information on the movement of the imaging device itself during imaging is also included.
[0102] The CPU 24 performs processes such as acquiring, recording, and displaying imaging information. Fig. 12 is a block diagram of functions realized by the CPU when acquiring and recording imaging information and when displaying imaging information. As shown in the figure, the CPU 24 functions as an imaging information acquisition unit 131, an imaging information recording unit 132, and an imaging information display unit 133.
[0103] The imaging information acquisition unit 131 acquires imaging information of the first video in synchronization with the imaging of the first video by the imaging unit 10. The imaging information acquired includes information on subject distance, focal length, and information on the movement of the imaging device (for example, the output of an acceleration sensor).
[0104] The imaging information recording unit 132 records the imaging information acquired by the imaging information acquisition unit 131 in the storage unit 18. The imaging information is recorded in association with the video (first video).
[0105] The imaging information display unit 133 acquires imaging information from the storage unit 18 based on an operation input from the operation unit 22, and displays it on the display unit 16. For example, the imaging information is displayed in chronological order.
[0106] The generation of a video with syllables using the imaging device of this embodiment is performed, for example, as follows.
[0107] First, a first video is captured by the imaging unit 10 and recorded in the storage unit 18 (video recording step). In synchronization with the capturing, imaging information of the first video is acquired (first information acquisition step), and is associated with the first video and recorded in the storage unit 18. Also, in synchronization with the capturing, a first sound and a second sound are collected, and are associated with the first video and recorded in the storage unit 18 (first sound recording step and second sound recording step).
[0108] Next, a video with audio (second video) is generated using the recorded first video, first audio, and second audio. First, imaging information of the first video is read from the storage unit 18 and displayed on the display unit 16 (first display step). The imaging information is displayed in chronological order of the first video. Displaying this imaging information makes it easier to identify the portions where audio editing is required. The user specifies, via the operation unit 22, the portions where the second audio is to be synthesized or the portions where the second audio is to be edited (portions where the second audio is to be emphasized, reduced, or canceled), and instructs the generation of a video with audio (second video).
[0109] When generation of a video with audio is instructed, the first audio is processed using the second audio to generate a third audio (audio generation process) based on an operation input from the operation unit 22. Then, the generated third audio is associated with the first video to generate a video with audio (second video) (video generation process).
[0110] In this way, according to the present embodiment, the imaging information of the first video is acquired and recorded, which makes it possible to easily identify the portions where audio editing is required when generating video with audio.
[0111] [Modification of the Fourth Embodiment] The uses of the acquired imaging information are not limited to the above examples. For example, the imaging information can be used to automatically identify locations requiring audio editing, automatically process the first audio, and generate the third audio. For example, locations where the imaging device itself moves significantly (locations where the movement is above a threshold) or locations where the subject moves significantly (locations where the change in subject distance is above a threshold) can be identified based on the imaging information and automatically processed. Also, for example, the imaging information can be used to automatically change the audio editing method. When reducing the second audio throughout the entire video, the reduction intensity can be automatically changed based on the imaging information (e.g., changing the intensity to locations where the imaging device itself moves significantly or locations where the subject moves significantly). Also, for example, the imaging information can be used to detect locations requiring audio editing and automatically cue the video.
[0112] [Fifth embodiment] In this embodiment, information from the first microphone 12A that collects the first sound and the second microphone 14A that collects the second sound is acquired and recorded in the storage unit 18. The recorded information from the first microphone 12A and the second microphone 14A is used when generating a video with sound (second video). Specifically, when generating a video with sound, the information is displayed on the display unit 16. The user uses the information displayed on the display unit 16 to set the intensity of the second sound, etc. Here, the microphone information includes, for example, information on whether or not there is a windshield, information on the type of windshield (sponge type, fur type, cage type, etc.) if there is a windshield, information on whether or not there is directionality, etc. In addition, microphone information can also include information on various performance elements of the microphone (for example, information on the microphone's directional characteristics (information on how sensitivity changes depending on the direction from which sound arrives), information on frequency characteristics (information on how sensitivity changes depending on the pitch of the sound), information on maximum sound pressure level (the loudest sound level that the microphone can pick up), information on equivalent noise level (input conversion noise level), information on output impedance, information on open circuit sensitivity, etc.).
[0113] The user inputs information from the first microphone 12A and the second microphone 14A to the imaging device 1 via the operation unit 22. The CPU 24 stores the information from the first microphone 12A and the second microphone 14A input via the operation unit 22 in the memory unit 18. When generating a video with sound, the information from the first microphone 12A and the second microphone 14A stored in the memory unit 18 is displayed on the display unit 16.
[0114] 13 is a block diagram of the functions realized by the CPU when acquiring and recording microphone information and when displaying the microphone information. As shown in the figure, the CPU 24 functions as a microphone information acquisition unit 141, a microphone information recording unit 142, and a microphone information display unit 143.
[0115] The microphone information acquisition unit 141 acquires information about the first microphone 12A and the second microphone 14A. As described above, the information about the first microphone 12A and the second microphone 14A is input by the user via the operation unit 22. The user inputs information about the first microphone 12A and the second microphone 14A, such as whether or not they have a windshield, and if so, the type of windshield (sponge type, fur type, cage type, etc.), whether or not they have directionality, etc.
[0116] The microphone information recording unit 142 records the information of the first microphone 12A and the second microphone 14A acquired by the microphone information acquisition unit 141 in the storage unit 18. The information of the first microphone 12A and the second microphone 14A is recorded in association with the video (first video).
[0117] The microphone information display unit 143 acquires information about the first microphone 12A and the second microphone 14A from the storage unit 18 based on an operation input from the operation unit 22, and displays the information on the display unit 16.
[0118] The generation of video with audio using the imaging device of this embodiment is carried out, for example, as follows.
[0119] First, a first video is captured by the imaging unit 10 and recorded in the storage unit 18 (video recording step). Furthermore, a first sound and a second sound are collected in synchronization with the capturing of the video and are recorded in the storage unit 18 in association with the first video (first sound recording step and second sound recording step). Furthermore, information on the first microphone 12A and the second microphone 14A used during the capturing of the video is input (second information acquisition step) and recorded in the storage unit 18.
[0120] Next, a video with audio (second video) is generated using the recorded first video, first audio, and second audio. First, information from the first microphone 12A and the second microphone 14A is read from the storage unit 18 and displayed on the display unit 16 (second display step). The user sets the intensity when synthesizing the second audio based on the information from the first microphone 12A and the second microphone 14A. The user also sets processing conditions when processing the first audio based on the information from the first microphone 12A and the second microphone 14A. For example, the intensity is set depending on the presence or absence of a windshield and its type.
[0121] When generation of a video with audio is instructed, the first audio is processed using the second audio to generate a third audio (audio generation process) based on an operation input from the operation unit 22. Then, the generated third audio is associated with the first video to generate a video with audio (second video) (video generation process).
[0122] As described above, according to the present embodiment, information from the first microphone 12A and the second microphone 14A is acquired and recorded. This allows the third audio to be generated more appropriately when generating video with audio. For example, loss of the main audio can be minimized.
[0123] [Modification of the fifth embodiment] The use of the acquired information from the first microphone 12A and the second microphone 14A is not limited to the above example. For example, a configuration may be adopted in which the first audio is automatically processed to generate a third audio using the information from the first microphone 12A and the second microphone 14A. For example, when generating a third audio by synthesizing a second audio with the first audio, the intensity of the second audio when synthesizing the second audio may be automatically set depending on the type of the second microphone 14A. Furthermore, when generating a third audio by processing audio components of the first audio with the same frequency as the second audio, the frequency to be processed may be automatically changed depending on the type of the second microphone 14A.
[0124] Furthermore, in the above example, the user is configured to input information from the first microphone 12A and the second microphone 14A into the imaging device 1, but it is also possible to configure the imaging device 1 to automatically collect information from the first microphone 12A and the second microphone 14A.
[0125] [Sixth embodiment] In this embodiment, the timing at which the second audio is recorded is detected and recorded while the first video is being captured. The recorded information is used when generating the video with audio (second video). Specifically, when generating the video with audio, the information is displayed on the display unit 16. The user uses the information displayed on the display unit 16 to specify the section (scene) in which the audio is to be edited.
[0126] The timing at which the second voice is recorded is detected by the CPU 24. Fig. 14 is a block diagram of functions realized by the CPU when detecting and recording the timing at which the second voice is recorded and when displaying the recorded information. As shown in the figure, the CPU 24 functions as a second voice detection unit 151, a timing information recording unit 152, and a timing information display unit 153.
[0127] The second audio detection unit 151 detects the timing at which the second audio was recorded based on the audio signal of the second audio input from the second audio input unit 14. That is, the second audio detection unit 151 detects the input of the audio signal of the second audio and detects the timing at which the second audio was recorded.
[0128] The timing information recording unit 152 records information (timing information) about the recording timing of the second audio detected by the second audio detection unit 151 in the storage unit 18. The timing information is recorded in association with the video (first video).
[0129] The timing information display unit 153 acquires timing information from the storage unit 18 based on an operation input from the operation unit 22, and displays it on the display unit 16. For example, it displays the timing at which the second audio was recorded on the time axis.
[0130] The generation of a video with syllables using the imaging device of this embodiment is performed, for example, as follows.
[0131] First, a first video is captured by the imaging unit 10 and recorded in the storage unit 18 (video recording process). Timing information of the first video is acquired in synchronization with the capturing (first information acquisition process), and is recorded in association with the first video in the storage unit 18. Also, a first sound and a second sound are collected in synchronization with the capturing, and are recorded in association with the first video in the storage unit 18 (first sound recording process and second sound recording process). Also, the timing at which the second sound was recorded is detected (second sound detection process). The detected information is recorded in association with the first video in the storage unit 18 (association process).
[0132] Next, a video with audio (second video) is generated using the recorded first video, first audio, and second audio. First, timing information is read from storage unit 18 and displayed on display unit 16. For example, the timing at which the second audio was recorded is displayed on a time axis. The user specifies, via operation unit 22, the point at which the second audio is to be synthesized or the point at which the second audio is to be edited, and instructs generation of a video with audio (second video).
[0133] When generation of a video with audio is instructed, the first audio is processed using the second audio to generate a third audio (audio generation process) based on an operation input from the operation unit 22. Then, the generated third audio is associated with the first video to generate a video with audio (second video) (video generation process).
[0134] As described above, according to the present embodiment, the timing at which the second audio is recorded is detected and recorded, which makes it possible to easily identify the locations (scenes) where audio editing is required when generating video with audio.
[0135] [Modification of the sixth embodiment] The use of the acquired timing information is not limited to the above example, and for example, it may be configured to automatically identify portions that require audio editing based on the timing information.
[0136] The timing at which the second audio was recorded may be detected after the first video has been captured, i.e., the audio data of the second audio may be analyzed after the first video has been captured to detect the timing at which the second audio was recorded.
[0137] [Other embodiments] As described above, it is preferable to use a microphone with omnidirectional sound collection characteristics for the first microphone 12A that collects the first sound, which is the main sound. It is also preferable to use a microphone with directional sound collection characteristics (such as a gun microphone) for the second microphone 14A that collects the second sound. This improves the accuracy of recording the first sound and the second sound. It is also possible to adjust the sound while maintaining the sound characteristics of a specific voice, for example.
[0138] Furthermore, when collecting the second sound using a microphone with directional sound collection characteristics, if the directivity can be adjusted, it is more preferable to change the microphone's directivity in response to changes in the position of the sound source of the second sound. For example, the orientation of the second microphone 14A is changed in response to changes in the position of the sound source of the second sound within the video. The change in the position of the sound source is detected, for example, by analyzing the video. For example, the subject that is the sound source of the second sound is identified within the video, and the position of the subject is detected using image recognition or the like to identify the position of the sound source.
[0139] In the above embodiment, the present invention has been described with reference to an imaging device, but the devices and systems for implementing the present invention are not limited to this. For example, the present invention can also be implemented in a portable electronic device (e.g., a smartphone, a tablet computer, a laptop computer, etc.) equipped with imaging and audio recording functions. It is also possible to import the recorded first video, first audio, and second audio into a computer (e.g., a personal computer, etc.), generate the third audio, and generate the video with audio (the second video).
[0140] The control unit that executes the function of generating the third audio and the function of generating the second video can be realized using various processors. The various processors include, for example, a CPU, which is a general-purpose processor that executes software (programs) to implement various functions. The various processors also include a GPU (Graphics Processing Unit), which is a processor specialized for image processing, and a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacturing. Furthermore, the various processors also include dedicated electrical circuits, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processing.
[0141] The control unit may be realized by a single processor, or by multiple processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU). Furthermore, multiple functions may be realized by a single processor. Examples of multiple functions configured by a single processor include: a first configuration in which a single processor is configured by combining one or more CPUs and software, as typified by computers such as servers, and this processor realizes multiple functions; a second configuration in which a processor is used to realize the functions of the entire system on a single IC (Integrated Circuit) chip, as typified by system-on-chip (SoC). In this way, various functions are configured using one or more of the above-mentioned various processors as a hardware structure. Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit combining circuit elements such as semiconductor devices. These electrical circuits may realize the above-mentioned functions using logical operations such as logical sum, logical product, logical negation, exclusive OR, and combinations of these.
[0142] When the processor or electrical circuit executes software (programs), processor-readable code for the software to be executed is stored in a non-transitory recording medium such as a ROM, and the processor references the software. Software stored in a non-transitory recording medium includes programs for image input, analysis, display control, and the like. Code may be recorded in a non-transitory recording medium such as various types of magneto-optical recording devices or semiconductor memory, rather than in a ROM. When processing using the software, RAM, for example, is used as a temporary storage area, and data stored in an EEPROM (Electronically Erasable and Programmable Read Only Memory), not shown, may also be referenced. [Explanation of symbols]
[0143] 1. Imaging device 10. Imaging unit 10A Imaging optical system 10B image sensor 10C Image signal processing unit 12 First audio input unit 12A 1st microphone 12B First audio signal processing section 14 Second audio input unit 14A Second microphone 14B Second audio signal processing section 16 Display 18 Memory section 20 Audio output section 22 Control section 24 CPU 26 ROM 28 RAM 30 Motion detection unit 101 Imaging control unit 102 Video output section 103 First Video Recording Unit 104 First Audio Recording Unit 105 Second Audio Recording Unit 106 Motion Recording Unit 111 Video playback unit 112 Audio playback unit 121 First Image Acquisition Unit 122 First audio acquisition unit 123 Second audio acquisition unit 124 Third speech generation unit 124A Strength adjustment section 124B Synthesis Department 124C Frequency Detector 124D Audio Processing Unit 124E Second Audio Processing Unit 125 Intensity setting section 126 Image Generation Unit 127 Second Video Recording Unit 128 Processing condition setting section 131 Imaging information acquisition unit 132 Imaging information recording unit 133 Imaging information display section 141 Microphone information acquisition unit 142 Microphone information recording unit 143 Microphone information display 151 Second voice detection unit 152 Timing information recording unit 153 Timing information display section
Claims
1. a video recording step of recording a first video captured by the imaging unit; a first audio recording step of recording a first audio in association with the first video; a second audio recording step of recording a second audio different from the first audio in association with the first video; a second audio detection step of detecting a timing at which the second audio is recorded in the first video; a step of displaying the timing at which the second audio was recorded on a time axis and displaying information on the timing at which the second audio was recorded in the first video on a display unit; playing the first video and the first audio; receiving, by using the information displayed on the display unit, a designation of a section on a time axis that requires editing by combining the second audio with the first video and the first audio; an intensity setting step of receiving an instruction to set the intensity of the second sound; a sound generating step of synthesizing the second sound, which has been emphasized or reduced based on the setting instruction, for a specified section of the first sound to generate a third sound; an image generating step of generating a second image by associating the first image with the third audio; An image generation method comprising:
2. the first voice includes a common component that is a voice component common to the second voice, the sound generating step performs a process of emphasizing or reducing the common component on the first sound using the second sound to generate the third sound. The image generation method according to claim 1 .
3. a processing condition setting step of setting processing conditions for the common component before the sound generating step, the sound generating step performs processing on the first sound to emphasize or reduce the common component in accordance with the processing conditions set in the processing condition setting step; The image generation method according to claim 2 .
4. a detection step of detecting a movement of the imaging device body including the imaging unit, the sound generating step, when the motion is detected in the detecting step, performs a predetermined process on the first sound or the second sound to generate the third sound. The image generating method according to any one of claims 1 to 3.
5. a first information acquisition step of acquiring imaging information of the first video by the imaging unit; a first display step of displaying the imaging information; The image generating method according to claim 1 , further comprising:
6. The imaging information includes at least one of information on the movement of the imaging device body including the imaging unit and information on the focal length. The image generation method according to claim 5 .
7. a second information acquisition step of acquiring information about a sound collection unit that collects the first sound and the second sound, a second display step of displaying information about the sound collection unit; The image generating method according to claim 1 , further comprising:
8. the first sound recording step includes recording the first sound via a first sound collection unit; the second sound recording step includes recording the second sound via a second sound collection unit different from the first sound collection unit. The image generation method according to any one of claims 1 to 7.
9. the second sound collection unit has directional sound collection characteristics, The first sound collection unit has a sound collection characteristic with lower directivity than the second sound collection unit. The image generation method according to claim 8.
10. the second sound collection unit has directional sound collection characteristics, the second sound recording step includes detecting a position of a sound source of the second sound, directing the second sound collection unit in the direction of the detected sound source, and recording the second sound.
10. The image generating method according to claim 8 or 9.
Citation Information
Patent Citations
digital camera with voice recording
JP2002537691A
Video processing apparatus
JP2008178090A
Video camera device
JP2009124414A
Video camera
JP2010004339A
Imaging apparatus and program
JP2012129854A