Imaging apparatus

The imaging device addresses the challenge of collecting sound from multiple subjects by automatically adjusting sound collection area and directionality using face recognition, ensuring clear audio capture in diverse shooting scenarios.

JP2025126341AActive Publication Date: 2025-08-28PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025110743
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-28
Estimated Expiration
2040-06-23

AI Technical Summary

Technical Problem

Existing imaging devices struggle to easily collect sound from a subject in accordance with a user's intention, particularly when multiple subjects are involved in a conversation.

Method used

An imaging device that includes an imaging unit, an audio acquisition unit, a setting unit, and a control unit, which allows for automatic adjustment of sound collection area based on user instructions and the device's shooting state, using face recognition to determine the sound collection area and directionality.

Benefits of technology

Enables easy and accurate sound collection from subjects in various shooting conditions, emphasizing the voices of multiple subjects in a conversation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126341000001_ABST
    Figure 2025126341000001_ABST
Patent Text Reader

Abstract

To provide an imaging apparatus that takes an image while acquiring voice, the apparatus making it easier to collect voice of an object according to an intention of a user.SOLUTION: The imaging apparatus includes: an imaging unit for imaging an object and generating image data; a voice acquisition unit for acquiring a voice signal showing voice collected during imaging of the imaging unit; a setting unit for setting an own device in an auto-mode as an operation mode for automatically changing the directivity of the voice acquisition unit in response to an instruction of a user; and a control unit for controlling a voice collecting area for collecting voice from an object in the voice signal. The control unit controls the voice collecting area so that the area will include the object, by changing the directivity of the voice acquisition unit in association with the imaging state of the own device when the setting unit sets the own device in the auto-mode.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an imaging device that captures an image while acquiring sound. [Background technology]

[0002] Patent Document 1 discloses a video camera with a face detection function. The video camera in Patent Document 1 changes the microphone's directional angle depending on the zoom ratio and the size of the person's face in the captured image. This allows the video camera to control the microphone's directional angle in relation to the distance between the video camera and the subject's video, thereby achieving control to change the microphone's directional angle so as to more reliably capture the subject's voice while maintaining consistency between the video and audio. In this case, the video camera detects the position and size of the person's (subject's) face, displays the detected face with a frame (face detection frame), and uses information on the size of the face detection frame (face size). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-283706 Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure provides an imaging device that captures an image while acquiring sound, and that can easily collect sound from a subject in accordance with a user's intention. [Means for solving the problem]

[0005] In the present disclosure, an imaging device includes an imaging unit that images a subject and generates image data, an audio acquisition unit that acquires an audio signal indicating audio picked up during imaging by the imaging unit, a setting unit that sets the device to auto mode, an operating mode that automatically changes the directionality of the audio acquisition unit, in response to a user's instruction, and a control unit that controls a sound collection area that collects audio from the subject in the audio signal, and when the setting unit is set to auto mode, the control unit controls the sound collection area to include the subject by changing the directionality of the audio acquisition unit in conjunction with the shooting state of the device. [Effects of the Invention]

[0006] According to the imaging device according to the present disclosure, in an imaging device that captures an image while acquiring sound, it is possible to easily collect the sound of a subject in accordance with the user's intention. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 shows a configuration of a digital camera 100 according to a first embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating the rear view of a digital camera 100. [Figure 3] FIG. 10 is a diagram illustrating a state of the digital camera 100 when taking a selfie; [Figure 4] FIG. 10 is a diagram illustrating the state of the digital camera 100 when taking a vertical shot. [Figure 5] FIG. 1 is a diagram illustrating the configuration of a beam forming section 172 in a digital camera 100. [Figure 6] FIG. 1 is a diagram illustrating an example of a sound collection area in a digital camera 100. [Figure 7] FIG. 10 shows an example of a setting menu display in the digital camera 100. [Figure 8] FIG. 10 is a diagram illustrating an example of a further sound collection area in the digital camera 100. [Figure 9] FIG. 1 is a diagram for explaining an outline of the operation of the digital camera 100 in auto mode. [Figure 10] 1 is a flowchart illustrating an example of the operation of the digital camera 100 in focus mode according to the first embodiment; [Figure 11] FIG. 1 is a diagram for explaining an outline of the operation of the focus mode of the digital camera 100. [Figure 12] 10 is a flowchart illustrating the process of selecting sound collection targets (S3 in FIG. 10) of the digital camera 100 according to the first embodiment. [Figure 13] FIG. 10 is a diagram for explaining the selection process of sound collection targets in the digital camera 100. [Figure 14] 10 is a flowchart illustrating the process of determining the sound collection area in the digital camera 100 (S4 in FIG. 10). [Figure 15] FIG. 10 is a diagram illustrating a process for determining a sound collection area in the digital camera 100. [Figure 16] 10 is a flowchart illustrating sound collection control using face recognition in the digital camera 100 (S5 in FIG. 10). [Figure 17] FIG. 10 is a diagram for explaining management information obtained by the sound collection area determination process. [Figure 18] FIG. 10 is a diagram illustrating a relationship for calculating gain from the horizontal angle of view and the focal distance in the digital camera 100. [Figure 19] A flowchart illustrating sound collection control (S6 in FIG. 10) without using face recognition in the digital camera 100. [Figure 20] 1 is a flowchart illustrating an example of an operation in auto mode of the digital camera 100 according to the first embodiment; [Figure 21] FIG. 10 is a diagram for explaining sound collection control during horizontal and vertical shooting in auto mode according to the first embodiment; [Figure 22] FIG. 10 shows a display example of a digital camera 100 according to a second embodiment. [Figure 23] FIG. 10 is a diagram for explaining manual operations in the digital camera 100 according to the second embodiment. [Figure 24] 10 is a flowchart illustrating an example of the operation of the digital camera 100 of the second embodiment during manual operation. [Figure 25] FIG. 10 is a diagram illustrating an example of the arrangement of a microphone 161A in a digital camera 100A according to a modified example. [Figure 26]FIG. 10 is a diagram illustrating sound collection control during vertical shooting in auto mode according to a modified example. [Figure 27] FIG. 10 is a diagram illustrating an example of sound collection control operation linked to face recognition in the digital camera 100. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments will be described in detail with reference to the drawings as appropriate. However, more detailed description than necessary may be omitted. For example, detailed description of already well-known matters or redundant description of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the inventor(s) provide the accompanying drawings and the following description to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.

[0009] (Embodiment 1) In embodiment 1, as an example of an imaging device according to the present disclosure, a digital camera is described which detects a subject based on image recognition technology, controls the sound collection area according to the size of the detected subject, and controls the sound collection gain to emphasize the sound to be collected.

[0010] [1-1. Structure]

[0011] FIG. 1 is a diagram showing the configuration of a digital camera 100 according to this embodiment. The digital camera 100 of this embodiment includes an image sensor 115, an image processing engine 120, a display monitor 130, and a controller 135. The digital camera 100 also includes a buffer memory 125, a card slot 140, a flash memory 145, an operation unit 150, and a communication module 160. The digital camera 100 also includes a microphone 161, an analog-to-digital (A / D) converter 165 for the microphone, and an audio processing engine 170. The digital camera 100 also includes, for example, an optical system 110 and a lens driver 112. The digital camera 100 also includes, for example, a magnetic sensor 132 and an acceleration sensor 137.

[0012] Fig. 2 illustrates an example of the rear surface of digital camera 100. Fig. 2 illustrates the three axial directions X, Y, and Z of digital camera 100, as well as the direction of gravity G. The X, Y, and Z axes correspond to the horizontal and vertical angle of view of digital camera 100 and the optical axis direction of the lens in optical system 110, respectively. In the example of Fig. 2, the Y axis direction of digital camera 100 is oriented along the direction of gravity G, i.e., horizontally.

[0013] The digital camera 100 of this embodiment can be used by the user to take a selfie, or to take a portrait shot by holding the digital camera 100 in a portrait orientation. Fig. 3 shows an example of the state of the digital camera 100 when taking a selfie. Fig. 4 shows an example of the state of the digital camera 100 when taking a portrait shot.

[0014] Returning to FIG. 1, the optical system 110 includes a focus lens, a zoom lens, an optical image stabilization lens (OIS), an aperture, a shutter, etc. The focus lens is a lens for changing the focus state of the subject image formed on the image sensor 115. The zoom lens is a lens for changing the magnification of the subject image formed by the optical system. The focus lens, etc. are each composed of one or more lenses.

[0015] Lens driver 112 drives the focus lens and the like in optical system 110. Lens driver 112 includes a motor, and moves the focus lens along the optical axis of optical system 110 under the control of controller 135. The configuration for driving the focus lens in lens driver 112 can be implemented by a DC motor, a stepping motor, a servo motor, an ultrasonic motor, or the like.

[0016] The image sensor 115 captures an image of a subject formed via the optical system 110 and generates imaging data. The imaging data constitutes image data that represents an image captured by the image sensor 115. The image sensor 115 generates new frame image data at a predetermined frame rate (e.g., 30 frames per second). The timing of generating imaging data and the operation of the electronic shutter in the image sensor 115 are controlled by the controller 135. The image sensor 115 can be any of a variety of image sensors, such as a CMOS image sensor, a CCD image sensor, or an NMOS image sensor.

[0017] The image sensor 115 performs operations for capturing moving images, still images, through images, etc. Through images are mainly moving images, and are displayed on the display monitor 130 so that the user can determine the composition for capturing a still image, for example. The through image, the moving image, and the still image are each an example of a captured image in this embodiment. The image sensor 115 is an example of an imaging unit in this embodiment.

[0018] The image processing engine 120 performs various processes on the imaging data output from the image sensor 115 to generate image data, and performs various processes on the image data to generate an image to be displayed on the display monitor 130. The various processes include, but are not limited to, white balance correction, gamma correction, YC conversion processing, electronic zoom processing, compression processing, and expansion processing. The image processing engine 120 may be configured with a hardwired electronic circuit, or may be configured with a microcomputer, processor, or the like using a program.

[0019] In this embodiment, the image processing engine 120 includes a face recognition unit 122 that realizes a function for detecting a subject, such as a human face, through image recognition of a captured image. The face recognition unit 122 performs face detection, for example, through rule-based image recognition processing, and outputs detection information. Face detection may be performed using various image recognition algorithms. The detection information includes position information corresponding to the subject detection result. The position information is defined, for example, by the horizontal and vertical positions on the image Im to be processed, and indicates, for example, a rectangular area surrounding a human face as the detected subject (see FIG. 11).

[0020] The display monitor 130 is an example of a display unit that displays various information. For example, the display monitor 130 displays an image (through image) represented by image data captured by the image sensor 115 and processed by the image processing engine 120. The display monitor 130 also displays a menu screen or the like that allows the user to make various settings for the digital camera 100. The display monitor 130 can be configured, for example, by a liquid crystal display device or an organic EL device.

[0021] The digital camera 100 of this embodiment is configured with a movable display monitor 130 that can change its position, as shown in, for example, Figures 2 and 3. In the example of Figure 2, the display monitor 130 is positioned with its display surface facing the rear side (-Z side) of the digital camera 100. This position of the display monitor 130 is hereinafter referred to as the "normal position." In the example of Figure 3, the display monitor 130 is positioned with its display surface facing the front side (+Z side) of the digital camera 100, i.e., the subject side. This position of the display monitor 130 is hereinafter referred to as the "selfie position."

[0022] The magnetic sensor 132 is an example of a detection unit that detects whether the display monitor 130 is in the normal position or the selfie position. The magnetic sensor 132 outputs a detection signal indicating the detection result of the position of the display monitor 130 to the controller 135, for example.

[0023] For example, a vari-angle type or a tilt type can be used as the movable display monitor 130. For example, a hinge 131 is provided to rotatably connect the display monitor 130 to the body of the digital camera 100. The magnetic sensor 132 is provided, for example, inside the hinge 131, and is configured by a switch or the like having two states corresponding to FIGS.

[0024] Acceleration sensor 137 detects acceleration in one or more of three axial directions, for example, X, Y, and Z, and outputs a detection signal to controller 135. Acceleration sensor 137 is an example of an orientation detection unit that detects whether the orientation of digital camera 100 is landscape as shown in FIG. 2 or portrait as shown in FIG. 4 based on the detected state of gravitational acceleration.

[0025] Operation unit 150 is a general term for hard keys such as operation buttons and operation levers provided on the exterior of digital camera 100, and accepts operations by the user. Operation unit 150 includes, for example, a release button, a mode dial, a touch panel, cursor buttons, and a joystick. When operation unit 150 accepts an operation by the user, it transmits an operation signal corresponding to the user operation to controller 135. For example, as shown in FIG. 2, operation unit 150 includes a release button 151, a selection button 152, an enter button 153, function buttons 154, a touch panel 155, and the like.

[0026] The controller 135 controls the overall operation of the digital camera 100. The controller 135 includes a CPU and other components, and the CPU executes programs (software) to implement predetermined functions. Instead of a CPU, the controller 135 may include a processor configured with dedicated electronic circuits designed to implement predetermined functions. That is, the controller 135 can be implemented with various processors, such as a CPU, MPU, GPU, DSU, FPGA, or ASIC. The controller 135 may be configured with one or more processors. Furthermore, the controller 135 may be configured on a single semiconductor chip together with the image processing engine 120 and other components.

[0027] The buffer memory 125 is a recording medium that functions as a work memory for the image processing engine 120 and the controller 135. The buffer memory 125 is realized by a DRAM (Dynamic Random Access Memory) or the like. The flash memory 145 is a non-volatile recording medium. Although not shown, the controller 135 may have various types of internal memory, for example, a built-in ROM. The ROM stores various programs executed by the controller 135. The controller 135 may also have a built-in RAM that functions as a work area for the CPU.

[0028] The card slot 140 is a means for inserting a removable memory card 142. The card slot 140 can electrically and mechanically connect the memory card 142. The memory card 142 is an external memory equipped with a recording element such as a flash memory inside. The memory card 142 can store data such as image data generated by the image processing engine 120.

[0029] The communication module 160 is a communication module (circuit) that performs communication in accordance with the communication standard IEEE802.11 or the Wi-Fi standard. The digital camera 100 can communicate with other devices via the communication module 160. The digital camera 100 may communicate with other devices directly via the communication module 160, or may communicate via an access point. The communication module 160 may be connectable to a communication network such as the Internet.

[0030] The microphone 161 is an example of a sound collection unit that collects sound. The microphone 161 converts the collected sound into an analog signal, which is an electrical signal, and outputs the analog signal. The microphone 161 of this embodiment includes three microphone elements 161L, 161C, and 161R. The microphone 161 may be composed of two or four or more microphone elements.

[0031] The microphone A / D converter 165 converts the analog signal from the microphone 161 into digital audio data. The microphone A / D converter 165 is an example of an audio acquisition unit in this embodiment. The microphone 161 may include a microphone element external to the digital camera 100. In this case, the digital camera 100 includes an interface circuit for the external microphone 161 as the audio acquisition unit.

[0032] The audio processing engine 170 receives audio data output from an audio acquisition unit such as the microphone A / D converter 165, and performs various audio processes on the received audio data. The audio processing engine 170 is an example of an audio processing unit in this embodiment.

[0033] The sound processing engine 170 of this embodiment includes a beam forming unit 172 and a gain adjusting unit 174, as shown in FIG. 1, for example. The beam forming unit 172 realizes a function of controlling sound directivity. The beam forming unit 172 will be described in detail later. The gain adjusting unit 174 amplifies the sound by performing a multiplication process of multiplying the input sound data by a sound collection gain set by the controller 135, for example. The gain adjusting unit 174 may perform a process of suppressing the sound by multiplying the input sound data by a negative gain. The sound collection gain adjusting unit 174 may further have a function of changing the frequency characteristics and stereo characteristics of the input sound data. The sound collection gain setting will be described in detail later.

[0034] [1-1-1. Beam forming section] The beam forming section 172 in this embodiment will be described in detail below.

[0035] The beam forming unit 172 performs beam forming to control the directivity of the sound picked up by the microphone 161. An example of the configuration of the beam forming unit 172 in this embodiment is shown in FIG.

[0036] 5, the beam forming unit 172 includes, for example, filters D1 to D3 and an adder 173, and adjusts the delay periods of the sounds collected by the microphone elements 161L, 161C, and 161R, and outputs the weighted sum. The beam forming unit 172 controls the direction and range of the sound collection directivity of the microphone 161, and can set the physical range in which the microphone 161 collects sound.

[0037] In the figure, the beam forming unit 172 outputs one channel using one adder 173, but it may also be configured to have two or more adders and output different signals for each channel, such as a stereo output. Furthermore, a subtractor may be used in addition to the adder 173 to form directivity with a blind spot, which is a direction with particularly low sensitivity, in a specific direction, or adaptive beamforming may be performed, which changes processing depending on the environment. Different processing may also be applied depending on the frequency band of the audio signal.

[0038] 5 shows an example in which the microphone elements 161L, 161C, and 161R are linearly arranged, but the arrangement of the microphone elements is not limited to this. For example, even in the case of a triangular arrangement, the sound collection directivity of the microphone 161 can be controlled by appropriately adjusting the delay periods and weights of the filters D1 to D3. The beam forming unit 172 may also apply a known method to control the sound collection directivity. For example, an audio processing technology such as OZO Audio may be used to perform a process of forming directivity and also to perform a process of suppressing noise in the audio.

[0039] The sound collection area of ​​the digital camera 100 that can be set by the beam forming section 172 as described above will be described.

[0040] [1-1-2. About the recording area] Fig. 6 shows an example of a sound collection area defined in digital camera 100. Fig. 6 illustrates the sound collection area as a sector of a circle centered on digital camera 100. In digital camera 100 of this embodiment, the horizontal angle of view direction coincides with the direction in which microphone elements 161R, 161C, and 161R are aligned.

[0041] FIG. 6(A) shows a "front center sound collection area" 41 that faces the sound collection area in front of the digital camera 100 (i.e., the shooting direction) within an angle range 401 (e.g., 70°). FIG. 6(B) shows a "left half sound collection area" 42 that faces the sound collection area to the left of the digital camera 100 within the angle range 401. FIG. 6(C) shows a "right half sound collection area" 43 that faces the sound collection area to the right of the digital camera 100 within the angle range 401. FIG. 6(D) shows a "front sound collection area" 44 that faces the sound collection area in front of the digital camera 100 within an angle range 402 (e.g., 160°) that is larger than the angle range 401. These sound collection areas are examples of the plurality of predetermined areas in this embodiment, and the angle ranges 401 and 402 are examples of a first angle range and a second angle range.

[0042] When the subject is located in the center of the captured image, the digital camera 100 of this embodiment uses the front center sound collection area 41 in Fig. 6(A). When the subject is located in the left half of the captured image, the left half sound collection area 42 in Fig. 6(B) is used, and when the subject is located in the right half of the captured image, the right half sound collection area 43 in Fig. 6(C) is used. When the subject is located over the entire captured image, the digital camera 100 mainly uses the front sound collection area 44 in Fig. 6(D).

[0043] In the example of Fig. 11(B), subjects R1 and R3 to be collected are located in the center of the captured image, so a front center sound collection area 41 is used. In the example of Fig. 11(C), subjects R1 and R2 to be collected are located in the left half of the captured image, so a left half sound collection area 42 is used.

[0044] In a sound collection section of an imaging device, such as microphone 161 of digital camera 100, the number and arrangement of microphone elements are limited by factors such as the space available for mounting the elements. For example, in a shooting scene where a user wants to record audio from multiple subjects, the limited number of microphone elements may make it impossible to sufficiently narrow the sound collection directionality. Even in such a case, digital camera 100 of this embodiment can provide a sound collection area that meets the user's intentions by defining a sound collection area based on the user's expected shooting scene and determining the sound collection area using face recognition.

[0045] [1-1-3. Microphone settings] The settings relating to the sound collection area of ​​microphone 161 in digital camera 100 will be described with reference to FIGS.

[0046] 7 shows an example of a setting menu display in digital camera 100. Digital camera 100 of this embodiment has, for example, modes such as "auto," "surround," "front," "focus," and "narration" as operation modes (i.e., sound collection modes) that control the sound collection area of ​​microphone 161, as shown in FIG.

[0047] The focus mode is an operating mode that automatically changes the directionality of the microphone 161 and adjusts the sound collection area in conjunction with face recognition and the angle of view by the digital camera 100. For example, the focus mode can be realized by switching between the various sound collection areas 41 to 44 described above. By intentionally roughly realizing the focus mode using the four sound collection areas 41 to 44, it is possible to avoid situations in which the sound collection directionality frequently changes due to slight movements of the subject, and to reduce the auditory annoyance experienced by the user. Additional sound collection areas in the digital camera 100 are shown in FIG. 8.

[0048] 8(A) illustrates sound collection area 45 in surround mode. Surround mode is an operating mode for collecting sounds over a wide range, including the left and right, front and rear of digital camera 100. Sound collection area 45 in surround mode has an angular range of, for example, 360° all around the XZ plane.

[0049] 8(B) illustrates sound collection area 46 in front mode. Front mode is an operating mode for collecting sounds in front of digital camera 100. Sound collection area 46 in front mode faces the +Z side from digital camera 100 and has an angular range equal to or larger than that of front sound collection area 44 described above, for example.

[0050] 8(C) illustrates sound collection area 47 in narration mode. Narration mode is an operating mode for collecting sounds behind digital camera 100. Sound collection area 47 in narration mode is formed toward the -Z side from digital camera 100. When narration mode is set and it is detected that display monitor 130 is in the selfie position, digital camera 100 may perform the operation of focus mode.

[0051] The auto mode is an operating mode that automatically changes the directionality of the microphone 161 and adjusts the sound collection area in conjunction with the shooting state of the digital camera 100. The shooting state of the digital camera 100 includes, for example, whether a self-portrait is being taken and whether the image is being taken vertically or horizontally, in addition to face recognition and the like that are taken into account in the focus mode.

[0052] The microphone settings for setting the various sound collection modes as described above are provided, for example, as one of the video menus in the setting menu of the digital camera 100. The user can select a desired sound collection mode from the setting menu by touching the touch panel 155 or pressing the various buttons 152 and 153. The microphone settings may also be assigned in advance to the function buttons 154, etc. The setting of a specific sound collection mode such as auto mode may also be assigned to the function buttons 154, etc.

[0053] [1-2. Operation] The following describes the operation of the digital camera 100 configured as above. The following describes the operation of the digital camera 100 when shooting moving images.

[0054] The digital camera 100 sequentially captures the subject image formed via the optical system 110 with the image sensor 115 to generate imaging data. The image processing engine 120 performs various processes on the imaging data generated by the image sensor 115 to generate image data and records the image data in the buffer memory 125. Furthermore, the face recognition unit 122 of the image processing engine 120 detects the area of ​​the subject based on the image indicated by the imaging data and outputs detection information to the controller 135, for example.

[0055] The digital camera 100 of this embodiment has a face recognition mode, which is an operating mode in which faces are detected by image recognition processing in the captured image input to the face recognition unit 122, and a subject to be subjected to autofocus (AF) control is identified based on the detection information.

[0056] Simultaneously with the above imaging operation, the digital camera 100 collects sound using the microphone 161. The collected sound data is output from the microphone A / D converter 165 and processed by the audio processing engine 170. The audio processing engine 170 records the processed audio data Aout in the buffer memory 125.

[0057] The controller 135 synchronizes the image data received from the image processing engine 120 and the audio data received from the audio processing engine 170 via the buffer memory 125, and records the moving image on the memory card 142. The controller 135 also sequentially displays a through image on the display monitor 130. The user can check the composition of the shot at any time by viewing the through image on the display monitor 130. The moving image shooting operation is started / ended in response to a user operation on the operation unit 150.

[0058] When shooting video with the digital camera 100 as described above, there are various situations in which the user may want to record sound in various shooting conditions. For example, there may be cases where the user focuses on a group of subjects having a conversation among themselves, such as the photographer and their companion. In such cases, there may be a need to clearly record the vocalizations of the group of subjects.

[0059] In the digital camera 100 of this embodiment, for example, in the above-described focus mode operation, a subject is detected using detection information from the face recognition unit 122 in the image processing engine 120, and when the subject to be AF-targeted is determined, the sound processing engine 170 executes processing to emphasize the sounds picked up for that subject and for subjects around that subject in the shooting space. In this way, by linking the face recognition of the image processing engine 120 with the sound emphasis and the like of the sound processing engine 170, accurate sound pickup is realized in which the sounds of a group of subjects having a conversation as described above are emphasized.

[0060] Furthermore, in addition to the focus mode operation described above, digital camera 100 of this embodiment also realizes appropriate sound collection control in accordance with various shooting conditions as an auto mode operation. An overview of the auto mode operation will be described using FIG. 9.

[0061] 9 illustrates the correspondence between the auto mode and various sound collection modes of the digital camera 100. When face recognition is performed in landscape mode, for example, the digital camera 100 in auto mode performs the same operation as in focus mode.

[0062] On the other hand, when face recognition is not performed, for example, when a selfie is not being taken (see FIG. 2), digital camera 100 operates in the same way as in surround mode, that is, it employs surround mode sound collection area 45. Also, when face recognition is not performed and a selfie is being taken (see FIG. 3), digital camera 100 operates in the same way as in front mode.

[0063] In addition, in the case of vertical shooting, the operation when face recognition is not performed is the same as in the case of horizontal shooting described above. On the other hand, when face recognition is performed in vertical shooting, the digital camera 100 of this embodiment performs the same operation as in front mode instead of focus mode.

[0064] According to the above-described auto mode operation, by combining the operations of various sound collection modes according to various shooting conditions, as shown in FIG. 9, it is possible to easily achieve appropriate sound collection control in each shooting condition.

[0065] [1-2-1. Focus mode operation] An overview of the operation of the focus mode of the digital camera 100 according to this embodiment will be described with reference to FIGS.

[0066] Fig. 10 is a flowchart illustrating the operation of the focus mode of the digital camera 100 according to this embodiment. Each process shown in the flowchart of Fig. 10 is repeatedly executed at a predetermined cycle, for example, when the digital camera 100 is set to the focus mode. The predetermined cycle is, for example, the frame cycle of a moving image. Fig. 11 is a diagram for explaining an overview of the operation of the focus mode of the digital camera 100 according to this embodiment.

[0067] The controller 135 identifies an AF target based on detection information from the face recognition unit 122 and performs AF control (S1). The AF target indicates an area on an image of a subject that is the target of AF control. FIG. 11(A) illustrates a captured image Im including face areas R1, R2, and R3 that indicate areas where the subject is detected in the detection information from the face recognition unit 122. The face areas R1, R2, and R3 are examples of subject areas in this embodiment. For example, the face area R1 is identified as the face area 60 of the AF target.

[0068] Next, the controller 135 determines whether or not a face area identified as an AF target exists (S2). Specifically, the controller 135 determines whether or not a face area has been detected and whether or not the AF target is a face area.

[0069] If there is a face area 60 to be an AF target (YES in S2), the controller 135 executes a process of selecting a sound collection target for the microphone 161 from the subjects in the detection information (S3). A sound collection target is a subject whose sound is to be emphasized and collected by the microphone 161. The face area R1 (60) identified as the AF target becomes the sound collection target. FIG. 11(B) shows an example in which, based on the detection information shown in FIG. 11(A), face areas R1 and R3 are determined to be sound collection targets, while face area R2 is not determined to be a sound collection target.

[0070] In the sound collection target selection process (S3), the digital camera 100 of this embodiment determines, in addition to the face area R1 (60) that is the AF target, a face R3 that is approximately the same size as the face area R1 in the captured image Im as an additional sound collection target. On the other hand, a face area R2 that is a different size from the face area R1 is excluded from the sound collection targets. This reflects the fact that person 21 and person 23 are at approximately the same distance from the digital camera 100 (i.e., the difference in distance in the Z axis direction is small), while person 22 is at a different distance, and therefore, for example, a group of subjects having a conversation among themselves can be set as a sound collection target. The sound collection target selection process (S3) will be described in detail later.

[0071] Next, the controller 135 performs a process of determining a sound collection area based on the determined sound collection targets (S4). The sound collection area determination process (S4) determines a sound collection area that includes all of the determined sound collection targets. In the example of FIG. 11(B), the sound collection area is determined to be the front center sound collection area 41 (FIG. 6(A)) so as to include face areas R1 and R3 of the sound collection targets. The sound collection area determination process (S4) will be described in detail later.

[0072] Next, the controller 135 controls sound collection using face recognition based on the determined sound collection target and sound collection area (S5). The sound collection control using face recognition (S5) is performed by setting sound collection parameters including the sound collection target, sound collection area, and sound collection gain determined by the controller 135 in the sound processing engine 170. The sound processing engine 170 realizes sound collection directivity and sound collection gain according to the sound collection parameters.

[0073] On the other hand, if there is no face area 60 to be an AF target (NO in S2), for example, if no face area is detected during operation in the face recognition mode, the controller 135 performs sound collection control (S6) without using face recognition. Sound collection control with and without using face recognition (S5, S6) will be described in detail later.

[0074] After executing the sound collection control in step S5 or S6, the controller 135 repeats the processing from step S1 onwards.

[0075] According to the above process, the digital camera 100 of this embodiment selects sound collection targets from subjects detected by face recognition, determines a sound collection area that includes all of the sound collection targets, and performs sound collection control using face recognition. As a result, it is possible to collect sound with emphasis on the voices of a group of subjects having a conversation among themselves, for example.

[0076] In the AF control by face recognition (S1), the AF target can be identified based on the detection information, for example, by displaying a frame indicating the face area on the through image displayed on the display monitor 130, and receiving an operation by the user to select the frame display using the operation unit 150.

[0077] FIG. 11C shows an example of a captured image Im in which people 21 to 23 are located in positions different from those shown in FIGS. 11A and 11B. As in the example of FIG. 11B, the digital camera 100 first identifies, for example, face area R1 as face area 60 to be the AF target (S1) and determines it as the sound collection target. In the example of FIG. 11C, the sound collection target selection process (S3) determines face area R2, which has a face of approximately the same size as face area R1 on the captured image Im, as the sound collection target, and excludes face area R3 from the sound collection target. The sound collection area determination process (S4) determines the left half sound collection area 42 (FIG. 6B) including face areas R1 and R2 determined as the sound collection targets as the sound collection area. The sound collection control using face recognition (S5) is performed by controlling the directivity of the left half sound collection area 42 and setting sound collection parameters so as to clearly collect the voices of people 21 and 22.

[0078] [1-2-2. Sound recording selection process] The details of the sound collection target selection process in step S3 of FIG. 10 will be described with reference to FIGS.

[0079] 12 is a flowchart illustrating the process (S3) of selecting sound collection targets of the digital camera 100. The processes in the flowchart shown in FIG. 12 are executed by, for example, the controller 135 of the digital camera 100 when the process proceeds to YES in step S11 of FIG.

[0080] 13 is a diagram for explaining the process of selecting a sound collection target (S3) in the digital camera 100. The operation of determining a sound collection target will be explained below using the examples of FIGS.

[0081] 12, the controller 135 determines, as a sound collection target, a subject corresponding to the face area of ​​the AF target identified in step S1 of Fig. 10 (S10). At this time, the controller 135 sets the size of the face area of ​​the AF target (i.e., face width W) as a criterion for selecting the sound collection target from other subjects, based on the detection information acquired from the face recognition unit 122.

[0082] 13(A) illustrates a case where a sound collection target is selected in the example of FIGS. 11(A) and (B). Face widths W1, W2, and W3 indicate the size of face areas R1, R2, and R3 in the captured image Im as widths in the X-axis direction. In the example of FIG. 13(A), the controller 135 sets the face width W1 of the face area R1 of the AF target to the reference face width W (S10). The set face width W is stored in, for example, the RAM of the controller 135.

[0083] Next, the controller 135 determines whether or not there is a detected subject other than the AF target (S11). Specifically, the controller 135 determines whether or not the detection information of the face recognition unit 122 includes a face area other than the face area of ​​the AF target.

[0084] If there is a detected subject other than the AF target (YES in S11), the controller 135 selects one subject i as a sound collection candidate that is a candidate for a sound collection target (S12). In the example of Fig. 13(A), the detection information is such that other face areas R2 and R3 than the AF target face area R1 are sequentially associated with the subject i of the sound collection candidate for each step S12 and selected.

[0085] The controller 135 performs a calculation to compare the face width Wi of the selected subject i with the reference face width W (S13). Specifically, the controller 135 calculates the ratio Wi / W of the face width Wi of the subject i to the reference face width W. In the example of Fig. 13(A), when the face area R2 is selected as a sound collection candidate (S12), the ratio W2 / W for the face width W2 is calculated (S13).

[0086] The controller 135 determines whether the ratio Wi / W between the face width Wi of the sound collection candidate and the reference face width W is within a predetermined range (S14). The predetermined range is defined by an upper limit value greater than "1" and a lower limit value less than "1", for example, from the viewpoint of defining a range in which the face width Wi of the sound collection candidate is considered to be relatively similar to the reference face width Wi. Note that a user interface for setting the predetermined range may be provided, and the predetermined range set by the user via the operation unit 150 may be stored in the buffer memory 125 or the like, for example.

[0087] When the controller 135 determines that the face width ratio Wi / W is within a predetermined range (YES in S14), it determines that the subject i is to be the sound collection target (S15).

[0088] On the other hand, if the controller 135 determines that the face width ratio W i / W is not within the predetermined range (NO in S14), the controller 135 determines that the subject i is not to be included in the sound collection target (S16). In the example of Fig. 13(A), the ratio W2 / W is below the lower limit of the predetermined range, and it is determined that the face area R2 is not to be included in the sound collection target.

[0089] When the controller 135 determines whether or not to make the subject i a sound collection target (S15 or S16), the controller 135 records the result of the determination for the subject i, for example, in the buffer memory 125 (S17). Next, the controller 135 performs the processes from step S11 onward again for a subject other than the subject that has already been selected as a sound collection candidate.

[0090] In the example of FIG. 13(A), in addition to face region R2, face region R3 is included in the detection information (YES in S11). When controller 135 selects a subject corresponding to face region R3 (S12), it calculates the ratio W3 / W of face width W3 to the reference face width W, as in the case of face region R2 (S13). In the example of FIG. 13(A), the ratio W3 / W is calculated to be close to "1." Controller 135 determines that the calculated face width ratio W3 / W is within a predetermined range of the sound collection target (YES in S14), and determines the subject corresponding to face region R3 as the sound collection target (S15).

[0091] The controller 135 repeats the processes of steps S11 to S17 until there are no subjects that have not been selected as sound collection candidates (NO in step S11). After that, the controller 135 ends the sound collection target selection process (S3) and proceeds to step S4 in FIG.

[0092] According to the above process, for subjects detected by face recognition, the relative sizes of face areas R2 and R3 are compared using face area R1 identified as the AF target as a reference. This makes it possible to select subjects whose relative face area R3 is about the same size as the AF target face area R1 and determine them as sound collection targets.

[0093] Fig. 13(B) illustrates a case where a sound collection target is selected in the example of Fig. 11(C). In the example of Fig. 13(B), the face area R1 is identified as the AF target, similar to the example of Fig. 13(A). Therefore, the controller 135 determines the face area R1 as the sound collection target and sets the face width W1 as the reference face width W (S10).

[0094] In the example of FIG. 13(B), the face width W2 of face region R2 is approximately the same as the reference face width W (= W1). On the other hand, the face width W3 of face region R3 is larger than the other face widths W1 and W2. In this example, the controller 135 determines that the ratio W2 / W is within a predetermined range (YES in S14), and determines the subject in face region R2 as a sound collection target (S15). On the other hand, because the ratio W3 / W exceeds the upper limit of the predetermined range (NO in S14), it determines that the subject in face region R3 is not a sound collection target (S16). Therefore, the two subjects corresponding to face regions R1 and R2 are determined as sound collection targets in this example (see FIG. 11(C)).

[0095] FIG. 13(C) illustrates a case where a face region R3 is identified as the face region 60 of the AF target in a captured image Im similar to that of FIG. 11(C) (S1 in FIG. 10). The controller 135 determines the face region R3 as the sound collection target and sets the face width W3 as the reference face width W (S10). In the example of FIG. 13(C), because the ratios W2 / W and W1 / W are below the lower limit of the predetermined range (NO in S14), it is determined that the subjects corresponding to the face regions R1 and R2 are not to be the sound collection target (S16). Therefore, the sound collection target in this example is determined to be one subject corresponding to the face region R3.

[0096] As described above, the digital camera 100 of this embodiment can be used to determine a sound collection area in accordance with the user's intentions, as described below, by determining, from among multiple subjects detected by image recognition, a subject that is approximately the same size as the AF target as the sound collection target.

[0097] [1-2-3. Sound pickup area determination process] The details of the sound collection area determination process in step S4 of FIG. 10 will be described with reference to FIGS.

[0098] 14 is a flowchart illustrating the sound collection area determination process (S4) in the digital camera 100 of this embodiment. Each process in the flowchart shown in FIG. 14 is executed by, for example, the controller 135 of the digital camera 100 after executing step S3 in FIG.

[0099] FIG. 15 is a diagram illustrating the sound collection area determination process (S4) in the digital camera 100. FIGS. 15(A) and 15(B) illustrate an example of determining a sound collection area following the example of FIGS. 13(A) and 13(B), respectively. FIG. 15(C) illustrates a case different from FIGS. 15(A) and 15(B). In FIGS. 15(A) to 15(C), the center position x0 indicates the center position of the captured image Im in the X-axis direction, and the image width Wh indicates the width of the captured image Im in the X-axis direction. The image range is defined as the range x0±xh from X coordinate -xh to xh on the captured image Im, with the center position x0 as the base. The X coordinate xh is defined as xh=Wh / 2 (>0).

[0100] 14, the controller 135 determines whether the position of the center of the face region or the like of each sound collection target is within the central range of the captured image Im (S20). The central range is a range associated with the front central sound collection area 41 in the captured image Im.

[0101] 15A, the central range is defined as a range x0±xe from X coordinate −xe to xe on the captured image Im, with the center position x0 as the reference. The X coordinate xe is defined, for example, as xe=xh×θe / θh (>0) based on a predetermined angle of view θe and a horizontal angle of view θh corresponding to the image width Wh. The predetermined angle of view θe is set in advance, for example, from the perspective of including one person, and is, for example, 30°. The controller 135 obtains the current horizontal angle of view θh from, for example, the zoom magnification of the zoom lens of the optical system 110, and calculates the central range x0±xe.

[0102] In wide-angle shooting with a large horizontal angle of view θh, the X coordinate xe is small and the central range x0±xe is narrow. On the other hand, in telephoto shooting with a small horizontal angle of view θh, the X coordinate xe is large and the central range x0±xe is wide. This makes it easier to determine the sound collection area that corresponds to the physical range and distance to be captured.

[0103] If the positions of the face areas of all sound collection targets are within the central range (YES in S20), the controller 135 determines the sound collection area to be the front central sound collection area 41 (S21). In the example of FIG. 15(A), the sound collection targets correspond to face areas R1 and R3. The center positions x1 and x3 of the face areas R1 and R3 are both within the range of x0±xe (YES in S20). Therefore, the sound collection area is determined to be the front central sound collection area 41 (S21, see FIG. 11(B)).

[0104] On the other hand, if the position of the face area of ​​at least one sound collection target is not within the central range (NO in S20), a sound collection area other than the front central sound collection area 41 is used. In this case, the controller 135 determines, for all sound collection targets, for example, whether the position of the face area is within only the left or right half of the captured image Im (S22). The left half is the range whose X coordinate is smaller than the central position x0 in the X-axis direction, and the right half is the range whose X coordinate is larger than the central position x0.

[0105] If the positions of the face areas of all the sound collection targets are within only the left half or right half of the captured image Im (YES in S22), the controller 135 further determines whether the positions of the face areas of all the sound collection targets are within the left half of the captured image Im (S23).

[0106] If the positions of the face regions of all sound collection targets are within the left half of the captured image Im (YES in S23), the controller 135 determines the sound collection area to be the left half sound collection area 42 (S24). In the example of FIG. 15(B), the sound collection targets correspond to face regions R1 and R2. Since the position x1 of face region R1 and the position x2 of face region R2 are on the left side (i.e., the X coordinate is smaller) of the center position x0 in the X-axis direction (YES in S23), the sound collection area is determined to be the left half sound collection area 42 (S24, see FIG. 11(C)).

[0107] On the other hand, if the positions of all face areas of the sound collection targets are within the right half of the captured image Im but not within the left half (NO in S23), the controller 135 determines the sound collection area to be the right half sound collection area 43 (S25).

[0108] 6(D) and 6(A), the front sound collection area 44 has an angle range 402 wider than the angle range 401 of the front central sound collection area 41. In other words, the front sound collection area 44 includes subjects of the sound collection targets located over a wide range in the X-axis direction in the captured image Im.

[0109] 15(C), the sound collection targets correspond to facial areas R1, R2, and R3. The central positions x1, x2, and x3 of the facial areas R1 to R3 include positions x1 and x2 outside the central range x0±xe (NO in S20), and also include position x1 within the left half range and positions x2 and x3 within the right half range (NO in S22 and S23). Therefore, in this example, the sound collection area is determined to be the front sound collection area 44 (S26).

[0110] When the controller 135 determines the sound collection area (S21, S24 to S26), it records the determined sound collection area as management information in the buffer memory 125 or the like (S27). This ends the sound collection area determination process (S4), and the process proceeds to step S5 in FIG. 10.

[0111] According to the above process, a sound collection area is determined from a plurality of predefined sound collection areas in accordance with the position on the captured image of the subject determined as the sound collection target so as to include all of the sound collection targets. This makes it possible to determine a sound collection area in video shooting so as to include the subject as the sound collection target in accordance with the user's intention.

[0112] Fig. 17 is a diagram for explaining management information obtained by the sound collection area determination process (S4). Fig. 17(A) illustrates management information obtained at the stage when the sound collection target selection process (S3) and the sound collection area determination process (S4) are executed in the examples of Fig. 13(A) and Fig. 15(A). Fig. 17(B) illustrates management information in the examples of Fig. 13(B) and Fig. 15(B).

[0113] The management information manages, for example, the "sound collection target" determined by the sound collection target selection process (S3), the "sound collection area" determined by the sound collection area determination process (S4), the "horizontal angle of view," and the "focus distance" in association with each other. The focus distance is acquired, for example, when executing AF control (S1) using face recognition. For example, the controller 135 may acquire the corresponding focus distance based on the positions or focal lengths of the various lenses of the optical system 110 when focusing. The digital camera 100 may also detect the focus distance using DFD (Depth from Defocus) technology or measurement using a distance measurement sensor.

[0114] It should be noted that digital camera 100 of this embodiment can set the angle of view θe of the central range used in determining the front central sound collection area (S20), and this is recorded, for example, in a ROM or the like of controller 135. Also, a user interface for setting the angle of view θe may be provided, and the value set by the user via operation unit 150 may be stored in buffer memory 125 or the like.

[0115] [1-2-4. Sound pickup control] (1) Step S5 in Figure 10 The details of the sound collection control using face recognition in step S5 of FIG. 10 will be described with reference to FIGS.

[0116] In sound collection control using sound collection parameter settings, the digital camera 100 of this embodiment sets the sound collection gain so as to emphasize video audio for a subject corresponding to the face region of the AF target, for example. The sound collection gain has, for example, frequency filter characteristics and stereo separation characteristics. The digital camera 100 calculates the sound collection gain based on the horizontal angle of view and the focal distance when the digital camera 100 focuses on the face region of the AF target while shooting a video. The sound collection gain is specified so that, for example, the larger the calculated value, the more it suppresses frequency bands other than human voices and controls the stereo effect, thereby creating a sound collection zoom effect.

[0117] Fig. 16 is a flowchart illustrating sound collection control (S5) using face recognition. Each process shown in the flowchart of Fig. 16 is executed by, for example, the controller 135 of the digital camera 100 after executing step S4 of Fig. 10.

[0118] The digital camera 100 starts the process of step S5 with the management information shown in FIG. 17 held.

[0119] The controller 135 acquires the horizontal angle of view from, for example, the buffer memory 125, and calculates the gain Gh based on the horizontal angle of view (S30). Fig. 18A illustrates an example of the relationship for calculating the gain Gh from the horizontal angle of view. In the example of Fig. 18A, the gain Gh increases as the horizontal angle of view becomes smaller between a predetermined maximum value Gmax and minimum value Gmin of the gain. This allows the gain to be increased during sound collection as the horizontal angle of view becomes smaller due to zooming, etc., thereby emphasizing the sound of a subject captured at a telephoto end.

[0120] Controller 135 acquires the focus distance as in step S30 and calculates gain Gd based on the focus distance (S31). Figure 18(B) illustrates an example of the relationship for calculating gain Gd from focus distance. In the example of Figure 18(B), gain Gd increases as the focus distance increases, between a predetermined maximum value Gmax and minimum value Gmin of the gain. This allows the gain to be increased when focusing on an object farther from digital camera 100, thereby emphasizing the sound of the object that is farther away.

[0121] The controller 135 compares the calculated sound collection gain Gh based on the horizontal angle of view with the sound collection gain Gd based on the focal distance, and sets the larger gain as the sound collection gain G (S32). In this way, the sound collection gain G can be calculated so as to emphasize the sound of the subject in accordance with the intention of the user who is taking a picture at a telephoto horizontal angle of view or a long focal distance, for example.

[0122] The controller 135 determines whether the sound collection gain G calculated a predetermined number of times in the past (for example, five times) and the determined sound collection area are the same (S33). For example, the sound collection gain G is stored together with the above management information each time it is calculated within a predetermined number of times in the execution cycle of steps S1 to S5 in Fig. 10. If the controller 135 determines that the sound collection gain G and sound collection area calculated a predetermined number of times in the past are the same (YES in S33), the controller 135 proceeds to step S34.

[0123] The controller 135 sets the sound collection target determined by the sound collection target selection process in step S3, the sound collection area determined by the sound collection area determination process in step S4, and the sound collection gain G calculated in step S32 as sound collection parameters in the sound processing engine 170 (S34). The sound processing engine 170 realizes the sound collection area and sound collection gain according to the set sound collection parameters by the beam forming unit 172 and the gain adjusting unit 174.

[0124] After setting the sound collection parameters (S34), the controller 135 ends the sound collection control process using face recognition (S5). Furthermore, if the controller 135 determines that the sound collection gain G and sound collection area are not the same for a predetermined number of times in the past (NO in S33), it ends the process of step S5 in Fig. 10 without performing the process of step S34. Thereafter, the processes from step S1 onwards in Fig. 10 are repeated.

[0125] According to the above processing, the calculated sound collection gain and the sound collection target and sound collection area determined based on face recognition can be set as sound collection parameters, thereby realizing a sound collection area and sound collection gain that makes it easier to clearly collect the sound of the sound collection target subject, including the AF target.

[0126] The order of execution of steps S30 and S31 is not limited to the order shown in this flowchart. For example, gain Gd may be calculated in step S31 and then gain Gh may be calculated in step S30, or steps S30 and S31 may be executed in parallel.

[0127] Furthermore, according to the above step S33, the process (S34) of setting the sound collection parameters is executed only if the sound collection area and sound collection gain G have not changed a predetermined number of times (for example, five times). This prevents the sound collection area and sound collection gain G from being changed too frequently due to the movement of the subject, etc., and makes it possible to accurately realize sound collection control (S5) using face recognition in line with the user's intentions.

[0128] (2) Step S6 in Figure 10 The details of the sound collection control (S6) without using face recognition in step S6 of FIG. 10 will be described with reference to FIG.

[0129] Fig. 19 is a flowchart illustrating sound collection control (S6) without using face recognition. The processes shown in the flowchart of Fig. 19 are executed by, for example, controller 135 of digital camera 100 when there is no face area to be AF-targeted in step S2 of Fig. 10 (NO in S2), such as when a face area is not detected.

[0130] First, the controller 135 determines the sound collection area to be, for example, the front sound collection area 44 (S40).

[0131] Next, the controller 135 calculates a gain Gh based on the horizontal angle of view in the same manner as in step S30, and sets the calculated gain Gh as the sound collection gain G (S41). Furthermore, the controller 135 determines whether the sound collection gain G calculated a predetermined number of times in the past and the determined sound collection area are the same as each other (S42), in the same manner as in step S33.

[0132] If the controller 135 determines that the sound collection gain G and sound collection area are the same for the predetermined number of times in the past (YES in S42), it sets the sound collection area and sound collection gain G as sound collection parameters (S43) and ends the sound collection control without face recognition (S6). If the controller 135 determines that the sound collection gain G and sound collection area are not the same for the predetermined number of times in the past (NO in S42), it ends step S6 in Fig. 10 without performing the process of step S43. After step S6 ends, the processes from step S1 onwards are repeated.

[0133] According to the above processing, even if there is no face area to be AF-targeted, sound from a wide range in front of the digital camera 100 can be picked up, and the smaller the horizontal angle of view due to zooming, etc., the larger the sound pickup gain, making it easier to clearly pick up sound from the range to be captured.

[0134] Note that a total sound collection area having an angular range of 360° around digital camera 100 may be defined depending on the operation mode of digital camera 100, and may be determined as the total sound collection area in step S40. At this time, for example, only the total sound collection area may be set as the sound collection parameter.

[0135] [1-2-5. Auto mode operation] The operation of the digital camera 100 according to this embodiment in auto mode will be described in detail with reference to FIGS.

[0136] Fig. 20 is a flowchart illustrating an example of the operation in auto mode of digital camera 100 according to embodiment 1. Each process shown in the flowchart in Fig. 20 is executed by controller 135 when digital camera 100 is set to auto mode, for example, as in Fig. 10.

[0137] 20, in digital camera 100 in auto mode, controller 135 determines whether display monitor 130 is in the selfie position based on, for example, the detection signal of magnetic sensor 132 (S51). If controller 135 determines that display monitor 130 is not in the selfie position (NO in S51), it sets the sound collection area for non-face recognition to sound collection area 45 in surround mode (S52). On the other hand, if controller 135 determines that display monitor 130 is in the selfie position (YES in S51), it sets the sound collection area for non-face recognition to sound collection area 46 in front mode (S53).

[0138] The controller 135 performs face recognition processing (S1, S2) in the same manner as in the focus mode described above. For example, if a face area targeted for AF is detected (YES in S2), the controller 135 determines whether the digital camera 100 is in a vertical shooting position based on the detection signal from the acceleration sensor 137 (S54). If the controller 135 determines that the digital camera 100 is not in a vertical shooting position (NO in S54), it performs the same processing as in steps S3 to S5 in the focus mode and executes sound collection control. On the other hand, if the controller 135 determines that the digital camera 100 is in a vertical shooting position (YES in S54), it employs the sound collection area 46 in the front mode and executes sound collection control (S55).

[0139] If the face area of ​​the AF target is not detected (NO in S2), the controller 135 performs sound collection control without face recognition based on the setting results of steps S52 and S53 (S6A). The sound collection control in step S6A is performed in the same manner as in step S6 described above, using the sound collection area set as the sound collection area when face recognition is not performed.

[0140] The above processing makes it possible to realize an auto mode operation that adjusts the directivity of the microphone 161 in conjunction with various shooting conditions. Sound collection control in the auto mode during horizontal and vertical shooting will be further described with reference to FIG.

[0141] Fig. 21(A) illustrates the relationship between the captured image Im and the sound collection areas 41 to 43 when taken horizontally. Fig. 21(B) illustrates the case when step S55 is not performed when taken vertically. Fig. 21(C) illustrates the case when step S55 is performed when taken vertically.

[0142] As shown in Fig. 21(A), during horizontal capture, the sound collection areas, facial areas R1-R3, and sound collection areas 41-43 are all coaxial, and switching between the sound collection areas 41-43 allows for intended sound collection control to match the positions of the desired facial areas R1-R3. However, during vertical capture, as shown in Fig. 21(B), the relationships between the sound collection areas 41-43 and the facial areas R1-R3 do not match. This can lead to situations where sound collection control cannot be performed as intended for the positions of the desired facial areas R1-R3, and instead unintended sound collection control occurs.

[0143] Therefore, in this embodiment, during vertical shooting, the sound collection area is fixed to the sound collection area 46 in front mode, as shown in Fig. 21(C). This ensures that the range in which sound can be collected is the entire range during imaging, and makes it possible to avoid situations in which unintended sound collection control occurs.

[0144] [1-3. Effects, etc.] In this embodiment, the digital camera 100 includes an image sensor 115 as an example of an imaging unit, a microphone 161 as an example of an audio acquisition unit, an operation unit 150 as an example of a setting unit, and a controller 135 as an example of a control unit. The image sensor 115 captures an image of a subject and generates image data. The microphone 161 acquires an audio signal indicating audio picked up during image capture by the imaging unit. The operation unit 150 receives a user instruction and sets the digital camera 100 to an auto mode, an operating mode that automatically changes the directionality of the audio acquisition unit. The controller 135 controls the audio pickup area for picking up audio from the subject in the audio signal. When set to the auto mode, the controller 135 controls the audio pickup area to include the subject by changing the directionality of the microphone 161 in conjunction with the imaging state of the digital camera 100. This enables appropriate audio pickup control according to various imaging conditions, making it easier to pick up the audio of the subject according to the user's intentions when capturing an image while acquiring audio.

[0145] The digital camera 100 of this embodiment includes a face recognition unit 122, which is an example of a face detection unit that detects the facial area of ​​a subject in image data. When the digital camera 100 is set to auto mode, the controller 135 determines the subject to be collected in the audio signal based on the facial area detected by the face recognition unit 122, and controls the sound collection area so that the determined subject is included in the sound collection target. This allows sound collection control to be performed according to the shooting conditions, such as various facial recognition results of the subject, making it easier to collect sound in line with the user's intentions.

[0146] In this embodiment, when the controller 135 is set to the auto mode, the controller 135 controls the sound collection area so as to change the directivity of the sound acquisition unit in conjunction with the shooting state, whether the device is in portrait or landscape orientation. This allows sound collection control to be performed according to the shooting state, such as portrait or landscape, making it easier to collect sound in line with the user's intentions.

[0147] In this embodiment, when the auto mode is set, the controller 135 controls the sound collection area so as to change the directionality of the sound capture unit in conjunction with the shooting state, i.e., whether or not the photographer is shooting himself / herself. This allows sound collection control to be performed according to the shooting state, such as whether or not the photographer is shooting a self-portrait, making it easier to collect sound in line with the user's intentions.

[0148] The digital camera 100 of this embodiment further includes a display monitor 130, which is an example of a display unit, and a magnetic sensor 132, which is an example of a detection unit. The display monitor 130 has a display surface that displays an image of a subject, etc., and is configured to be able to move the display surface toward the subject. The magnetic sensor 132 detects whether the display monitor 130 has moved the display surface toward the subject. The setting unit of this embodiment may set the digital camera 100 to auto mode when the magnetic sensor 132 detects that the display monitor 130 has moved the display surface toward the subject. For example, the controller 135 may automatically set the digital camera 100 to auto mode when the display monitor 130 is in the selfie position, in response to a detection signal from the magnetic sensor 132.

[0149] In the digital camera 100 of this embodiment, the setting unit can set the audio capture unit to at least one of a plurality of operation modes, each having a different directivity, in addition to the auto mode, in response to a user instruction. For example, the operation mode can be set to a surround mode, a front mode, or a navigation mode, and may also be set to a focus mode.

[0150] In the digital camera 100 of this embodiment, when the setting unit is set to the auto mode and image capturing is started by the image sensor 115, the controller 135 may cause the display monitor 130 to display information indicating that the auto mode is set along with the subject. For example, the controller 135 may cause the display monitor 130 to display an icon or the like dedicated to the auto mode.

[0151] (Embodiment 2) Hereinafter, a second embodiment will be described with reference to the drawings. In the first embodiment, a digital camera 100 that selects and determines a sound collection target when shooting a moving image or the like has been described. In the second embodiment, a digital camera 100 that visualizes information about the determined sound collection target to the user during operation similar to that of the first embodiment will be described.

[0152] Below, the digital camera 100 according to this embodiment will be described, omitting descriptions of the same configurations and operations as those of the digital camera 100 according to the first embodiment as appropriate.

[0153] [2-1. Overview] An overview of the operation of the digital camera 100 according to this embodiment to display various types of information will be described using FIG.

[0154] Fig. 22 shows a display example of the digital camera 100 according to this embodiment. The display example of Fig. 22 shows an example of what is displayed in real time on the display unit 130 when the digital camera 100 has determined a sound collection target as shown in Fig. 11(B) as an example. In this display example, the digital camera 100 displays on the display monitor 130 an AF frame 11 indicating the AF target subject and a detection frame 13 indicating a detected subject other than the AF target, as well as a sound collection icon 12 indicating the sound collection target subject, superimposed on the captured image Im.

[0155] The digital camera 100 of this embodiment uses the sound collection icon 12 in combination with the AF frame 11 and the detection frame 13 to visualize to the user whether a main subject such as an AF target and other detected subjects have been determined as AF targets and / or sound collection targets.

[0156] For example, in the display example of Fig. 22, the digital camera 100 has determined that the subject corresponding to face area R1 (60) in the example of Fig. 11(B) is to be an AF target and a sound collection target, and therefore displays an AF frame 11 and a sound collection icon 12 for person 21. Furthermore, the digital camera 100 has determined that the subject corresponding to face area R3 in the example of Fig. 11(B) is to be a sound collection target other than an AF target, and therefore displays a detection frame 13 and a sound collection icon 12 for person 23. Furthermore, by displaying the detection frame 13 without the sound collection icon 12, the digital camera 100 visualizes to the user that the subject other than an AF target corresponding to face area R2 in the example of Fig. 11(B) has been determined not to be a sound collection target.

[0157] In the digital camera 100 of this embodiment, the user can confirm whether or not a detected subject is an AF target by the display of either the AF frame 11 or the detection frame 13. The user can also confirm whether or not the subject is a sound collection target by the presence or absence of the sound collection icon 12. The combination of the AF frame 11 and the sound collection icon 12 is an example of first identification information in this embodiment. The combination of the detection frame 13 and the sound collection icon 12 is an example of second identification information in this embodiment. The detection frame 13 is an example of third identification information.

[0158] As described above, the digital camera 100 according to this embodiment displays a distinction between the sound collection target and the AF target determined from the subjects included in the detection information. This allows the user to identify the sound collection target subjects among the subjects detected by the digital camera 100 and to check, for example, whether the intended subject has been determined as the sound collection target.

[0159] 23(A) and (B) are diagrams illustrating manual operations in the digital camera 100 of this embodiment. FIG. 23(A) illustrates a state in which a specific person 22 has not been detected by the face recognition unit 122. For example, even if the photographer wants to collect the voice of the person 22, a case can be assumed in which face recognition is not performed because the face of the person 22 is facing sideways or backwards relative to the digital camera 100. It can also be assumed that the sound source that the photographer wants to collect is not a person. Taking such cases into consideration, the digital camera 100 of this embodiment operates so that a manual operation can be input to manually set the sound collection area.

[0160] Fig. 23(B) illustrates an example of manual operation of the sound collection area in digital camera 100. For example, manual operation of the sound collection area is implemented as a touch operation. Fig. 24 is a flowchart illustrating an example of operation during manual operation in digital camera 100. The processing shown in this flowchart may be executed independently of the processing in the auto mode or focus mode described above, or may be executed by interrupting the processing in the auto mode or focus mode.

[0161] First, the controller 135 of the digital camera 100 accepts a manual operation by a user such as a photographer (S61). During manual operation, the controller 135 displays a designated range 48 of the sound collection area on the display monitor 130, as shown in Fig. 23(B), for example. This example is a case where a manual operation is performed by interrupting processing in the auto mode or focus mode.

[0162] In the example of Fig. 23(B), the photographer inputs a manual operation by touch operation to adjust the designated range 48 of the sound collection area so that it includes the person 22 whose face was not recognized. The controller 135 determines the designated range 48 of the sound collection area based on the input manual operation (S62). If the start point and end point of the sound collection range are input by touch operation as a manual operation, the sound collection range 48 of a predetermined size including the start point and end point is displayed. In this state, if the confirm button displayed on the display monitor 130 is touched, the sound collection range 48 is confirmed.

[0163] The controller 135 reflects the determined designated range 48 of the sound collection area in the sound collection control of the microphone 161 (S63). As a result, the sound collection control is executed so as to emphasize the sound from the sound collection area corresponding to the designated range 48.

[0164] [2-3. Effects, etc.] As described above, digital camera 100 of this embodiment includes image sensor 115 that captures an image of a subject and generates image data, microphone 161 that acquires an audio signal indicating audio picked up during image capture by image sensor 115, and display monitor 130 that displays an image of the subject. Digital camera 100 of this embodiment also includes an input unit such as operation unit 150 that inputs a user operation to set the subject displayed on display monitor 130 as a sound collection area for collecting audio from the subject, and controller 135 that controls the sound collection area for audio signals. When a user operation to set the sound collection area is input, controller 135 controls the sound collection area by changing the directionality of microphone 161 based on the user operation so that the subject is included. This manual operation makes it possible to control the sound collection area, making it easier to collect audio as intended by the user.

[0165] In the digital camera 100 of this embodiment, the display monitor 130 may be configured to be able to shift the display surface toward the subject, as in the first embodiment. A user operation may be performed on the display monitor 130 in a state in which the display surface is shifted toward the subject. That is, when taking a selfie with the digital camera 100, the above-described manual operation may be input.

[0166] (Other embodiments) As described above, the above embodiments have been described as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these, and can be applied to embodiments in which appropriate modifications, substitutions, additions, omissions, etc. are made. Furthermore, it is also possible to combine the components described in the above embodiments to create new embodiments.

[0167] In the first embodiment, an example has been described in which three microphone elements 161L, 161C, and 161R are used in microphone 161. A modified example in which four microphone elements are used will be described with reference to FIGS.

[0168] 25 shows an example of the arrangement of microphone 161A in digital camera 100A of this modification. In this modification, microphone 161A of digital camera 100A includes a fourth microphone element 161B in addition to three microphone elements 161L, 161C, and 161R that are mutually on the XZ plane. Fourth microphone element 161B is arranged so that its position in the Y direction is different from that of the other microphone elements 161L to 161R.

[0169] FIG. 26 illustrates the relationship between the captured image Im and the sound collection areas 41A to 43A during vertical shooting in this modified example. According to the configuration of the microphone 161A described above, the sound collection area can also be changed in the Y-axis direction of the digital camera 100. Therefore, by using the sound collection area using the fourth microphone element 161B during sound collection control during vertical shooting, sound collection control that follows face recognition can be achieved even during vertical shooting, as shown in FIG. 26, for example. For example, the controller 135 of the digital camera 100A of this modified example performs sound collection control using the fourth microphone element 161B as described above, instead of step S55, in the same process as in FIG. 20. This enables intended sound collection control, such as using the sound collection areas 41A to 43A to align with the desired positions of the facial areas R1 to R3, even during vertical shooting.

[0170] Furthermore, in the sound collection control described in each of the above embodiments, by varying the speed of the transition of the sound collection area in conjunction with face recognition, it is possible to further reduce the sense of discomfort felt by the listener. An example of this operation will be described with reference to FIG.

[0171] 27 shows an example of operation in which the width of sound collection directivity (i.e., the angular range of the sound collection area) is changed in conjunction with the presence or absence of face recognition in digital camera 100. In this example of operation, the face of a subject is detected by face recognition unit 122 of digital camera 100 at time t1. Then, erroneous detection prevention control (i.e., chattering) is performed (see S33 in FIG. 16).

[0172] For example, if face recognition is performed only for a moment or if the recognized subject turns to the side, changing the sound collection directionality may cause the user listening to the collected sound to experience a sense of discomfort. In response to this, the above-mentioned erroneous detection prevention control constantly monitors the position of the recognized subject's face and changes the sound collection directionality when the subject is in the sound collection area for a certain period of time. This makes it possible to avoid the above-mentioned sense of discomfort.

[0173] Furthermore, after chattering from time t1, controller 135 of digital camera 100 transitions the sound collection area to narrow the sound collection directivity. The transition period when narrowing the sound collection directivity is set to be relatively short, for example. As a result, when a subject's face is detected in image recognition, the sound collection directivity is quickly changed to focus on the detected face, allowing a user listening to the collected sound to have a good auditory impression. Furthermore, transitions between sound collection areas of the same angle range, such as front center sound collection area 41 and left half sound collection area 42, are also performed relatively quickly in the same manner as above.

[0174] In the example of FIG. 27, for example, the subject moves or turns their face sideways, and therefore, at time t2, the face of the subject is no longer detected by the face recognition unit 122. In this case, the digital camera 100 also transitions the sound collection directivity after chattering control. Here, the transition period when widening the range of sound collection directivity is set longer than the transition period when narrowing the range. This makes it possible to avoid giving the user an unpleasant auditory sensation that may occur when sounds from a wide range suddenly start to be heard. Rather, by executing the transition when widening the range of sound collection directivity slowly, it is possible to suppress the unpleasant auditory sensation felt by the user.

[0175] Furthermore, in this example, face recognition is performed again at time t3 during the transition period in which the sound collection directivity is widened. In this case, digital camera 100 switches to interrupt control to narrow the sound collection directivity again before it is fully widened. This allows for quick control to orient the sound collection directivity toward the subject whose face has been recognized, even when face recognition of the subject is intermittent, further reducing the sense of discomfort felt by the user.

[0176] In each of the above embodiments, the face recognition unit 122 is used to detect the sound collection target. In this embodiment, the detection of the sound collection target is not limited to the face recognition unit 122, and for example, instead of or in addition to the face recognition unit 122, human body recognition that performs image recognition of the entire or at least a part of a human may be used. Furthermore, the sound collection target does not necessarily have to be a person, and may be, for example, various animals. In this case, the sound collection target may be detected by image recognition of a part or the entire animal.

[0177] Furthermore, the first identification information, second identification information, and third identification information according to the second embodiment identify whether or not a subject is a main subject based on the presence or absence of the AF frame 11, and identify whether or not a subject is a sound collection target based on the presence or absence of the sound collection icon 12. In this embodiment, the first to third identification information are not limited to this, and may be, for example, three types of frame displays. FIG. 22 illustrates three types of frame displays according to this embodiment. In the example of FIG. 22, the display of the AF target and other subjects and the display of the sound collection target are integrated by frame display 11A indicating a subject that is both an AF target and a sound collection target, frame display 13A indicating a subject that is not an AF target but a sound collection target, and frame display 13B indicating a subject that is not a sound collection target.

[0178] In the above first and second embodiments, the operation in auto mode and the manual operation have been described separately, but these may be combined. That is, the digital camera 100 of this embodiment may include a display monitor 135 as a display unit that displays an image of a subject, and an operation unit 150 as an input unit that inputs a user operation to set the subject displayed on the display monitor 135 as a sound collection area that collects sound from the subject. When a user operation to set the sound collection area is input, the controller 135 may control the sound collection area so as to include the subject by changing the directionality of the microphone 161 based on the user operation. In such a case, the display monitor 135 does not need to be movable, and may be fixed, for example, in the normal position described above.

[0179] In the first and second embodiments, an example of operation has been described in the flowchart of FIG. 10 in which sound collection control (S5 or S6) is performed with or without face recognition for the microphone 161 built into the digital camera 100. The digital camera 100 of the present embodiment may be provided with an external microphone (hereinafter referred to as "microphone 161a") instead of the built-in microphone 161. The microphone 161a includes microphone elements external to the digital camera 100 and has three or more microphone elements. In this embodiment, the controller 135 can execute step S5 or S6 in the same way as in the first embodiment by storing information about the arrangement of the microphone elements for the microphone 161a in advance in the buffer memory 125 or the like. In this case, too, it is possible to easily obtain the sound of the subject clearly according to the sound collection target and / or sound collection area determined in the same way as in the first embodiment.

[0180] In addition, in the first and second embodiments, an example of operation has been described in the flowchart of FIG. 16 in which gain Gh is calculated (S30) based on the horizontal angle of view corresponding to the imaging range of digital camera 100. The horizontal angle of view in this case is the same as the horizontal angle of view θh used in determining the front center sound collection area (S20) in the flowchart of FIG. 14. In this embodiment, a horizontal angle of view different from the horizontal angle of view θh used in step S20 may be used to calculate gain Gh. For example, the angle range corresponding to the width in the X-axis direction that includes all sound collection target subjects on the captured image is set as the horizontal angle of view in step S30. In this way, gain Gh can be calculated according to the angle of view at which the sound collection target is captured so as to more clearly collect voices of distant subjects.

[0181] In the first and second embodiments, the face recognition unit 122 detects a human face. In the present embodiment, the face recognition unit 122 may detect, for example, an animal face. It is conceivable that animal faces vary in size depending on the type of animal. Even in this case, it is possible to select sound collection targets in the same way as in the first embodiment, for example, by expanding the predetermined range (see S14) for selecting sound collection targets. Furthermore, the face recognition unit 122 may detect a face for each type of animal, and set the predetermined range in step S14 according to the type.

[0182] In addition, in the first and second embodiments, the digital camera 100 has been described as including the face recognition unit 122. In the present embodiment, the face recognition unit 122 may be provided in an external server. In this case, the digital camera 100 may transmit image data of a captured image to the external server via the communication module 160, and receive detection information of the processing result by the face recognition unit 122 from the external server. In such a digital camera 100, the communication module 160 functions as a detection unit.

[0183] Furthermore, in the first and second embodiments, the digital camera 100 is illustrated as including the optical system 110 and the lens driving unit 112. The imaging device of this embodiment does not have to include the optical system 110 and the lens driving unit 112, and may be, for example, an interchangeable lens camera.

[0184] Although a digital camera has been described as an example of an imaging device in the first and second embodiments, the imaging device is not limited to this. The imaging device of the present disclosure may be any electronic device having an image capturing function (for example, a video camera, a smartphone, a tablet terminal, etc.).

[0185] As described above, the embodiments have been described as examples of the technology in the present disclosure, and for that purpose, the accompanying drawings and detailed description have been provided.

[0186] Therefore, the components shown in the accompanying drawings and detailed description may include not only essential components for solving the problem, but also components that are not essential for solving the problem in order to illustrate the above technology. Therefore, the fact that these non-essential components are shown in the accompanying drawings or detailed description should not be interpreted as immediately indicating that these non-essential components are essential.

[0187] Furthermore, since the above-described embodiments are intended to illustrate the technology of the present disclosure, various modifications, substitutions, additions, omissions, etc. may be made within the scope of the claims or their equivalents. [Industrial Applicability]

[0188] The present disclosure is applicable to an imaging device that captures images while acquiring audio. [Explanation of symbols]

[0189] 100 digital cameras 115 Image Sensor 120 Image Processing Engine 122 Face Recognition Unit 125 buffer memory 130 Display Monitor 132 Magnetic Sensor 135 Controller 137 Acceleration Sensor 145 Flash Memory 150 Operation section 11 AF frame 12 Audio pickup icon 13 Detection Frame

Claims

1. an imaging unit that captures an image of a subject and generates image data; an audio acquisition unit that acquires an audio signal indicating audio picked up during imaging by the imaging unit; a setting unit that receives an instruction from a user and sets the device to an auto mode, which is an operation mode in which the directivity of the voice acquisition unit is automatically changed; a control unit that controls a sound collection area for collecting the sound from the subject in the audio signal; Equipped with When the setting unit is set to the auto mode, the control unit controls the sound collection area to include the subject by changing the directionality of the sound acquisition unit in conjunction with the shooting state of the device.

2. a face detection unit that detects a face area of ​​the subject in the image data, 2. The imaging device of claim 1, wherein when the setting unit is set to the auto mode, the control unit determines a subject to be collected in the audio signal based on the face area detected by the face detection unit, and controls the sound collection area so that the determined subject is included in the sound collection target.

3. The imaging device of claim 1, wherein when the setting unit is set to the auto mode, the control unit controls the sound collection area so as to change the directivity of the sound acquisition unit in accordance with the shooting state of the device, whether the device is in portrait or landscape orientation.

4. The imaging device according to claim 1, wherein when the setting unit is set to the auto mode, the control unit controls the sound collection area so as to change the directivity of the sound acquisition unit in conjunction with a shooting state of whether or not the photographer is shooting himself / herself.

5. a display unit having a display surface for displaying an image of the subject, the display surface being movable toward the subject; a detection unit that detects whether the display unit has displaced the display surface toward the subject; Furthermore, The imaging device according to claim 1 , wherein the setting unit sets the imaging device to the auto mode when the detection unit detects that the display surface of the display unit has been displaced toward the subject.

6. The imaging device according to claim 1 , wherein the setting unit is capable of setting the sound acquisition unit to at least one of a plurality of operation modes, each having a different directivity from the other, in addition to the auto mode, in response to an instruction from the user.

7. a display unit that displays an image of the subject; an input unit for inputting a user operation to set the subject displayed on the display unit as a sound collection area for collecting sound from the subject; Furthermore, The imaging device according to claim 1 , wherein when a user operation to set the sound collection area is input, the control unit controls the sound collection area so as to include the subject by changing the directivity of the sound acquisition unit based on the user operation.

8. an imaging unit that captures an image of a subject and generates image data; an audio acquisition unit that acquires an audio signal indicating audio picked up during imaging by the imaging unit; a display unit that displays an image of the subject; an input unit for inputting a user operation to set the subject displayed on the display unit as a sound collection area for collecting sound from the subject; a control unit that controls the sound collection area for the audio signal, When a user operation to set the sound collection area is input, the control unit controls the sound collection area so as to include the subject by changing the directivity of the sound acquisition unit based on the user operation.

9. the display unit is configured to be able to displace a display surface toward the subject, The imaging device according to claim 8 , wherein the user operates the display unit in a state in which the display surface is displaced toward the subject.

Citation Information

Patent Citations

  • Picture voice recording apparatus and collecting voice direction adjustment method

    JP2006287735A

  • Camera, reproducing device, and reproducing method

    JP2010093603A

  • Imaging apparatus

    JP2010200253A

  • Imaging device

    JP2011114769A

  • Imaging device

    JP2016189584A