Imaging device

The imaging device adjusts sound pickup area automatically based on user instructions and face recognition, improving sound collection accuracy and user experience.

JP7829159B2Active Publication Date: 2026-03-13PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing imaging devices struggle to easily collect the voice of a subject in accordance with the user's intentions while capturing images and sound.

Method used

The imaging device includes an audio acquisition unit that sets the device to auto mode, automatically adjusting the sound pickup area based on user instructions and the shooting state, using face recognition to control the directivity of the audio acquisition unit.

Benefits of technology

This allows for easier capture of the subject's sound in accordance with the user's intentions, enhancing sound collection accuracy and reducing auditory annoyance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007829159000001
    Figure 0007829159000001
  • Figure 0007829159000002
    Figure 0007829159000002
  • Figure 0007829159000003
    Figure 0007829159000003
Patent Text Reader

Abstract

To provide an imaging apparatus that takes an image while acquiring voice, the apparatus making it easier to collect voice of an object according to an intention of a user.SOLUTION: The imaging apparatus includes: an imaging unit for imaging an object and generating image data; a voice acquisition unit for acquiring a voice signal showing voice collected during imaging of the imaging unit; a setting unit for setting an own device in an auto-mode as an operation mode for automatically changing the directivity of the voice acquisition unit in response to an instruction of a user; and a control unit for controlling a voice collecting area for collecting voice from an object in the voice signal. The control unit controls the voice collecting area so that the area will include the object, by changing the directivity of the voice acquisition unit in association with the imaging state of the own device when the setting unit sets the own device in the auto-mode.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an imaging device that performs imaging while acquiring sound.

Background Art

[0002] Patent Document 1 discloses a video camera having a face detection function. The video camera of Patent Document 1 changes the directivity angle of a microphone according to the zoom ratio and the size of a person's face in the captured screen. Thereby, the video camera controls the directivity angle of the microphone in association with the distance between the video camera and the subject video, and changes the directivity angle of the microphone so as to more reliably capture the voice of the subject while achieving matching between the video and the sound. At this time, the video camera detects the position and size of the face of a person (subject), attaches a frame (face detection frame) to the detected face portion and displays it, and uses the information on the size of the face detection frame (the size of the face).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present disclosure provides an imaging device that can easily collect the voice of a subject in accordance with the intention of a user in an imaging device that performs imaging while acquiring sound.

Means for Solving the Problems

[0005] In this disclosure, the imaging device comprises an imaging unit that images a subject and generates image data, an audio acquisition unit that acquires an audio signal indicating sound picked up during imaging by the imaging unit, a setting unit that sets the device to auto mode, which is an operating mode that automatically changes the directivity of the audio acquisition unit in response to user instructions, and a control unit that controls the sound pickup area for picking up sound from the subject in the audio signal. When the setting unit is set to auto mode, the control unit controls the sound pickup area to include the subject by changing the directivity of the audio acquisition unit in conjunction with the shooting state of the device. [Effects of the Invention]

[0006] According to the imaging device described herein, in an imaging device that captures images while acquiring sound, it is possible to make it easier to capture the sound of the subject in accordance with the user's intentions. [Brief explanation of the drawing]

[0007] [Figure 1] This figure shows the configuration of the digital camera 100 according to Embodiment 1 of this disclosure. [Figure 2] Diagram illustrating the back of digital camera 100. [Figure 3] A diagram illustrating the state of a digital camera 100 when taking a selfie. [Figure 4] A diagram illustrating the state of digital camera 100 when shooting vertically. [Figure 5] A diagram illustrating the configuration of the beamforming unit 172 in the digital camera 100. [Figure 6] Diagram illustrating the sound pickup area in digital camera 100. [Figure 7] This diagram shows an example of the display of the settings menu in digital camera 100. [Figure 8] Diagram illustrating an additional sound pickup area in digital camera 100. [Figure 9] A diagram illustrating the operation of the auto mode on the digital camera 100. [Figure 10] A flowchart illustrating the operation of the focus mode of the digital camera 100 according to Embodiment 1. [Figure 11] Figure for explaining the outline of the operation of the focus mode of the digital camera 100 [Figure 12] Flowchart exemplifying the selection process of the sound collection target (S3 in FIG. 10) of the digital camera 100 according to Embodiment 1 [Figure 13] Figure for explaining the selection process of the sound collection target in the digital camera 100 [Figure 14] Flowchart exemplifying the determination process of the sound collection area (S4 in FIG. 10) in the digital camera 100 [Figure 15] Figure for explaining the determination process of the sound collection area in the digital camera 100 [Figure 16] Flowchart exemplifying the sound collection control using face recognition (S5 in FIG. 10) in the digital camera 100 [Figure 17] Figure for explaining the management information obtained by the determination process of the sound collection area [Figure 18] Figure exemplifying the relationship for obtaining the gain from the horizontal angle of view and the focal length in the digital camera 100 [Figure 19] Flowchart exemplifying the sound collection control without using face recognition (S6 in FIG. 10) in the digital camera 100 [Figure 20] Flowchart exemplifying the operation of the auto mode of the digital camera 100 according to Embodiment 1 [Figure 21] Figure for explaining the sound collection control during horizontal shooting and vertical shooting in the auto mode of Embodiment 1 [Figure 22] Figure showing an example of display of the digital camera 100 according to Embodiment 2 [Figure 23] Figure for explaining the manual operation in the digital camera 100 of Embodiment 2 [Figure 24] Flowchart exemplifying the operation during manual operation in the digital camera 100 of Embodiment 2 [Figure 25] Figure exemplifying the arrangement of the microphone 161A in the digital camera 100A of the modification [Figure 26]A diagram for explaining the sound collection control during vertical shooting in the auto mode of the modified example [Figure 27] A diagram for explaining an operation example of sound collection control linked to face recognition in the digital camera 100

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments will be described in detail with reference to the drawings as appropriate. However, a more detailed description than necessary may be omitted. For example, a detailed description of well-known matters and a redundant description of substantially the same configuration may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art. The inventors provide the accompanying drawings and the following description so that those skilled in the art can fully understand the present disclosure, and do not intend to limit the subject matter described in the claims thereby.

[0009] (Embodiment 1) In Embodiment 1, as an example of the imaging device according to the present disclosure, a digital camera that detects a subject based on image recognition technology, controls the sound collection area according to the size of the detected subject, and controls the sound collection gain to emphasize the collected sound will be described.

[0010] [1-1. Configuration]

[0011] Figure 1 shows the configuration of the digital camera 100 according to this embodiment. The digital camera 100 of this embodiment includes an image sensor 115, an image processing engine 120, a display monitor 130, and a controller 135. Furthermore, the digital camera 100 includes a buffer memory 125, a card slot 140, a flash memory 145, an operation unit 150, and a communication module 160. The digital camera 100 also includes a microphone 161, an analog-to-digital (A / D) converter 165 for the microphone, and an audio processing engine 170. The digital camera 100 also includes, for example, an optical system 110 and a lens drive unit 112. Furthermore, the digital camera 100 includes, for example, a magnetic sensor 132 and an acceleration sensor 137.

[0012] Figure 2 illustrates the back of the digital camera 100. In Figure 2, the three axes X, Y, and Z of the digital camera 100 are shown along with the direction of gravity G. The X, Y, and Z axes correspond to the horizontal field of view, vertical field of view, and optical axis of the lens in the optical system 110, respectively. In the example in Figure 2, the Y axis of the digital camera 100 is aligned with the direction of gravity G, i.e., sideways.

[0013] The digital camera 100 of this embodiment can be used by the user to take selfies or to take vertical photos using the digital camera 100 in a vertical orientation. Figure 3 illustrates the state of the digital camera 100 when taking a selfie. Figure 4 illustrates the state of the digital camera 100 when taking vertical photos.

[0014] Returning to Figure 1, the optical system 110 includes a focusing lens, a zoom lens, an optical image stabilization (OIS) lens, an aperture, a shutter, etc. The focusing lens is a lens for changing the focus state of the subject image formed on the image sensor 115. The zoom lens is a lens for changing the magnification of the subject image formed by the optical system. The focusing lens, etc., are each composed of one or more lenses.

[0015] The lens drive unit 112 drives the focus lens and other components in the optical system 110. The lens drive unit 112 includes a motor and moves the focus lens along the optical axis of the optical system 110 based on the control of the controller 135. The configuration for driving the focus lens in the lens drive unit 112 can be implemented using a DC motor, stepping motor, servo motor, or ultrasonic motor, etc.

[0016] The image sensor 115 captures an image of a subject formed through the optical system 110 and generates imaging data. The imaging data constitutes image data representing the image captured by the image sensor 115. The image sensor 115 generates image data of a new frame at a predetermined frame rate (e.g., 30 frames / second). The timing of image data generation and the operation of the electronic shutter in the image sensor 115 are controlled by the controller 135. The image sensor 115 can use various image sensors, such as a CMOS image sensor, a CCD image sensor, or an NMOS image sensor.

[0017] The image sensor 115 performs operations such as capturing moving images, still images, and through images. Through images are mainly moving images and are displayed on the display monitor 130 for the user to determine the composition for, for example, still image capture. The through image, moving image, and still image are examples of captured images in this embodiment. The image sensor 115 is an example of an imaging unit in this embodiment.

[0018] The image processing engine 120 performs various processes on the imaging data output from the image sensor 115 to generate image data, and also performs various processes on the image data to generate an image for display on the display monitor 130. Examples of various processes include, but are not limited to, white balance correction, gamma correction, YC conversion, electronic zoom, compression, and decompression. The image processing engine 120 may be composed of hardwired electronic circuits, or it may be composed of a microcomputer or processor using a program.

[0019] In this embodiment, the image processing engine 120 includes a face recognition unit 122 that performs a function to detect subjects such as human faces by image recognition of the captured image. The face recognition unit 122 performs face detection, for example, by rule-based image recognition processing and outputs detection information. Face detection may be performed by various image recognition algorithms. The detection information includes position information corresponding to the detection result of the subject. The position information is defined, for example, by the horizontal and vertical positions on the image Im to be processed, and indicates, for example, a rectangular area surrounding a human face as the detected subject (see Figure 11).

[0020] The display monitor 130 is an example of a display unit that displays various information. For example, the display monitor 130 displays an image (through image) represented by image data captured by the image sensor 115 and processed by the image processing engine 120. The display monitor 130 also displays a menu screen or the like for the user to make various settings for the digital camera 100. The display monitor 130 can be made of, for example, a liquid crystal display device or an organic EL device.

[0021] The digital camera 100 of this embodiment is configured to have a movable display monitor 130 whose position can be changed, as shown in Figures 2 and 3, for example. In the example in Figure 2, the display monitor 130 is positioned with its display surface facing the back side (-Z side) of the digital camera 100. This position of the display monitor 130 will be referred to below as the "normal position". In the example in Figure 3, the display monitor 130 is positioned with its display surface facing the front side (+Z side) of the digital camera 100, i.e., the subject side. This position of the display monitor 130 will be referred to below as the "selfie position".

[0022] The magnetic sensor 132 is an example of a detection unit that detects whether the display monitor 130 is in its normal position or in the selfie position. The magnetic sensor 132 outputs a detection signal to the controller 135 that indicates the detection result of the position of the display monitor 130.

[0023] The movable display monitor 130 can be, for example, a vari-angle type or a tilt type. For example, a hinge 131 is provided to rotatably connect the display monitor 130 to the body of the digital camera 100. The magnetic sensor 132 is provided, for example, inside the hinge 131 and consists of a switch or the like having two states corresponding to Figures 2 and 3.

[0024] The acceleration sensor 137 detects one or more accelerations in the three axes X, Y, and Z, for example, and outputs a detection signal to the controller 135. The acceleration sensor 137 is an example of an attitude detection unit that detects whether the orientation of the digital camera 100 is horizontal, as illustrated in Figure 2, or vertical, as illustrated in Figure 4, based on the detected state of gravitational acceleration.

[0025] The operation unit 150 is a general term for hard keys such as operation buttons and operation levers provided on the exterior of the digital camera 100, and accepts operations from the user. The operation unit 150 includes, for example, a shutter release button, a mode dial, a touch panel, cursor buttons, and a joystick. When the operation unit 150 accepts an operation from the user, it transmits an operation signal corresponding to the user operation to the controller 135. The operation unit 150 includes, for example, a shutter release button 151, a selection button 152, a confirmation button 153, a function button 154, and a touch panel 155, as shown in Figure 2.

[0026] The controller 135 provides overall control over the operation of the digital camera 100. The controller 135 includes a CPU, and the CPU executes programs (software) to realize predetermined functions. Instead of a CPU, the controller 135 may include a processor consisting of dedicated electronic circuits designed to realize predetermined functions. In other words, the controller 135 can be realized with various processors such as a CPU, MPU, GPU, DSU, FPGA, and ASIC. The controller 135 may consist of one or more processors. Alternatively, the controller 135 may be configured on a single semiconductor chip together with the image processing engine 120, etc.

[0027] The buffer memory 125 is a recording medium that functions as work memory for the image processing engine 120 and the controller 135. The buffer memory 125 is implemented using DRAM (Dynamic Random Access Memory) or the like. The flash memory 145 is a non-volatile recording medium. Although not shown in the diagram, the controller 135 may also have various internal memories, such as built-in ROM. The ROM stores various programs that the controller 135 executes. The controller 135 may also have built-in RAM that functions as a work area for the CPU.

[0028] The card slot 140 is a means for inserting a removable memory card 142. The card slot 140 can electrically and mechanically connect to the memory card 142. The memory card 142 is an external memory equipped with recording elements such as flash memory. The memory card 142 can store data such as image data generated by the image processing engine 120.

[0029] The communication module 160 is a communication module (circuit) that performs communication compliant with the IEEE 802.11 or Wi-Fi standard, etc. The digital camera 100 can communicate with other devices via the communication module 160. The digital camera 100 may communicate directly with other devices via the communication module 160, or it may communicate via an access point. The communication module 160 may be connectable to a communication network such as the Internet.

[0030] Microphone 161 is an example of a sound-collecting unit. Microphone 161 converts the collected sound into an analog signal, which is an electrical signal, and outputs it. Microphone 161 in this embodiment includes three microphone elements 161L, 161C, and 161R. Microphone 161 may be composed of two or more microphone elements.

[0031] The A / D converter 165 for the microphone converts the analog signal from the microphone 161 into digital audio data. The A / D converter 165 for the microphone is an example of the audio acquisition unit in this embodiment. The microphone 161 may include a microphone element located outside the digital camera 100. In this case, the digital camera 100, as the audio acquisition unit, includes an interface circuit for the external microphone 161.

[0032] The audio processing engine 170 receives audio data output from an audio acquisition unit such as an A / D converter 165 for a microphone, and performs various audio processing on the received audio data. The audio processing engine 170 is an example of an audio processing unit in this embodiment.

[0033] The audio processing engine 170 of this embodiment includes, for example, a beamforming unit 172 and a gain adjustment unit 174, as shown in Figure 1. The beamforming unit 172 implements a function to control the directivity of the sound. Details of the beamforming unit 172 will be described later. The gain adjustment unit 174 amplifies the sound by performing a multiplication process on the input audio data, for example, by a sound pickup gain set by the controller 135. The gain adjustment unit 174 may also perform a process to suppress the sound by multiplying the input audio data by a negative gain. The sound pickup gain adjustment unit 14 may further have a function to change the frequency characteristics and stereo characteristics of the input audio data. Details on setting the sound pickup gain will be described later.

[0034] [1-1-1. About the beam forming section] Details of the beamforming unit 172 in this embodiment will be described below.

[0035] The beamforming unit 172 performs beamforming to control the directivity of the sound picked up by the microphone 161. An example of the configuration of the beamforming unit 172 in this embodiment is shown in Figure 5.

[0036] As shown in Figure 5, the beamforming unit 172 includes, for example, filters D1 to D3 and an adder 173, and adjusts the delay period of the sound picked up by each microphone element 161L, 161C, and 161R to output a weighted sum. The beamforming unit 172 controls the direction and range of the sound pickup directivity of the microphone 161, thereby setting the physical range in which the microphone 161 picks up sound.

[0037] The beamforming unit 172, as shown in the figure, outputs one channel using one adder 173, but it may also be configured with two or more adders to produce different outputs for each channel, such as stereo output. In addition, subtractors may be used in addition to the adder 173 to form directivity with a blind spot in a specific direction, which is a direction with particularly low sensitivity, or adaptive beamforming may be performed, which changes the processing according to the environment. Furthermore, different processing may be applied depending on the frequency band of the audio signal.

[0038] Figure 5 shows an example in which microphone elements 161L, 161C, and 161R are arranged in a straight line, but the arrangement of each microphone element is not limited to this. For example, even when they are arranged in a triangular shape, the sound pickup directivity of the microphone 161 can be controlled by appropriately adjusting the delay period and weights of filters D1 to D3. Furthermore, the beamforming unit 172 may apply known methods to control the sound pickup directivity. For example, sound processing technology such as OZO Audio may be used to perform processing to form directivity, and at the same time, processing to suppress sound noise may be performed.

[0039] The sound pickup area of ​​the digital camera 100, which can be set by the beamforming unit 172 described above, will be explained.

[0040] [1-1-2. Regarding the sound pickup area] Figure 6 shows an example of the sound pickup area defined in the digital camera 100. Figure 6 illustrates the sound pickup area as a sector-shaped region of a circle centered on the digital camera 100. In the digital camera 100 of this embodiment, the horizontal field of view direction coincides with the direction in which the microphone elements 161R, 161C, and 161R are aligned.

[0041] Figure 6(A) shows a "front central sound pickup area" 41 that faces forward (i.e., in the shooting direction) of the digital camera 100 within an angular range of 401 (e.g., 70°). Figure 6(B) shows a "left half sound pickup area" 42 that faces left of the digital camera 100 within an angular range of 401. Figure 6(C) shows a "right half sound pickup area" 43 that faces right of the digital camera 100 within an angular range of 401. Figure 6(D) shows a "front sound pickup area" 44 that faces forward of the digital camera 100 within an angular range of 402 (e.g., 160°), which is greater than the angular range of 401. These sound pickup areas are examples of a plurality of predetermined areas in this embodiment, and angular ranges 401 and 402 are examples of the first and second angular ranges.

[0042] In this embodiment, the digital camera 100 uses the front central sound pickup area 41 shown in Figure 6(A) when the subject is located in the center of the captured image. When the subject is located in the left half of the captured image, the left half sound pickup area 42 shown in Figure 6(B) is used, and when the subject is located in the right half of the captured image, the right half sound pickup area 43 shown in Figure 6(C) is used. Furthermore, when the subject is located in the entirety of the captured image, the front sound pickup area 44 shown in Figure 6(D) is mainly used.

[0043] In the example in Figure 11(B), the subjects R1 and R3 to be sound-recorded are located in the center of the captured image, so the front central sound-recording area 41 is used. In the example in Figure 11(C), the subjects R1 and R2 to be sound-recorded are located in the left half of the captured image, so the left half sound-recording area 42 is used.

[0044] In the sound-collecting section of an imaging device such as the microphone 161 of a digital camera 100, the number and arrangement of microphone elements are constrained by factors such as the mounting space for the elements. For example, in a shooting scenario where a user wants to record audio from multiple subjects, the number of microphone elements may be limited, making it difficult to sufficiently narrow the sound-collecting directivity. Even in such cases, the digital camera 100 of this embodiment can provide a sound-collecting area that aligns with the user's intentions by defining a sound-collecting area based on the user's shooting scenario and determining the sound-collecting area using face recognition.

[0045] [1-1-3. Microphone Settings] The settings for the sound pickup area of ​​the microphone 161 in the digital camera 100 will be explained using Figures 7 and 8.

[0046] Figure 7 shows an example of the display of the settings menu in the digital camera 100. The digital camera 100 of this embodiment has operating modes (i.e., sound pickup modes) for controlling the sound pickup area of ​​the microphone 161, such as "Auto," "Surround," "Front," "Focus," and "Narration," as shown in Figure 7.

[0047] The focus mode is an operating mode that automatically changes the directivity of the microphone 161 and adjusts the sound pickup area in conjunction with face recognition and the field of view by the digital camera 100. For example, the focus mode can be achieved by switching between the various sound pickup areas 41 to 44 described above. By deliberately implementing the focus mode in a broad manner using the four sound pickup areas 41 to 44, it is possible to avoid situations where the sound pickup directivity changes frequently due to slight movements of the subject, thereby reducing the auditory annoyance for the user. Further sound pickup areas in the digital camera 100 are illustrated in Figure 8.

[0048] Figure 8(A) illustrates the sound pickup area 45 in surround mode. Surround mode is an operating mode for capturing sound from a wide range of directions, including the left, right, front, and back of the digital camera 100. The sound pickup area 45 in surround mode has an angular range such as 360° around the entire XZ plane.

[0049] Figure 8(B) illustrates the sound pickup area 46 in front mode. Front mode is an operating mode for picking up sound from the front of the digital camera 100. The sound pickup area 46 in front mode is oriented towards the +Z side from the digital camera 100 and has an angular range greater than, for example, the front sound pickup area 44 described above.

[0050] Figure 8(C) illustrates the sound pickup area 47 in narration mode. Narration mode is an operating mode for picking up sound from behind the digital camera 100. The sound pickup area 47 in narration mode is formed facing the -Z side from the digital camera 100. When narration mode is set and the display monitor 130 is detected to be in the selfie position, the digital camera 100 may perform the operation of focus mode.

[0051] Auto mode is an operating mode that automatically changes the directivity of the microphone 161 and adjusts the sound pickup area in conjunction with the shooting state of the digital camera 100. The shooting state of the digital camera 100 includes, for example, face recognition considered in focus mode, as well as whether it is a selfie or not, and whether it is a vertical or horizontal shot.

[0052] The microphone settings, which allow for the selection of various sound pickup modes as described above, are provided, for example, as one of the video menus in the settings menu of the digital camera 100. The user can select the desired sound pickup mode from the settings menu by touching the touch panel 155 or by pressing various buttons 152 and 153. The microphone settings may also be pre-assigned to function buttons 154, etc. The setting of a specific sound pickup mode, such as auto mode, may also be assigned to function buttons 154, etc.

[0053] [1-2. Operation] The operation of the digital camera 100, configured as described above, will now be explained. Below, the operation of the digital camera 100 during video recording will be described.

[0054] The digital camera 100 sequentially captures the subject image formed through the optical system 110 with the image sensor 115 and generates imaging data. The image processing engine 120 performs various processes on the imaging data generated by the image sensor 115 to generate image data and records it in the buffer memory 125. The face recognition unit 122 of the image processing engine 120 detects the area of ​​the subject based on the image shown in the imaging data and outputs detection information, for example, to the controller 135.

[0055] The digital camera 100 of this embodiment includes a face recognition mode, which is an operating mode in which the face recognition unit 122 detects faces in the captured image through image recognition processing and identifies the subject to be controlled by autofocus (AF) based on the detection information.

[0056] Simultaneously with the imaging operation described above, the digital camera 100 captures sound using the microphone 161. The audio data obtained from the microphone's A / D converter 165 is processed by the audio processing engine 170. The audio processing engine 170 records the processed audio data Aout in the buffer memory 125.

[0057] The controller 135 synchronizes the image data received from the image processing engine 120 and the audio data received from the audio processing engine 170 via the buffer memory 125 to record the video onto the memory card 142. The controller 135 also sequentially displays a pass-through image on the display monitor 130. The user can check the composition of the shot at any time using the pass-through image on the display monitor 130. The video recording operation is started and stopped according to the user's operation on the control unit 150.

[0058] In video recording with the Digital Camera 100, various situations can arise where the user wants to capture sound. For example, the focus may be on a group of subjects, such as the photographer and their companions, conversing amongst themselves. In this case, there is a need to clearly capture the voices of the group of subjects.

[0059] In this embodiment, the digital camera 100, for example, as part of the operation of the focus mode described above, detects a subject using detection information from the face recognition unit 122 in the image processing engine 120. When the subject to be AF is determined, the audio processing engine 170 performs processing to enhance the sound to be captured of that subject and subjects surrounding that subject in the shooting space. In this way, by linking the face recognition of the image processing engine 120 with the sound enhancement of the audio processing engine 170, accurate sound capture with enhanced sound from a group of subjects having a conversation as described above is achieved.

[0060] Furthermore, in addition to the operation of the focus mode described above, the digital camera 100 of this embodiment also implements appropriate sound capture control according to various shooting conditions as part of its auto mode operation. The overview of the auto mode operation will be explained using Figure 9.

[0061] Figure 9 illustrates the correspondence between the auto mode and various sound recording modes of the digital camera 100. In auto mode, for example, when face recognition is performed while shooting horizontally, the digital camera 100 operates in the same way as the focus mode.

[0062] On the other hand, when face recognition is not performed, for example, when not taking a selfie (see Figure 2), the digital camera 100 operates in the same way as in surround mode, that is, it adopts the surround mode sound pickup area 45. Also, when face recognition is not performed and a selfie is being taken (see Figure 3), the digital camera 100 operates in the same way as in front mode.

[0063] Furthermore, in the case of vertical shooting, the operation when face recognition is not performed is the same as in the case of horizontal shooting described above. On the other hand, when face recognition is performed in vertical shooting, the digital camera 100 of this embodiment performs the same operation as in front mode instead of focus mode.

[0064] As shown in Figure 9, the operation of the auto mode described above makes it easier to achieve appropriate sound pickup control in each shooting condition by combining the operation of various sound pickup modes according to various shooting conditions.

[0065] [1-2-1. Operation of Focus Mode] Figures 10 and 11 will be used to illustrate the general operation of the focus mode of the digital camera 100 according to this embodiment.

[0066] Figure 10 is a flowchart illustrating the operation of the focus mode of the digital camera 100 according to this embodiment. Each process shown in the flowchart of Figure 10 is repeatedly executed at a predetermined interval when the digital camera 100 is set to focus mode. The predetermined interval is, for example, the frame period of a video. Figure 11 is a diagram illustrating the overview of the operation of the focus mode of the digital camera 100 according to this embodiment.

[0067] The controller 135 identifies the AF target based on the detection information from the face recognition unit 122 and performs AF control (S1). The AF target indicates the area of ​​the subject to be AF controlled in the image. Figure 11(A) illustrates an image Im that includes face regions R1, R2, and R3, which indicate the areas where the subject was detected in the detection information from the face recognition unit 122. Face regions R1, R2, and R3 are examples of subject areas in this embodiment. For example, face region R1 is identified as the face region 60 that is the AF target.

[0068] Next, the controller 135 determines whether or not a face region identified as an AF target exists (S2). Specifically, the controller 135 determines whether or not a face region has been detected and whether or not the AF target is a face region.

[0069] If there is a face region 60 that is an AF target (YES in S2), the controller 135 performs the process of selecting the subject to be picked up by the microphone 161 from the detected subject (S3). The subject to be picked up is the subject whose sound will be amplified and picked up by the microphone 161. Face region R1(60) identified as an AF target becomes a subject to be picked up. Figure 11(B) shows an example in which face regions R1 and R3 are decided to be subjects to be picked up, while face region R2 is not, based on the detection information shown in Figure 11(A).

[0070] In this embodiment, the digital camera 100, during the sound recording target selection process (S3), determines, in addition to the face region R1 (60) targeted for AF, to also select face R3, which shows a face size similar to face region R1 in the captured image Im, as an additional sound recording target. On the other hand, face region R2, which is of a different size than face region R1, is excluded from the sound recording target. This allows the camera to capture a group of subjects, for example, a group of people conversing among themselves, reflecting that person 21 and person 23 are at roughly the same distance from the digital camera 100 (i.e., the difference in distance in the Z-axis direction is small), while person 22 is at a different distance. Details of the sound recording target selection process (S3) will be described later.

[0071] Next, the controller 135 performs a process to determine the sound collection area based on the determined sound collection targets (S4). The sound collection area determination process (S4) determines a sound collection area that includes all the determined sound collection targets. In the example in Figure 11(B), the sound collection area is determined to be the front central sound collection area 41 (Figure 6(A)) which includes the face regions R1 and R3 of the sound collection targets. Details of the sound collection area determination process (S4) will be described later.

[0072] Next, the controller 135 controls sound collection using face recognition based on the determined sound collection target and sound collection area (S5). Sound collection control using face recognition (S5) is performed by setting the sound collection parameters, including the sound collection target, sound collection area, and sound collection gain determined by the controller 135, to the voice processing engine 170. The voice processing engine 170 realizes the sound collection directivity and sound collection gain according to the sound collection parameters.

[0073] On the other hand, if there is no face region 60 to be AF targeted, for example, when a face region is not detected during operation in face recognition mode (NO in S2), the controller 135 performs sound pickup control without face recognition (S6). Details of sound pickup control with and without face recognition (S5, S6) will be described later.

[0074] After executing the sound collection control in step S5 or S6, the controller 135 repeats the processing from step S1 onward.

[0075] According to the above process, the digital camera 100 of this embodiment selects subjects to be sound-collected from those detected by face recognition, determines a sound-collecting area that includes all of the subjects to be sound-collected, and performs sound-collecting control using face recognition. This makes it possible to emphasize and collect the voices of, for example, a group of subjects having a conversation among themselves.

[0076] In addition, in AF control by face recognition (S1), the identification of the AF target based on the detection information can be performed, for example, by displaying a frame indicating the face area on the through image displayed on the display monitor 130, and receiving an operation from the user via the operation unit 150 to select the display of the frame.

[0077] Figure 11(C) shows an example of an captured image Im when people 21-23 are in different positions than in Figures 11(A) and (B). Similar to the example in Figure 11(B), the digital camera 100 first identifies, for example, face region R1 as the face region 60 to be AF-targeted (S1) and decides it to be the sound-collecting target. In the example in Figure 11(C), the sound-collecting target selection process (S3) determines face region R2, which is about the same size as face region R1 on the captured image Im, to be the sound-collecting target, and excludes face region R3 from the sound-collecting target. The sound-collecting area determination process (S4) determines the left half sound-collecting area 42 (Figure 6(B)), which includes face regions R1 and R2 that have been determined to be the sound-collecting targets, to be the sound-collecting area. Sound-collecting control using face recognition (S5) is performed by setting sound-collecting parameters to control the directionality of the left half sound-collecting area 42 so that the voices of people 21 and 22 are clearly collected.

[0078] [1-2-2. Selection process for sound collection targets] The details of the sound collection target selection process in step S3 of Figure 10 will be explained using Figures 12 and 13.

[0079] Figure 12 is a flowchart illustrating the selection process (S3) for sound recording targets of the digital camera 100. Each process shown in the flowchart in Figure 12 is executed, for example, by the controller 135 of the digital camera 100 when the step S11 in Figure 10 is answered with YES.

[0080] Figure 13 is a diagram illustrating the selection process (S3) for selecting the sound to be recorded in the digital camera 100. Below, the operation of determining the sound to be recorded will be explained using the examples in Figures 11(A) and (B).

[0081] In the flowchart of Figure 12, the controller 135 determines the subject corresponding to the face region identified as the AF target in step S1 of Figure 10 as the sound collection target (S10). At this time, the controller 135 sets the size of the face region of the AF target (i.e., face width W) as a criterion for selecting the sound collection target from other subjects, based on the detection information obtained from the face recognition unit 122.

[0082] Figure 13(A) illustrates the case where the sound collection target is selected in the examples of Figures 11(A) and (B). Face widths W1, W2, and W3 indicate the size of face regions R1, R2, and R3 in the captured image Im as widths in the X-axis direction. In the example of Figure 13(A), the controller 135 sets the face width W1 of the face region R1 targeted for AF to the reference face width W (S10). The set face width W is stored, for example, in the RAM of the controller 135.

[0083] Next, the controller 135 determines whether there are any other subjects detected besides the AF target (S11). Specifically, the controller 135 determines whether the detection information from the face recognition unit 122 includes any face areas other than the face area of ​​the AF target.

[0084] If there are other subjects detected besides the AF target (YES in S11), the controller 135 selects one subject i as a candidate for sound pickup (S12). In the example in Figure 13(A), the detected information includes the other face regions R2 and R3 of the AF target face region R1, which are sequentially selected in each step S12, corresponding to the candidate subject i for sound pickup.

[0085] The controller 135 performs a calculation to compare the face width Wi of the selected subject i with the reference face width W (S13). Specifically, the controller 135 calculates the ratio Wi / W of the face width Wi of subject i to the reference face width W. In the example in Figure 13(A), when the face region R2 is selected as a sound collection candidate (S12), the ratio W2 / W for its face width W2 is calculated (S13).

[0086] The controller 135 determines whether the ratio Wi / W between the face width Wi of the sound-collecting candidate and the reference face width W is within a predetermined range (S14). The predetermined range is defined, for example, by an upper limit greater than "1" and a lower limit less than "1," from the viewpoint of defining the range in which the face width Wi of the sound-collecting candidate is considered to be relatively about the same as the reference face width Wi. A user interface for setting the predetermined range may be provided, and for example, the predetermined range set by the user by the operation unit 150 may be stored in the buffer memory 125 or the like.

[0087] When the controller 135 determines that the ratio of face width Wi / W is within a predetermined range (YES in S14), it decides to select subject i as the target for sound collection (S15).

[0088] On the other hand, if the controller 135 determines that the face width ratio Wi / W is not within a predetermined range (NO in S14), the controller 135 decides not to include subject i in the sound collection target (S16). In the example in Figure 13(A), the ratio W2 / W falls below the lower limit of the predetermined range, and it is decided not to include face region R2 in the sound collection target.

[0089] When the controller 135 decides whether or not to include subject i as a sound recording target (S15 or S16), it records the result of the decision regarding subject i in the buffer memory 125 (S17). Next, the controller 135 repeats the processing from step S11 onwards for subjects other than those already selected as sound recording candidates.

[0090] In the example in Figure 13(A), in addition to face region R2, face region R3 is also included in the detection information (YES in S11). The controller 135 selects the subject corresponding to face region R3 (S12) and, as in the case of face region R2, calculates the ratio W3 / W of face width W3 to the reference face width W (S13). In the example in Figure 13(A), the ratio W3 / W is calculated to be close to "1". The controller 135 determines that the calculated face width ratio W3 / W is within the predetermined range of the sound collection target (YES in S14) and decides to select the subject corresponding to face region R3 as the sound collection target (S15).

[0091] The controller 135 repeats the processes in steps S11 to S17 until there are no more subjects that have not been selected as candidates for sound collection (NO in step S11). After that, the controller 135 finishes the selection process for sound collection targets (S3) and proceeds to step S4 in Figure 10.

[0092] Through the above process, for subjects detected by face recognition, the relative sizes of face regions R2 and R3 are compared with respect to face region R1, which has been identified as the AF target. This allows subjects whose relative face region R3 is approximately the same size as the AF target face region R1 to be selected and designated as sound recording targets.

[0093] Figure 13(B) illustrates the case in which the sound pickup target is selected in the example of Figure 11(C). In the example of Figure 13(B), the face region R1 is identified as an AF target, similar to the example of Figure 13(A). Therefore, the controller 135 determines that the face region R1 is the sound pickup target and sets the face width W1 to the reference face width W (S10).

[0094] In the example in Figure 13(B), the face width W2 of face region R2 is about the same size as the reference face width W (=W1). On the other hand, the face width W3 of face region R3 is larger than the other face widths W1 and W2. In this example, the controller 135 determines that the ratio W2 / W is within a predetermined range (YES in S14) and decides to select the subject in face region R2 as the sound recording target (S15). On the other hand, since the ratio W3 / W exceeds the upper limit of the predetermined range (NO in S14), it is decided that the subject in face region R3 will not be selected as the sound recording target (S16). Therefore, the sound recording targets in this example are determined to be the two subjects corresponding to face regions R1 and R2 (see Figure 11(C)).

[0095] Figure 13(C) illustrates a case where, in the same captured image Im as in Figure 11(C), face region R3 is identified as the face region 60 to be AF-targeted (S1 in Figure 10). The controller 135 determines face region R3 to be the target for sound collection and sets face width W3 to the reference face width W (S10). In the example in Figure 13(C), since the ratios W2 / W and W1 / W are below the lower limit of the predetermined range (NO in S14), it is decided that subjects corresponding to face regions R1 and R2 will not be the target for sound collection (S16). Therefore, the target for sound collection in this example is determined to be one subject corresponding to face region R3.

[0096] As described above, the digital camera 100 of this embodiment can be used to determine a sound pickup area that aligns with the user's intentions, as described later, by selecting a subject of similar size to the AF target from among multiple subjects detected by image recognition as the sound pickup target.

[0097] [1-2-3. Determination of the sound pickup area] The details of the sound pickup area determination process in step S4 of Figure 10 will be explained using Figures 14 and 15.

[0098] Figure 14 is a flowchart illustrating the process of determining the sound pickup area (S4) in the digital camera 100 of this embodiment. Each process shown in the flowchart in Figure 14 is executed, for example, by the controller 135 of the digital camera 100 after step S3 in Figure 10 has been performed.

[0099] Figure 15 is a diagram illustrating the process of determining the sound pickup area (S4) in the digital camera 100. Figures 15(A) and (B) illustrate cases where the sound pickup area is determined following the examples in Figures 13(A) and (B), respectively. Figure 15(C) illustrates yet another case different from Figures 15(A) and (B). In Figures 15(A) to (C), the center position x0 indicates the position of the center of the captured image Im in the X-axis direction, and the image width Wh indicates the width of the captured image Im in the X-axis direction. The image range is defined as the range x0±xh from X-coordinate -xh to xh, with respect to the center position x0 on the captured image Im. The X-coordinate xh is defined as xh = Wh / 2 (>0).

[0100] In the flowchart of Figure 14, the controller 135 determines whether the position of the center of the face region or other location for all sound pickup targets is within the central range of the captured image Im (S20). The central range is the range in the captured image Im that corresponds to the front central sound pickup area 41.

[0101] The central range is defined as the range x0±xe from the X coordinate -xe to xe, with respect to the center position x0 on the captured image Im, as shown in Figure 15(A), for example. The X coordinate xe is defined, for example, as xe = xh × θe / θh (>0), based on a predetermined field of view θe and the horizontal field of view θh corresponding to the image width Wh. The predetermined field of view θe is set in advance, for example, from a viewpoint that includes one person, such as 30°. The controller 135 obtains the current horizontal field of view θh from, for example, the zoom magnification of the zoom lens of the optical system 110, and calculates the central range x0±xe.

[0102] In wide-angle photography, where the horizontal field of view θh is large, the X-coordinate xe becomes small, and the central range x0±xe is narrow. On the other hand, in telephoto photography, where the horizontal field of view θh is small, the X-coordinate xe becomes large, and the central range x0±xe is wide. This makes it easier to determine the sound pickup area corresponding to the physical range and distance of the image being captured.

[0103] If the position of all face regions to be sound-collected is within the central range (YES in S20), the controller 135 determines the sound-collecting area to be the front central sound-collecting area 41 (S21). In the example in Figure 15(A), the sound-collecting targets correspond to face regions R1 and R3. The central positions x1 and x3 of each face region R1 and R3 are both within the range of x0±xe (YES in S20). Therefore, the sound-collecting area is determined to be the front central sound-collecting area 41 (S21, see Figure 11(B)).

[0104] On the other hand, if the position of at least one face region of the sound pickup target is not within the central range (NO in S20), a sound pickup area other than the front central sound pickup area 41 is used. In this case, the controller 135 determines for all sound pickup targets whether, for example, the position of the face region is within only the left or right half of the captured image Im (S22). The left half is the range where the X coordinate is smaller than the center position x0 in the X axis direction, and the right half is the range where the X coordinate is larger than the center position x0.

[0105] If, for all sound pickup targets, the position of the face region is within the left half or right half of the captured image Im (YES in S22), the controller 135 further determines whether the position of the face region of all sound pickup targets is within the left half of the captured image Im (S23).

[0106] If the position of the face region to be sound-collected is within the left half of the captured image Im (YES in S23), the controller 135 determines the sound-collecting area to be the left half sound-collecting area 42 (S24). In the example in Figure 15(B), the sound-collecting targets correspond to face regions R1 and R2. Since the position x1 of face region R1 and the position x2 of face region R2 are to the left of the center position x0 in the X-axis direction (i.e., the X coordinate is small) (YES in S23), the sound-collecting area is determined to be the left half sound-collecting area 42 (S24, see Figure 11(C)).

[0107] On the other hand, if the position of the face region to be fully sound-collected is within the right half of the captured image Im, but not within the left half (NO in S23), the controller 135 determines the sound-collecting area to be the right half sound-collecting area 43 (S25).

[0108] Furthermore, if the position of the face region of all sound-collecting targets is not within either the left or right half of the captured image Im (NO in S22), the controller 135 determines the sound-collecting area to be the front sound-collecting area 44 (S26). As shown in Figures 6(D) and (A), the front sound-collecting area 44 has a wider angular range 402 than the angular range 401 of the front central sound-collecting area 41. That is, the front sound-collecting area 44 includes subjects of sound-collecting targets that are located in a wide area in the X-axis direction in the captured image Im.

[0109] In the example in Figure 15(C), the sound pickup targets correspond to facial regions R1, R2, and R3. The central positions x1, x2, and x3 of facial regions R1 to R3 include positions x1 and x2 outside the central range x0±xe (NO in S20), and also include position x1 within the left half and positions x2 and x3 within the right half (NO in S22 and S23). Therefore, in this example, the sound pickup area is determined to be the front sound pickup area 44 (S26).

[0110] Once the controller 135 determines the sound pickup area (S21, S24~S26), it records the determined sound pickup area as management information in the buffer memory 125 or the like (S27). This completes the sound pickup area determination process (S4), and the process proceeds to step S5 in Figure 10.

[0111] Through the above process, the sound pickup area is determined from a number of predefined sound pickup areas, based on the position of the subject selected as the sound pickup target on the captured image, so as to include all the sound pickup targets. This makes it possible to determine the sound pickup area in video recording so as to include the subject of the sound pickup target as intended by the user.

[0112] Figure 17 is a diagram illustrating the management information obtained by the sound collection area determination process (S4). Figure 17(A) illustrates the management information obtained at the stage when the sound collection target selection process (S3) and the sound collection area determination process (S4) are performed in the examples of Figures 13(A) and 15(A). Figure 17(B) illustrates the management information in the examples of Figures 13(B) and 15(B).

[0113] The management information is managed by associating the "sound collection target" determined by the sound collection target selection process (S3), the "sound collection area" determined by the sound collection area determination process (S4), the "horizontal angle of view," and the "focus distance." The focus distance is acquired, for example, when performing AF control by face recognition (S1). For example, the controller 135 may acquire the corresponding focus distance based on the position or focal length of the various lenses of the optical system 110 at the time of focusing. Alternatively, the digital camera 100 may detect the focus distance by DFD (Depth from Defocus) technology or measurement by a distance measuring sensor.

[0114] Furthermore, the digital camera 100 in this embodiment allows setting the field of view θe of the central range used in determining the front central sound pickup area (S20), and this setting is recorded, for example, in the ROM of the controller 135. In addition, a user interface for setting the field of view θe is provided, and the value set by the user via the operation unit 150 may be stored in the buffer memory 125 or the like.

[0115] [1-2-4. Sound Collection Control] (1) Regarding step S5 in Figure 10 The details of the sound collection control using face recognition in step S5 of Figure 10 will be explained using Figures 16 to 18.

[0116] In sound pickup control using sound pickup parameter settings, the digital camera 100 of this embodiment sets the sound pickup gain to, for example, emphasize the video audio for subjects corresponding to the face area of ​​the AF target. The sound pickup gain has, for example, frequency filter characteristics and stereo separation characteristics. The digital camera 100 calculates the sound pickup gain based on the horizontal angle of view and focus distance when the digital camera 100 focuses on the face area of ​​the AF target while shooting video. The sound pickup gain is defined such that, for example, a larger calculated value suppresses frequency bands other than human voices and controls the stereo effect, thereby producing a sound pickup zoom effect.

[0117] Figure 16 is a flowchart illustrating sound collection control (S5) using face recognition. Each process shown in the flowchart of Figure 16 is executed, for example, by the controller 135 of the digital camera 100, after step S4 in Figure 10 has been performed.

[0118] The digital camera 100 starts the process in step S5 with the management information shown in Figure 17 still in place.

[0119] The controller 135 obtains the horizontal field of view from the buffer memory 125, for example, and calculates the gain Gh based on the horizontal field of view (S30). Figure 18(A) illustrates the relationship between the horizontal field of view and the gain Gh. In the example in Figure 18(A), the gain Gh increases as the horizontal field of view decreases, between the predetermined maximum gain Gmax and minimum gain Gmin. This allows the gain to be increased during sound recording as the horizontal field of view decreases due to zooming, etc., thereby emphasizing the sound of subjects captured at the telephoto end.

[0120] The controller 135 obtains the focus distance in the same manner as in step S30 and calculates the gain Gd based on the focus distance (S31). Figure 18(B) illustrates the relationship between the focus distance and the gain Gd. In the example in Figure 18(B), the gain Gd increases as the focus distance increases, between the predetermined maximum gain Gmax and minimum gain Gmin. This allows the gain to be increased during sound recording when focusing on a subject farther away from the digital camera 100, thereby emphasizing the sound of distant subjects.

[0121] The controller 135 compares the calculated sound acquisition gain Gh based on the horizontal field of view with the sound acquisition gain Gd based on the focusing distance, and sets the larger of the two gains as the sound acquisition gain G (S32). This makes it possible to calculate the sound acquisition gain G in a way that emphasizes the sound of the subject in accordance with the intention of the user who is shooting with, for example, a telephoto horizontal field of view or a long focusing distance.

[0122] The controller 135 determines whether the sound pickup gain G and determined sound pickup area calculated over a predetermined number of past times (for example, 5 times) are the same (S33). For example, the sound pickup gain G is stored along with the above management information each time it is calculated within a predetermined number of times in the execution cycle of steps S1 to S5 in Figure 10. If the controller 135 determines that the sound pickup gain G and sound pickup area are the same for the predetermined number of past times (YES in S33), it proceeds to step S34.

[0123] The controller 135 sets the sound pickup target determined by the sound pickup target selection process in step S3, the sound pickup area determined by the sound pickup area determination process in step S4, and the sound pickup gain G calculated in step S32 as sound pickup parameters to the audio processing engine 170 (S34). The audio processing engine 170 uses the beamforming unit 172 and the gain adjustment unit 174 to realize the sound pickup area and sound pickup gain according to the set sound pickup parameters.

[0124] After setting the sound pickup parameters (S34), the controller 135 terminates the sound pickup control process using face recognition (S5). Also, if the controller 135 determines that the sound pickup gain G and sound pickup area are not the same as those of a predetermined number of previous instances (NO in S33), it terminates the process of step S5 in Figure 10 without performing the process of step S34. After that, the processes from step S1 onwards in Figure 10 are repeated.

[0125] Through the above process, the calculated sound pickup gain and the sound pickup target and sound pickup area determined based on face recognition can be set as sound pickup parameters to realize a sound pickup area and sound pickup gain that makes it easier to clearly capture the voice of the subject of the sound pickup target, including the AF target.

[0126] Note that the execution order of steps S30 and S31 is not limited to the order shown in this flowchart. For example, the gain Gd may be calculated in step S31 first, and then the gain Gh may be calculated in step S30, or steps S30 and S31 may be executed in parallel.

[0127] Furthermore, according to step S33 described above, the process of setting the sound pickup parameters (S34) is executed only if the sound pickup area and sound pickup gain G do not change for a predetermined number of times (for example, 5 times). This prevents the sound pickup area and sound pickup gain G from being changed excessively frequently due to the movement of the subject, and enables accurate sound pickup control using face recognition (S5) in accordance with the user's intentions.

[0128] (2) Regarding step S6 in Figure 10 The details of the sound pickup control (S6) without face recognition in step S6 of Figure 10 will be explained using Figure 19.

[0129] Figure 19 is a flowchart illustrating sound pickup control (S6) without face recognition. Each process shown in the flowchart of Figure 19 is executed, for example, by the controller 135 of the digital camera 100, when there is no face region to be AFed in step S2 of Figure 10 (NO in S2), such as when no face region is detected.

[0130] First, the controller 135 determines the sound pickup area to be, for example, the front sound pickup area 44 (S40).

[0131] Next, the controller 135 calculates the gain Gh based on the horizontal field of view in the same manner as in step S30, and sets it as the sound pickup gain G (S41). Furthermore, the controller 135 determines, in the same manner as in step S33, whether the sound pickup gain G and the determined sound pickup area calculated over a predetermined number of past times are the same as each other (S42).

[0132] If the controller 135 determines that the sound pickup gain G and sound pickup area are the same as those of a predetermined number of past attempts (YES in S42), it sets the sound pickup area and sound pickup gain G as sound pickup parameters (S43) and terminates the sound pickup control without face recognition (S6). If the controller 135 determines that the sound pickup gain G and sound pickup area are not the same as those of a predetermined number of past attempts (NO in S42), it terminates step S6 in Figure 10 without performing the processing in step S43. After the termination of step S6, the processing from step S1 onwards is repeated.

[0133] Through the above processing, even when there is no face area to be AF'd, the system can capture sound from a wide area in front of the digital camera 100, and the sound capture gain is increased as the horizontal angle of view becomes smaller due to zoom, etc., making it easier to clearly capture sound within the image area.

[0134] Furthermore, depending on the operating mode of the digital camera 100, an overall sound pickup area having an angular range of 360° around the digital camera 100 may be defined and determined as the overall sound pickup area in step S40. In this case, for example, only the overall sound pickup area may be set as the sound pickup parameter.

[0135] [1-2-5. Operation in Auto Mode] The details of the operation of the auto mode of the digital camera 100 according to this embodiment will be explained with reference to Figures 20 and 21.

[0136] Figure 20 is a flowchart illustrating the operation of the auto mode of the digital camera 100 according to Embodiment 1. Each process shown in the flowchart of Figure 20 is executed by the controller 135, for example, in the same way as in Figure 10, when the digital camera 100 is set to auto mode.

[0137] As shown in Figure 20, in the digital camera 100 in auto mode, the controller 135 determines whether the display monitor 130 is in a selfie position based on, for example, the detection signal from the magnetic sensor 132 (S51). If the controller 135 determines that the display monitor 130 is not in a selfie position (NO in S51), it sets the sound pickup area for non-face recognition to the sound pickup area 45 in surround mode (S52). On the other hand, if the controller 135 determines that the display monitor 130 is in a selfie position (YES in S51), it sets the sound pickup area for non-face recognition to the sound pickup area 46 in front mode (S53).

[0138] The controller 135 performs face recognition processing in the same way as the focus mode described above (S1, S2). For example, if a face area to be AF is detected (YES in S2), the controller 135 determines whether the digital camera 100 is in a vertical shooting position or not based on the detection signal from the acceleration sensor 137 (S54). If the controller 135 determines that it is not in a vertical shooting position (NO in S54), it performs the same steps S3 to S5 as in the focus mode and executes sound collection control. On the other hand, if the controller 135 determines that it is in a vertical shooting position (YES in S54), it adopts the sound collection area 46 of the front mode and executes sound collection control (S55).

[0139] Furthermore, if no face area is detected for AF (NO in S2), the controller 135 performs sound pickup control without face recognition based on the settings in steps S52 and S53 (S6A). The sound pickup control in step S6A is performed in the same way as in step S6 described above, using the sound pickup area set as the sound pickup area when no face recognition is performed.

[0140] The above process enables the operation of an auto mode that adjusts the directivity of the microphone 161 in conjunction with various shooting conditions. The sound pickup control for horizontal and vertical shooting in auto mode will be further explained with reference to Figure 21.

[0141] Figure 21(A) illustrates the relationship between the captured image Im and the sound pickup areas 41-43 during horizontal shooting. Figure 21(B) illustrates the case where step S55 is not performed during vertical shooting. Figure 21(C) illustrates the case where step S55 is performed during vertical shooting.

[0142] As illustrated in Figure 21(A), in horizontal shooting, the sound pickup area, face region R1-R3, and sound pickup areas 41-43 are coaxial, and by switching sound pickup areas 41-43, it is possible to control sound pickup as intended to align with the desired position of face region R1-R3. However, in vertical shooting, as shown in Figure 21(B), the relationship between sound pickup areas 41-43 and face region R1-R3 becomes misaligned. As a result, it is possible that sound pickup control cannot be performed as intended for the desired position of face region R1-R3, and instead, sound pickup control contrary to the intention may occur.

[0143] Therefore, in this embodiment, when shooting vertically, the sound pickup area is fixed to the sound pickup area 46 in front mode, as shown in Figure 21(C). This ensures a range in which sound can be picked up throughout the imaging range, thus avoiding situations where unintended sound pickup control occurs.

[0144] [1-3. Effects, etc.] In this embodiment, the digital camera 100 includes an image sensor 115 (an example of an imaging unit), a microphone 161 (an example of an audio acquisition unit), an operation unit 150 (an example of a setting unit), and a controller 135 (an example of a control unit). The image sensor 115 captures an image of a subject and generates image data. The microphone 161 acquires an audio signal indicating the sound picked up during imaging by the imaging unit. The operation unit 150, upon receiving instructions from the user, sets the device to auto mode, which is an operating mode that automatically changes the directivity of the audio acquisition unit. The controller 135 controls the sound pickup area in the audio signal for picking up sound from the subject. When set to auto mode, the controller 135 controls the sound pickup area to include the subject by changing the directivity of the microphone 161 in conjunction with the shooting state of the device. This enables appropriate sound pickup control according to various shooting conditions, making it easier to pick up the subject's sound according to the user's intentions when imaging while acquiring sound.

[0145] The digital camera 100 of this embodiment includes a face recognition unit 122, which is an example of a face detection unit that detects the face region of a subject in image data. When the controller 135 is set to auto mode, it determines the subject to be sound-collected in the audio signal based on the face region detected by the face recognition unit 122, and controls the sound-collecting area to include the subject determined to be sound-collected. This makes it easier to control sound collection according to the shooting conditions, such as various face recognitions of the subject, and to perform sound collection in accordance with the user's intentions.

[0146] In this embodiment, when the controller 135 is set to auto mode, it controls the sound pickup area so as to change the directivity of the sound acquisition unit in conjunction with the shooting state, whether the device is oriented vertically or horizontally. This allows for sound pickup control according to the shooting state, such as vertical or horizontal shooting, making it easier to acquire sound in accordance with the user's intentions.

[0147] In this embodiment, when the controller 135 is set to auto mode, it controls the sound pickup area so as to change the directivity of the sound acquisition unit in conjunction with the shooting state, such as whether or not the photographer is taking a picture of themselves. This makes it easier to control sound pickup according to the shooting state, such as whether or not it is a selfie, and to acquire sound in accordance with the user's intentions.

[0148] The digital camera 100 of this embodiment further comprises a display monitor 130, which is an example of a display unit, and a magnetic sensor 132, which is an example of a detection unit. The display monitor 130 has a display surface that displays an image of a subject, and is configured to be displaceable toward the subject. The magnetic sensor 132 detects whether or not the display surface of the display monitor 130 has been displaced toward the subject. The setting unit of this embodiment may set to auto mode when the magnetic sensor 132 detects that the display surface of the display monitor 130 has been displaced toward the subject. For example, the controller 135 may automatically set the digital camera 100 to auto mode when the display monitor 130 is in the selfie position in response to a detection signal from the magnetic sensor 132.

[0149] In the digital camera 100 of this embodiment, the setting unit can be set to at least one of a plurality of operating modes in which the directivity of the audio acquisition unit differs from that of the auto mode, in response to user instructions. For example, it can be set to surround mode, front mode, or navigation mode, and may also be set to focus mode.

[0150] In the digital camera 100 of this embodiment, when the setting unit is set to auto mode and imaging is started by the image sensor 115, the controller 135 may display information on the display monitor 130 indicating that the camera is set to auto mode along with the subject. For example, the controller 135 may display an icon or the like specifically for auto mode on the display monitor 130.

[0151] (Embodiment 2) Embodiment 2 will be described below with reference to the drawings. Embodiment 1 described a digital camera 100 that selects and determines the sound to be recorded when shooting video, etc. Embodiment 2 describes a digital camera 100 that visualizes information about the determined sound to be recorded to the user when operating as in Embodiment 1.

[0152] Hereinafter, descriptions of the configuration and operation similar to that of the digital camera 100 according to Embodiment 1 will be omitted as appropriate, and the digital camera 100 according to this embodiment will be described.

[0153] [2-1. Overview] Using Figure 22, an overview of the operation by which the digital camera 100 according to this embodiment displays various information will be explained.

[0154] Figure 22 shows an example of the display of the digital camera 100 according to this embodiment. The display example in Figure 22 shows an example of what is displayed in real time on the display unit 130 when the digital camera 100 determines a sound recording target as illustrated in Figure 11(B). In this display example, the digital camera 100 displays on the display monitor 130 an AF frame 11 indicating the subject to be AF, a detection frame 13 indicating detected subjects other than the AF target, and a sound recording icon 12 indicating the subject to be recorded, superimposed on the captured image Im.

[0155] In this embodiment, the digital camera 100 uses the sound pickup icon 12 in combination with the AF frame 11 and the detection frame 13 to visualize to the user whether the main subject, such as the AF target, and other detected subjects have been determined to be the AF target and / or the sound pickup target.

[0156] For example, in the display example in Figure 22, the digital camera 100 has determined that the subject corresponding to face region R1(60) in the example of Figure 11(B) is an AF target and a sound recording target, so it displays the AF frame 11 and the sound recording icon 12 on person 21. Also, the digital camera 100 has determined that the subject corresponding to face region R3 in the example of Figure 11(B) is a sound recording target other than an AF target, so it displays the detection frame 13 and the sound recording icon 12 on person 23. Furthermore, by displaying the detection frame 13 without the sound recording icon 12, the digital camera 100 visualizes to the user that it has decided not to record subjects other than the AF target corresponding to face region R2 in the example of Figure 11(B).

[0157] In the digital camera 100 of this embodiment, the user can confirm whether a detected subject is an AF target by the display of either the AF frame 11 or the detection frame 13. The user can also confirm whether a subject is a sound recording target by the presence or absence of the sound recording icon 12. The combination of the AF frame 11 and the sound recording icon 12 is an example of the first identification information in this embodiment. The combination of the detection frame 13 and the sound recording icon 12 is an example of the second identification information in this embodiment. The detection frame 13 is an example of the third identification information.

[0158] As described above, the digital camera 100 according to this embodiment displays a distinction between the subject to be sound-recorded and the subject to be autofocused, determined from the subject information included in the detection data. This allows the user to understand which subject to be sound-recorded among the subjects detected by the digital camera 100, and to confirm, for example, whether a subject that matches their intention has been determined to be sound-recorded.

[0159] Figures 23(A) and (B) illustrate the manual operation of the digital camera 100 in this embodiment. Figure 23(A) illustrates a state in which a specific person 22 is not detected by the face recognition unit 122. For example, even if the photographer wants to record the voice of person 22, face recognition may not be performed because person 22's face is facing sideways or backward to the digital camera 100. It is also conceivable that the sound source the photographer wants to record is not a person. Taking these cases into consideration, the digital camera 100 in this embodiment operates in a way that allows manual input to manually set the sound recording area.

[0160] Figure 23(B) illustrates the manual operation of the sound pickup area in the digital camera 100. For example, the manual operation of the sound pickup area is implemented as a touch operation. Figure 24 is a flowchart illustrating the operation during manual operation in the digital camera 100. The processes shown in this flowchart may be executed independently of the auto mode or focus mode processes described above, or they may be executed by interrupting the auto mode or focus mode processes.

[0161] First, the controller 135 of the digital camera 100 accepts manual operation from the user, such as the photographer (S61). As shown in Figure 23(B), for example, the controller 135 displays the specified range 48 of the sound pickup area on the display monitor 130 during manual operation. This example shows the case where manual operation is performed by interrupting the processing of auto mode or focus mode.

[0162] In the example shown in Figure 23(B), the photographer inputs a manual operation via touch to adjust the designated range 48 of the sound pickup area so that it includes the person 22 whose face was not recognized. The controller 135 determines the designated range 48 of the sound pickup area based on the input manual operation (S62). If the start and end points for the sound pickup area are input via touch as a manual operation, a predetermined size sound pickup area 48 including the start and end points will be displayed. If the confirmation button displayed on the display monitor 130 is touched in this state, the sound pickup area 48 will be confirmed.

[0163] The controller 135 reflects the determined specified range 48 of the sound pickup area in the sound pickup control of the microphone 161 (S63). As a result, the sound pickup control is performed in such a way that it emphasizes the sound coming from the sound pickup area corresponding to the specified range 48.

[0164] [2-3. Effects, etc.] As described above, the digital camera 100 of this embodiment includes an image sensor 115 that captures an image of a subject and generates image data, a microphone 161 that acquires an audio signal indicating the sound picked up during imaging by the image sensor 115, and a display monitor 130 that displays the image of the subject. The digital camera 100 of this embodiment includes an input unit such as an operation unit 150 that inputs a user operation to set the subject displayed on the display monitor 130 as a sound pickup area for picking up sound from the subject, and a controller 135 that controls the sound pickup area in the audio signal. When the controller 135 receives a user operation to set the sound pickup area, it controls the sound pickup area to include the subject by changing the directivity of the microphone 161 based on the user operation. This manual operation allows the sound pickup area to be controlled, making it easier to pick up sound in accordance with the user's intentions.

[0165] In the digital camera 100 of this embodiment, the display monitor 130 may be configured to be displaceable toward the subject, similar to Embodiment 1. User operations may be performed when the display monitor 130 is displaced toward the subject. That is, when taking a selfie with the digital camera 100, the manual operations described above may be input.

[0166] (Other embodiments) As described above, the embodiments described above have been presented as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited thereto and can be applied to embodiments that have been modified, replaced, added, or omitted as appropriate. Furthermore, it is possible to combine the components described in each of the embodiments above to create new embodiments.

[0167] Embodiment 1 described an example in which three microphone elements 161L, 161C, and 161R are used in the microphone 161. A modified example using four microphone elements will be described with reference to Figures 25 and 26.

[0168] Figure 25 illustrates the arrangement of the microphone 161A in the digital camera 100A of this modified example. In this modified example, the microphone 161A of the digital camera 100A includes a fourth microphone element 161B in addition to three microphone elements 161L, 161C, and 161R that are on the XZ plane relative to each other. The fourth microphone element 161B is positioned such that its position in the Y direction is different from that of the other microphone elements 161L to 161R.

[0169] Figure 26 illustrates the relationship between the captured image Im and the sound pickup areas 41A to 43A during vertical shooting in this modified example. With the configuration of the microphone 161A described above, the sound pickup area can also be changed in the Y-axis direction of the digital camera 100. Therefore, when controlling sound pickup during vertical shooting, by utilizing the sound pickup area using the fourth microphone element 161B, it is possible to achieve sound pickup control that follows face recognition even during vertical shooting, as shown in Figure 26, for example. For example, the controller 135 of the digital camera 100A in this modified example performs sound pickup control using the fourth microphone element 161B as described above, instead of step S55, in the same process as in Figure 20. This makes it possible to perform sound pickup control as intended, even during vertical shooting, by using the sound pickup areas 41A to 43A to align with the desired face region R1 to R3.

[0170] Furthermore, in the sound pickup control described in each of the above embodiments, by varying the speed of the transition period between sound pickup areas in conjunction with face recognition, auditory discomfort can be further suppressed. An example of this operation will be explained with reference to Figure 27.

[0171] Figure 27 shows an example of operation in which the width of the sound pickup directivity (i.e., the angular range of the sound pickup area) is changed in conjunction with the presence or absence of face recognition in the digital camera 100. In this example, at time t1, the face recognition unit 122 of the digital camera 100 detects the subject's face. Then, false detection prevention control (i.e., chattering) is performed (see S33 in Figure 16).

[0172] For example, if face recognition occurs only for a moment, or if the recognized subject turns to the side, changing the sound pickup direction could cause the user listening to the sound recording to experience an auditory unease. In contrast, the false detection prevention control described above constantly monitors the position of the subject's face recognition and changes the sound pickup direction when the subject remains in the sound pickup area for a certain period of time. This makes it possible to avoid such auditory unease.

[0173] Furthermore, the controller 135 of the digital camera 100 transitions the sound pickup area to narrow the sound pickup directivity after chattering from time t1. The transition period when narrowing the sound pickup directivity is set to be relatively short, for example. This allows for a quick change so that when a subject's face is detected in image recognition, the sound pickup directivity is focused towards the detected face, giving the user listening to the sound pickup results a good auditory impression. Similarly, transitions between sound pickup areas within the same angular range, such as the front central sound pickup area 41 and the left half sound pickup area 42, are also performed relatively quickly.

[0174] In the example shown in Figure 27, at time t2, the face recognition unit 122 can no longer detect the subject's face, for example, due to the subject moving or turning their face sideways. In this case, after chatter control, the digital camera 100 transitions the sound pickup directivity. Here, the transition period when widening the range of sound pickup directivity is set to be longer than the transition period when narrowing it as described above. This avoids causing the user auditory discomfort that may occur when sounds suddenly start to be heard from a wide range. Rather, by delaying the transition when widening the range of sound pickup directivity, auditory discomfort for the user can be suppressed.

[0175] Furthermore, in this example, at time t3, face recognition is performed again during the transition period in which the sound pickup directivity is widened. In this case, the digital camera 100 switches to a control that narrows the sound pickup directivity again as an interrupt control before it has fully widened. This allows the control to quickly direct the sound pickup directivity towards the subject whose face recognition has been performed, even when face recognition of the subject is intermittent, further suppressing any auditory discomfort for the user.

[0176] In each of the embodiments described above, a face recognition unit 122 was used to detect the sound collection target. In this embodiment, the detection of the sound collection target is not limited to the face recognition unit 122; for example, a human body recognition system that recognizes the whole or at least part of a human being may be used instead or in addition to the face recognition unit 122. Furthermore, the sound collection target does not necessarily have to be a person; for example, it may be various animals. In this case, the sound collection target may be detected by image recognition of part or the whole of an animal.

[0177] Furthermore, the first, second, and third identification information in Embodiment 2 identifies whether a subject is the main subject based on the presence or absence of the AF frame 11, and whether a subject is a sound recording target based on the presence or absence of the sound recording icon 12. In this embodiment, the first to third identification information is not limited to these, and for example, there may be three types of frame displays. Figure 22 illustrates three types of frame displays in this embodiment. In the example in Figure 22, the display of AF subjects and other subjects and the display of sound recording targets are integrated by a frame display 11A that shows a subject that is both an AF target and a sound recording target, a frame display 13A that shows a subject that is not an AF target but is a sound recording target, and a frame display 13B that shows a subject that is not a sound recording target.

[0178] In embodiments 1 and 2 described above, the operation of the auto mode and manual operation were explained separately, but these may be combined. That is, the digital camera 100 of this embodiment may include a display monitor 135 as a display unit for displaying an image of a subject, and an operation unit 150 as an input unit for inputting user operations to set the subject displayed on the display monitor 135 as a sound pickup area for picking up sound from the subject. When a user operation to set the sound pickup area is input to the controller 135, the controller 135 may control the sound pickup area to include the subject by changing the directivity of the microphone 161 based on the user operation. In such a case, the display monitor 135 does not necessarily have to be movable, and may be fixed, for example, in the normal position described above.

[0179] In Embodiments 1 and 2, the flowchart in Figure 10 describes an example of operation in which sound pickup control (S5 or S6) is performed with or without face recognition for the microphone 161 built into the digital camera 100. The digital camera 100 in this embodiment may be equipped with an external microphone (hereinafter referred to as "microphone 161a") instead of the built-in microphone 161. Microphone 161a includes microphone elements located outside the digital camera 100 and comprises three or more microphone elements. In this embodiment, the controller 135 can execute step S5 or S6 in the same manner as in Embodiment 1 by pre-holding information regarding the arrangement of microphone elements in the buffer memory 125 or the like for microphone 161a. In this case as well, the sound of the subject can be easily obtained clearly according to the sound pickup target and / or sound pickup area determined in the same manner as in Embodiment 1.

[0180] Furthermore, in Embodiments 1 and 2, an example of operation in which the gain Gh is calculated (S30) based on the horizontal field of view corresponding to the imaging range of the digital camera 100 was described in the flowchart of Figure 16. In this case, the horizontal field of view is the same as the horizontal field of view θh used in determining the front central sound pickup area (S20) in the flowchart of Figure 14. In this embodiment, a different horizontal field of view θh from the horizontal field of view θh in step S20 may be used to calculate the gain Gh. For example, the horizontal field of view in step S30 may be the angular range corresponding to the width in the X-axis direction that includes all the sound pickup targets on the captured image. This makes it possible to calculate the gain Gh in such a way that the voices of distant subjects are picked up more clearly, depending on the field of view in which the sound pickup targets are visible.

[0181] Furthermore, in embodiments 1 and 2, the face recognition unit 122 detected a human face. In this embodiment, the face recognition unit 122 may detect, for example, an animal's face. The size of an animal's face may vary depending on the animal species. Even in this case, the sound collection target can be selected in the same way as in embodiment 1 by, for example, expanding the predetermined range for selecting the sound collection target (see S14). Moreover, the face recognition unit 122 may detect faces for each animal species and set the predetermined range in step S14 according to the species.

[0182] Furthermore, embodiments 1 and 2 described a digital camera 100 equipped with a face recognition unit 122. In this embodiment, the face recognition unit 122 may be provided on an external server. In this case, the digital camera 100 may transmit image data of the captured image to the external server via a communication module 160 and receive detection information of the processing result by the face recognition unit 122 from the external server. In such a digital camera 100, the communication module 160 functions as a detection unit.

[0183] Furthermore, embodiments 1 and 2 illustrate a digital camera 100 equipped with an optical system 110 and a lens drive unit 112. The imaging device in this embodiment does not necessarily have to be equipped with an optical system 110 and a lens drive unit 112; for example, it may be a camera with interchangeable lenses.

[0184] Furthermore, while digital cameras were described as examples of imaging devices in Embodiments 1 and 2, the invention is not limited to them. The imaging device of this disclosure may be any electronic device having an image capture function (for example, a video camera, smartphone, tablet terminal, etc.).

[0185] As described above, embodiments have been explained as examples of the technology in this disclosure. For this purpose, accompanying drawings and a detailed description have been provided.

[0186] Therefore, the components described in the attached drawings and detailed descriptions may include not only components essential for solving the problem, but also components that are not essential for solving the problem, provided that they illustrate the technology described above. For this reason, the mere presence of these non-essential components in the attached drawings and detailed descriptions should not be immediately assumed to mean that they are essential.

[0187] Furthermore, since the embodiments described above are for illustrative purposes of the technology described herein, various modifications, substitutions, additions, omissions, etc., can be made within the claims or their equivalents. [Industrial applicability]

[0188] This disclosure is applicable to imaging devices that perform imaging while acquiring sound. [Explanation of symbols]

[0189] 100 Digital Cameras 115 Image Sensor 120 Image Processing Engines 122 Face Recognition Unit 125 buffer memory 130 Display Monitor 132 Magnetic Sensor 135 Controller 137 Accelerometer 145 Flash Memory 150 Operation section 11 AF frame 12. Sound recording icon 13 detection frames

Claims

1. An imaging unit that captures images of a subject and generates image data, A subject detection unit for detecting a subject region corresponding to the subject in the image data, A sound acquisition unit that acquires an audio signal indicating sound picked up during imaging by the aforementioned imaging unit, A display unit that displays an image of the subject, An input unit for inputting user operations to set the subject displayed on the display unit as a sound pickup area for capturing sound from the subject, The system includes a control unit that controls the sound pickup area, The control unit determines the subject to be picked up in the audio signal based on the subject area detected by the subject detection unit, and controls the sound pickup area to include the subject determined to be picked up. When there are undetected subjects that have not been detected by the subject detection unit, the control unit displays information on the display unit indicating the subject determined to be the sound pickup target from the detected subject area, and when a user operation to set the undetected subject to be included in the sound pickup area is input to the input unit, the control unit controls the range of the sound pickup area to include the undetected subject by changing the directivity of the sound acquisition unit based on the user operation. Imaging device.

2. The display unit is configured to be able to displace its display surface toward the subject, The imaging apparatus according to claim 1, wherein the user operation is performed when the display unit is in a state in which the display surface is displaced toward the subject.

Citation Information

Patent Citations

  • Video camera

    JP2010283706A

  • Imaging device

    JP2011114769A

  • Imaging control device and control method therefor

    JP2017022560A

  • Mobile terminal and audio zooming method thereof

    US20130342730A1