Imaging device, control method, and program

The imaging device enhances audio recording by identifying and adjusting sound based on the photographer's selected subject, ensuring the recorded audio matches their intended experience, even when subjects move out of view.

JP7774997B2Active Publication Date: 2025-11-25CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021140207
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-30
Publication Date
2025-11-25
Estimated Expiration
2041-08-30

AI Technical Summary

Technical Problem

Existing imaging devices fail to accurately record audio in a way that matches the photographer's imagination, as they rely solely on positional sound reproduction, neglecting the influence of visual cues like mouth movements and gestures.

Method used

The imaging device includes detection means for identifying subjects, determining related subjects, and adjusting audio based on the photographer's selected main subject, even when related subjects move out of the field of view.

Benefits of technology

This approach ensures that the recorded audio aligns with the photographer's perception, providing a coherent audio-visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007774997000001
    Figure 0007774997000001
  • Figure 0007774997000002
    Figure 0007774997000002
  • Figure 0007774997000003
    Figure 0007774997000003
Patent Text Reader

Abstract

To record a moving image and audio as a photographer images.SOLUTION: An audio processing device has detection means which detects a subject from a moving image, decision means which decides the voice of the subject from the moving image, selection means which selects a main subject among subjects detected from the moving image, and determination means which determines a subject related to the main subject selected by the selection means, wherein the determination means determines, if the subject related to the main subject moves out of the angle of view of the moving image, whether the subject having moved out of the angle of view of the moving image is still related to the main subject based upon the voice of the subjects determined by the decision means.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice processing device that performs voice processing on a person's voice. [Background technology]

[0002] When shooting moving images with an imaging device, it is important to record the shooting situation exactly as imagined by the photographer, and this applies not only to video but also to audio.

[0003] Patent Document 1 discloses that an acoustic space with a sense of realism and stereoscopic effect is realized by extracting the sound of a subject and adjusting the extracted sound signal individually according to the position of the subject. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-138930 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when humans listen to a conversation, the accurately reproduced acoustic space does not necessarily match what humans imagine. For example, even when many people are chatting away, humans can naturally hear the conversation of the person they are interested in, or their own name. It is also said that humans use not only audio information but also visual information, and by visually confirming the speaker, they supplement their hearing with information obtained from the person's mouth movements and gestures. In other words, it is important to record the audio recorded on video so that it matches the conversational audio that is remembered (imaged) by humans.

[0006] However, since the purpose of Patent Document 1 is to accurately reproduce the acoustic space of the voice based on the positional relationship of the person (sound source), there is a risk that the resulting video may differ from the image of the photographer.

[0007] Therefore, an object of the present invention is to record video and audio in accordance with the photographer's imagination. [Means for solving the problem]

[0008] The imaging device of the present invention comprises a detection means for detecting a subject from a moving image, a determination means for determining a sound of the subject from the moving image, a selection means for selecting a main subject from the subjects detected from the moving image, and a determination means for determining a subject related to the main subject selected by the selection means, When a subject related to the main subject moves out of the field of view of the video, the determination means determines whether the subject that has moved out of the field of view of the video continues to be related to the main subject based on the sound of the subject determined by the determination means. [Effects of the Invention]

[0009] According to the present invention, it is possible to record video and audio in accordance with the photographer's imagination. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram of an imaging apparatus according to a first embodiment. [Figure 2] FIG. 2 is a block diagram of an imaging processing unit and an audio processing unit according to the first embodiment (during recording); [Figure 3] FIG. 2 is a block diagram of an imaging processing unit and an audio processing unit (during post-processing) according to the first embodiment. [Figure 4] FIG. 10 is a diagram illustrating a main target selection method according to the first embodiment. [Figure 5] FIG. 4 is a diagram showing an operation flow of a moving image recording sequence according to the first embodiment. [Figure 6]FIG. 1 is a diagram illustrating an assumed scene according to a first embodiment. [Figure 7] FIG. 2 is a diagram illustrating the content of audio processing according to the first embodiment. [Figure 8] FIG. 10 is a block diagram of an imaging processing unit and an audio processing unit according to a second embodiment. [Figure 9] FIG. 10 is a diagram showing an operation flow of a video recording sequence according to the second embodiment. [Figure 10] FIG. 10 is a diagram for explaining a problem of the second embodiment. [Figure 11] FIG. 10 is a diagram illustrating a problem scene in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0012] [First embodiment] In this embodiment, an audio processing device included in an imaging device will be described with reference to FIGS.

[0013] FIG. 1 is a block diagram showing the configuration of an image capturing apparatus 100 according to the first embodiment.

[0014] The imaging unit 101 converts an optical image of a subject captured by a photographing optical lens into an image signal using an imaging element, and the image processing unit 102 performs analog-to-digital conversion, image adjustment processing, and other processes to generate image data. The photographing optical lens may be a built-in optical lens or a detachable optical lens. The imaging element may be a photoelectric conversion element such as a CCD or CMOS. The audio input unit 103 uses a microphone built in or connected via an audio terminal to collect audio from the surroundings of the imaging device 100, and the analog-to-digital converted audio is subjected to various audio processes by the audio processing unit 104 to generate audio data. The microphone may be directional or omnidirectional. The memory 105 temporarily stores image data obtained by the imaging unit 101 and image processing unit 102, and audio data obtained by the audio input unit 103 and audio processing unit 104. The display control unit 106 displays images related to the image data obtained by the image processing unit 102, an operation screen of the imaging device 100, a menu screen, etc. on the display unit 107 or on an external display via a video terminal (not shown). The display unit 107 has a touch panel function, and the photographer can operate it to select menus, subjects, etc.

[0015] The encoding processor 108 reads image data and audio data temporarily stored in the memory 105 and performs predetermined encoding to generate compressed image data, compressed audio data, etc. The audio data may not be compressed. The compressed image data may be compressed using any compression method, such as MPEG2 or H.264 / MPEG4-AVC. The compressed audio data may be compressed using a compression method such as AC3(A)AC, ATRAC, or ADPCM. The recording / playback unit 109 records the compressed image data, compressed audio data, or audio data generated by the encoding processor 108, and various data, onto the recording medium 110, and reads them from the recording medium 110. The recording medium 110 may be any type of recording medium capable of recording image data, audio data, etc., such as a magnetic disk, optical disk, or semiconductor memory.

[0016] The control unit 111 controls each block of the imaging device 100 by transmitting control signals to the imaging unit 101 and each block of the imaging device 100, and includes a CPU and memory for executing various controls. The memory 105 used by the control unit 111 includes a ROM for storing various control programs, a RAM for arithmetic processing, and other memory external to the control unit 111. The operation unit 112 includes buttons, dials, and other controls and transmits instruction signals to the control unit 111 in response to user operations. In the imaging device of this embodiment, the operation unit 112 includes a shooting button for starting and stopping video recording, a zoom lever for optically or electronically zooming an image, a cross key for various adjustments, and an enter key. The audio output unit 113 outputs audio data and compressed audio data reproduced by the recording / playback unit 109, or audio data output by the control unit 111, to a speaker 114, an audio terminal, or the like. The external output unit 115 outputs compressed video data, compressed audio data, audio data, and the like reproduced by the recording / playback unit 109 to an external device. A data bus 116 supplies various data such as audio data and image data, and various control signals to each block of the image pickup device 100 .

[0017] Here, the normal operation of the imaging device 100 of this embodiment will be described.

[0018] In the imaging device 100 of this embodiment, when a user operates the operation unit 112 to issue an instruction to turn on the power, a power supply unit (not shown) supplies power to each block of the imaging device.

[0019] When power is supplied, the control unit 111 checks which mode, such as the shooting mode or the playback mode, the mode selector switch of the operation unit 112 is in, based on an instruction signal from the operation unit 112. In the moving image recording mode, image data (video data) obtained by the imaging unit 101 and the image processing unit 102 and audio data obtained by the audio input unit 103 and the audio processing unit 104 are saved as a moving image file. In the playback mode, the compressed image data recorded on the recording medium 110 is played back by the recording / playback unit 109 and displayed on the display unit 107.

[0020] In the moving image recording mode, first, the control unit 111 sends a control signal to each block of the imaging device 100 to transition to a shooting standby state, and causes the blocks to perform the following operations. The imaging unit 101 converts an optical image of a subject captured by the imaging optical lens into an image signal using an imaging element, and the image processing unit 102 performs image adjustment processing and the like to generate image data. The obtained image data is then sent to the display control unit 106, which displays it on the display unit 107. The user prepares for shooting while looking at the screen displayed in this way.

[0021] The audio input unit 103 converts analog audio signals obtained by multiple microphones into digital audio signals, processes the obtained digital audio signals, and generates multi-channel audio data. The obtained audio data is then sent to the audio output unit 113, which outputs the audio from a connected speaker 114 or earphones (not shown). The user can adjust the manual volume to determine the recording volume while listening to the audio output in this manner.

[0022] Next, when the user operates the recording button on the operation unit 112 to send an instruction signal to start shooting to the control unit 111, the control unit 111 sends an instruction signal to start shooting to each block of the imaging device 100, causing it to perform the following operations.

[0023] The imaging unit 101 converts an optical image of a subject captured by a photographing optical lens into an image signal using an imaging element, and performs image adjustment processing and the like in an image processing unit 102 to generate image data. The obtained image data is then sent to a display control unit 106, which displays the image on a display unit 107. The obtained image data is also sent to a memory 105.

[0024] The audio input unit 103 digitally converts analog audio signals obtained by multiple microphones, and the audio processing unit 104 processes the obtained digital audio signals to generate multi-channel audio data. The obtained audio data is then sent to the memory 105. When there is only one microphone, the obtained analog audio signals are digitally converted to generate audio data, and the audio data is sent to the memory 105.

[0025] The encoding processor 108 reads out the image data and audio data temporarily stored in the memory 105 and performs predetermined encoding to generate compressed image data, compressed audio data, and the like.

[0026] Then, the control unit 111 combines the compressed image data and compressed audio data to form a data stream, which is output to the recording / playback unit 109. If the audio data is not compressed, the control unit 111 combines the audio data stored in memory 105 with the compressed image data to form a data stream, which is output to the recording / playback unit 109. The recording / playback unit 109 writes the data stream as a single moving image file to the recording medium 110 under the management of a file system such as UDF or FAT. The above operations continue while shooting is in progress.

[0027] Then, when the user operates the recording button on the operation unit 112 to send an instruction signal to end shooting to the control unit 111, the control unit 111 sends an instruction signal to end shooting to each block of the imaging device 100, causing it to perform the following operations.

[0028] The imaging unit 101, image processing unit 102, audio input unit 103, and audio processing unit 104 stop generating image data and audio data, respectively. The encoding processing unit 108 reads the remaining image data and audio data stored in memory, performs predetermined encoding, and stops operation once it has finished generating compressed image data, compressed audio data, etc. If audio data is not compressed, it naturally stops operation once it has finished generating compressed image data.

[0029] The control unit 111 then combines the final compressed image data with the compressed audio data or the audio data to form a data stream and outputs it to the recording / playback unit 109. The recording / playback unit 109 writes the data stream as a single moving image file to the recording medium 110 under the management of a file system such as UDF or FAT. When the supply of the data stream stops, the control unit 111 completes the moving image file and stops the recording operation. When the recording operation stops, the control unit 111 transmits a control signal to each block of the imaging device 100 to transition to a shooting standby state, and the imaging device 100 returns to the shooting standby state.

[0030] Next, in the playback mode, the control unit 111 transmits a control signal to each block of the imaging device 100 to transition to the playback state, causing the blocks to perform the following operations: The recording / playback unit 109 reads a moving image file consisting of compressed image data and compressed audio data recorded on the recording medium 110, and sends the read compressed image data and compressed audio data to the encoding processing unit 108.

[0031] The encoding processing unit 108 decodes the compressed image data and compressed audio data and transmits them to the display control unit 106 and audio output unit 113, respectively. The display control unit 106 displays the decoded image data on the display unit 107. The audio output unit 113 outputs the decoded audio data from a built-in or attached external speaker.

[0032] As described above, the imaging device 100 of this embodiment can record and play back images and sounds.

[0033] In this embodiment, when obtaining an audio signal, the audio input unit 103 and the audio processing unit 104 perform processing such as level adjustment of the audio signal obtained by the microphone. This processing may be performed continuously after the device is started, or may be performed after a shooting mode is selected, or may be performed after a mode related to audio recording is selected. Furthermore, in a mode related to audio recording, the above processing may be performed in response to the start of audio recording. In this embodiment, the above processing is performed at the timing when video recording starts.

[0034] FIG. 2 is a block diagram showing an example of a detailed configuration of the imaging unit 101, image processing unit 102, audio input unit 103, and audio processing unit 104 of the imaging device 100 of this embodiment.

[0035] The imaging unit 101 includes an optical system such as an optical lens 201 that captures an optical image of a subject, and an imaging element 202 that converts the optical image of the subject captured by the optical lens 201 into an electrical signal (image signal). The imaging unit 101 also includes an optical lens control unit 203 that includes a position sensor for moving the optical lens 201 and a known driving mechanism such as a motor. While the imaging unit 101 is described in this embodiment as having the optical lens 201 and the optical lens control unit 203 built in, these may be detachable, interchangeable optical lenses. For example, when a user inputs an instruction such as a zoom operation or focus adjustment by operating the operation unit 112, the control unit 111 transmits a control signal (drive signal) to the optical lens control unit 203 to move the optical lens. In response to this control signal, the optical lens control unit 203 checks the position of the optical lens 201 using a position sensor and moves the optical lens 201 using a motor or the like.

[0036] The image processing unit 102 performs various image quality adjustment processes on the image signal converted by the image sensor 202 in the image adjustment unit 221 to form image data, and transmits the image data to the memory 105 via the data bus 116. Based on the image data formed here, the control unit 111 performs various adjustments such as focus adjustment and light amount adjustment.

[0037] Furthermore, in this embodiment, the image processing unit 102 has various detection functions. The person detection unit 222 extracts facial features, such as the eyes, nose, and mouth, from the image data formed by the image adjustment unit 221, and then detects the position and face size of the person in the image data. Information on these features is then stored in the memory 105, making it possible to individually recognize the subject person based on that information. The person detection unit 222 also has a person movement detection unit 223 that detects lip and head movements, and a person speech detection unit 224 that uses these movements to determine whether the person is speaking. The image processing unit 102 also has a main subject selection unit 225 that selects which person, from among the people detected by the person detection unit 222, will be the main subject (hereinafter also referred to as the main subject or main target) of audio processing. The main subject selection unit 225 selects the main target based on conditions set by the control unit 111. The conditions for main target selection by the main target selection unit 225 will be described later.

[0038] Furthermore, image processing unit 102 has a conversation group detection unit 226. Conversation group detection unit 226 detects a person who is conversing with the person selected by main subject selection unit 225 from among the people detected by person detection unit 222. This detection is determined based on the relative positions of the people, the direction of their faces, their movements, and the like. For example, conversation group detection unit 226 determines that the subject closest to the main subject is the person who is conversing with the main subject (a related person). Also, for example, conversation group detection unit 226 determines that the subject whose body, face, line of sight, etc. faces the main subject is the person who is conversing with the main subject. Also, when the main subject is moving, conversation group detection unit 226 determines that the subject in the direction of the movement is the person who is conversing with the main subject. This is because such a subject is likely to converse with the main subject in the near future.

[0039] If conversation group detection unit 226 determines that a person who is conversing with the main target has not conversed with the main target for a predetermined period of time, that person is determined to be a person who is not conversing with the main target (not related to the main target). In other words, if a person who is conversing with the main target has not conversed with the main target for a predetermined period of time, that person is determined to be a person who is conversing with the main target, even if it is determined that the person has not conversed with the main target.

[0040] Next, the audio input unit 103 and audio processing unit 104 will be described. The audio input unit 103 is a microphone 211 that converts audio vibrations into an electrical signal and outputs it as an audio signal. In this embodiment, the microphone 211 is a stereo type consisting of two channels, left and right channels (Lch / Rch), but it may also be a mono type with one channel, or a configuration having multiple microphones with two or more channels. The A / D conversion unit 212 is a means for converting an analog audio signal obtained by the microphone 211 into a digital audio signal.

[0041] The audio processing unit 104 is a block that performs various audio processes on the audio signal converted by the audio input unit 103. In this embodiment, the audio processing unit 104 includes an audio extraction unit 213, an audio adjustment unit 215, and an audio synthesis unit 217. The audio extraction unit 213 is capable of extracting (determining) audio into human audio and other audio (hereinafter referred to as "non-human audio"). Furthermore, the human audio extraction unit 214 is capable of extracting audio from human audio into individual audio based on information from the person detection unit 222. For example, the human audio extraction unit 214 extracts individual audio based on the audio frequency, volume, and intonation. Furthermore, in the first embodiment, the control unit 111 can associate audio with a subject based on the audio extracted by the human audio extraction unit 214 and the subject's movements detected by the image processing unit 102. For example, the subject's movements include the frequency of speech, the timing of speech, and mouth movements.

[0042] Furthermore, the audio adjustment unit 215 can individually perform audio processing for each frequency band using level adjustment, equalizer, etc. on the audio extracted by the audio extraction unit 213. In particular, the conversation audio adjustment unit 216 performs adjustment based on information from the conversation group detection unit 226, emphasizing the extracted audio to make it easier to hear, or reducing its volume to make it harder to hear. The details of this adjustment will be described later. Furthermore, the audio synthesis unit 217 synthesizes the audio individually adjusted by the audio adjustment unit 215 and returns it to a single audio signal. The amplitude of the synthesized audio signal is then adjusted to a predetermined level by an auto level controller (hereinafter, ALC 219). With the above configuration, the audio processing unit 104 performs predetermined processing on the audio signal, forms audio data, and transmits it to the memory 105.

[0043] FIG. 3 is a block diagram showing another example of the configuration of the image processing unit 102 and the audio processing unit 104 of the imaging device 100 of this embodiment. FIG. 3 differs from FIG. 2 in that the input sources of image data and audio data are different. In FIG. 2, the image signal is from the imaging unit 101, and the audio signal is from the audio input unit 103. On the other hand, in FIG. 3, the image and audio input sources are data stored in the memory 105. By using data temporarily stored (held) in the memory 105 in this way, the proposed method can be used not only for processing during shooting but also for post-processing after recording. Furthermore, the main subject selection unit 225 can also select a person to be the subject of audio processing from a series of video data.

[0044] An example of a method for selecting a main subject by the main subject selection unit 225 will now be described with reference to FIG. 4. In this embodiment, the main subject will be described as a person who is thought to be the subject of the photographer's attention. For example, in the case of FIG. 4(a), the focus mark 402 is a mark indicating the subject on which the image capture device 100 is focusing. In FIG. 4(a), the main subject 401 and the focus mark 402 match, so the image capture device 100 recognizes the main subject 401 as the main subject and focuses on the main subject 401. The main subject selection unit 225 determines this main subject 401 as the main subject. In this way, the person recognized as the main subject can be selected as the main subject.

[0045] 4(b) shows a method using a registered face image. The registered face image 403 is an image of a subject that has been registered in advance in the memory 105. The main subject selection unit 225 selects a person whose face is determined to match the face in the image as the main subject.

[0046] 4(c) shows a method for determining the main subject at the photographer's discretion. The photographer selects the main subject from among the people displayed on the display unit 107 by touching the touch panel of the display unit 107. The main subject selection unit 225 determines the subject selected by the photographer as the main subject.

[0047] 4(d) shows a method of using recorded video data. For example, when recorded video data 404 is stored in memory 105, main subject selection unit 225 determines person 405 who appears most frequently in video data 404 as the main subject. Alternatively, for example, main subject selection unit 225 may select a person who is frequently in focus.

[0048] Note that, for example, when the main object selection unit 225 selects a subject that is in focus as the main object, even if the main object loses focus, the main object selection unit 225 maintains the subject as the main object as long as the focus returns to the subject within a predetermined time. In other words, when the focus is lost from the main object for longer than a predetermined time, the main object selection unit 225 selects a new subject to be the main object.

[0049] Next, the operation of the imaging device 100 of this embodiment will be described with reference to FIGS.

[0050] 5 is a flowchart showing an example of a video recording sequence of the imaging device 100. The processing of the imaging device 100 is realized by software stored in a ROM (not shown) being loaded into the memory 105 and executed by the CPU. The processing of this flowchart is triggered when the imaging device 100 is powered on.

[0051] In step S501, the control unit 111 receives an instruction to start recording a moving image through an operation of the operation unit 112 by the user.

[0052] In step S502, the control unit 111 connects an audio path for recording audio.

[0053] In step S503, after the audio path is established, the control unit 111 performs initial setting of signal processing, including the control described in this embodiment, and starts signal processing for video recording. The following description focuses on the recording sequence. Until the signal processing for video recording is completed, the control unit 111 records the video to be recorded in the video.

[0054] In step S504, the person detection unit 222 of the image processing unit 102 detects the subject.

[0055] In step S505, the main object selecting unit 225 of the image processing unit 102 selects (determines) a main object from the subjects detected in step S504.

[0056] In step S506, conversation group detection unit 226 of image processing unit 102 determines the person (subject) who is conversing with the main subject selected in step S505.

[0057] In step S507, the voice extraction unit 213 of the voice processing unit 104 extracts the voice of a person.

[0058] Audio adjustment unit 215 of audio processing unit 104 performs adjustment processing on the audio extracted in step S507. The content of the audio adjustment processing differs depending on whether the subject (person) of the audio extracted in step S507 is a subject (person) that belongs to the main conversation group. Details of the audio adjustment processing will be described later using Figures 6 and 7, but will be explained simply in this flowchart.

[0059] In step S508, audio adjustment unit 215 of audio processing unit 104 determines whether the person whose voice was extracted in step S507 is a subject belonging to the main conversation group. If the person whose voice was extracted is a subject belonging to the main conversation group, the process proceeds to step S509. If the person whose voice was extracted is not a subject belonging to the main conversation group, the process proceeds to step S510.

[0060] In step S509, the audio adjustment unit 215 of the audio processing unit 104 adjusts the level of the extracted audio so that the volume of the audio is increased.

[0061] In step S510, the audio adjustment unit 215 of the audio processing unit 104 adjusts the level of the extracted audio so that the volume of the extracted audio is reduced. In step S511, the audio adjustment unit 215 of the audio processing unit 104 performs adjustment processing other than volume on the extracted audio.

[0062] In step S512, the voice synthesis unit 217 of the voice processing unit 104 synthesizes the extracted voices that have been individually adjusted to generate one piece of voice data.

[0063] In step S513, control unit 111 determines whether to end moving image recording. For example, control unit 111 determines to end moving image recording when a user operates operation unit 112 to instruct ending moving image recording, or when it is determined that the remaining capacity of recording medium 110 is low. If it is determined not to end moving image recording, control returns to the processing of step S504, and the sound recording sequence processing continues. If it is determined to end moving image recording, the processing of step S514 is executed.

[0064] Here, if it is determined not to end the video recording, the process returns to step S504. That is, during video recording, the main subject and the people who are conversing with the main subject are repeatedly determined. As a result, for example, even if the main subject disappears outside the angle of view or is out of focus, the control unit 111 can determine another subject as the main subject. Also, even if the number of people conversing with the main subject increases or decreases, the control unit 111 can determine the people who are conversing with the main subject accordingly.

[0065] In step S514, the control unit 111 disconnects the audio path and ends the signal processing.

[0066] The audio adjustment process will now be described with reference to FIGS.

[0067] 6 is a diagram showing an assumed scene for audio adjustment processing. Assume that four subjects (people), person 602 to person 605, are present within angle of view 601, and person 602 is conversing (speaking) with person 603, and person 604 is conversing (speaking) with person 605. In this case, if person 602 is selected as the main target of audio processing by main target selection unit 225, people 602 and 603 are detected as conversation group 610 from the image data by conversation group detection unit 226. In this case, the audio of people 602 and 603 is adjusted to be emphasized as noteworthy audio, and the audio of people 604 and 605 is adjusted to be unnecessary audio that is not to be emphasized.

[0068] 7(a) to 7(c) are diagrams showing the audio adjustment process. In Fig. 7, the person 602, the person 603, and the person 604 in Fig. 6 are represented as people A, B, and C, respectively (person 605 is not shown).

[0069] 7(a) shows the voice signals of persons A to C extracted by the person voice extraction unit 214. That is, signal 701 shows the extracted voice signal of person A, signal 702 shows the extracted voice signal of person B, and signal 703 shows the extracted voice signal of person C. In each signal, intervals with large amplitude indicate periods when the respective persons are speaking (utterances) (voiced timings), and intervals with small amplitude indicate periods when the respective persons are not speaking (silent timings). For example, comparing signals 704 and 705, since persons A and B are conversing, voiced timings and silent timings appear almost alternately. On the other hand, since person C is not a conversation partner of persons A and B, signal 706 does not alternate between voiced timings and silent timings as much as signals 704 and 705.

[0070] 7(b) shows the audio correction coefficient for each person. In this embodiment, a correction coefficient of 1.0 indicates that no level adjustment (gain adjustment) is performed. When the correction coefficient is greater than 1.0, the audio adjustment process is performed to emphasize the audio to make it easier to hear (to increase the volume), and when the coefficient is less than 1.0, the audio adjustment process is performed to make it harder to hear (to decrease the volume).

[0071] For example, a case will be described in which the conversation group detection unit 226 determines that person A and person B are conversing during period 710. In this case, since person A is the main target, the conversation audio adjustment unit 216 recognizes the voices of person A and person B as targets to be emphasized and sets large correction coefficients for each voice (coefficients 714 and 715). In this embodiment, the correction coefficients for person A and person B are set to the same value. This is because it is assumed that the user who is taking the picture is listening to both voices equally. On the other hand, the conversation audio adjustment unit 216 sets a small correction coefficient for the voice of person C, who is determined not to be conversing with person A, so that the voice of person C becomes relatively difficult to hear (coefficient 716). In this way, the conversation audio adjustment unit 216 emphasizes the voices of person A, who is the main target, and person B, who is his conversation partner, and reduces the volume of other voices. For example, the conversation audio adjustment unit 216 increases the gain and level of the voices of person A, who is the main target, and person B, who is his conversation partner, compared to those of other voices. This results in video data with images and sounds that match the image of the user who is taking the video.

[0072] 7(c) shows an audio signal that has been adjusted based on the correction coefficients of FIG. 7(b). For example, if the audio adjustment by the conversation audio adjustment unit 216 is achieved by gain adjustment, during the period 710, the audio of persons A and B who have been determined to be conversing (signals 724 and 725) will have a correction coefficient greater than 1.0, making them louder and easier for the user to hear. Meanwhile, the audio of person C who has not been determined to be conversing (signal 726) will have a correction coefficient less than 1.0, making them quieter and harder to hear. The extracted audio that has been individually adjusted in this way is synthesized by the audio synthesis unit 217, resulting in the generation of audio data in which only the conversation determined to be the target of attention is easy to hear.

[0073] In this embodiment, the sound related to the main subject is emphasized (corrected to be louder) and the sound unrelated to the main subject is made hard to hear (corrected to be quieter), but adjustment may be applied to only one of them. That is, it is sufficient if the correction coefficients for the main subject (person) and the subjects (people) who are the subjects of conversation with that subject are larger than the correction coefficients for the other subjects.

[0074] Furthermore, the emphasis method used by the conversation voice adjustment unit 216 is not limited to adjusting the overall gain as described above, but may be adjusted for each frequency in the frequency band of human voices using an equalizer or the like.

[0075] [Second embodiment] In the first embodiment, after selecting a main target, conversation groups of people who are conversing with the main target are detected based on their relative positions relative to the main target and their movements, and the audio of the conversation group is emphasized or other unnecessary audio is suppressed, thereby obtaining audio data that makes the conversation of interest easier to hear.

[0076] In the first embodiment, the method for detecting a conversation group is to determine the positional relationship, facial orientation, movement, etc. between people who are conversing with the person selected by main subject selection unit 225 from among the people detected by person detection unit 222. As such, in the first embodiment, the detection by conversation group detection unit 226 is performed based on people present within angle of view 601 of imaging device 100.

[0077] Now, suppose that person A, who is the main subject, and persons B and D (603, 606) within angle of view 601 are detected as a conversation group as shown in FIG. 10(a). If person B moves out of the angle of view due to a zoom or panning operation by the photographer, person B will be removed from the conversation group when the next conversation group is detected, as shown in FIG. 10(b), even if the conversation between persons A, B, and D continues. As a result, even if person B is participating in the conversation, conversation group detection unit 226 does not determine person B as part of the conversation group, which could result in the voice of person B not being emphasized and making the conversation difficult to hear.

[0078] The second embodiment aims to maintain the conversation group in the state it was in before it moved out of the field of view, and continue to acquire easy-to-listen audio, even if at least one member of the conversation group that was within the field of view moves out of the field of view, when it is determined that the conversation of the person who moved out of the field of view is continuing.

[0079] The second embodiment will be described in detail below with reference to the accompanying drawings. Note that the configuration of the image capture device 100 in Fig. 1 is the same as that of the first embodiment, and therefore a description thereof will be omitted.

[0080] 8 is a block diagram showing the detailed configuration of the imaging unit 101, image processing unit 102, audio input unit 103, and audio processing unit 104 of the imaging device 100 of this embodiment. Note that blocks having the same functions as those in FIG. 2 are assigned the same numbers, and their explanations will be omitted.

[0081] The feature extraction unit 801 associates the voice extracted by the person voice extraction unit 214 with the person corresponding to that voice. For example, the feature extraction unit 801 associates the extracted voice with the corresponding person based on the voice features and the movement of the subject within the field of view. For example, the voice features include frequency, volume, and intonation. For example, the movement of the subject includes the frequency of utterance, the timing of utterance, and the movement of the mouth. This association can improve the accuracy of identifying the speaker. This allows the control unit 111 to identify the speaker from the voice even if the person in the conversation group is out of the field of view.

[0082] Conversation group modification unit 802 determines whether a person out of the field of view is continuing a conversation, based on the voice characteristics associated with the person acquired by characteristic extraction unit 801. Control unit 111 modifies the conversation groups based on this result and the detection result of conversation group detection unit 226, so that the conversation groups take into account the person out of the field of view.

[0083] In the second embodiment, a feature extraction unit 801 and a conversation group correction unit 802 are added to the block diagram shown in FIG. 2, but the operation is the same when the conversation group correction unit 802 is added to the block diagram shown in FIG. 3.

[0084] Next, the operation of the imaging device 100 of the second embodiment will be described with reference to FIGS.

[0085] Fig. 9 is a flowchart illustrating a series of recording operations of the imaging device 100. In Fig. 9, the same step numbers as in Fig. 5 are assigned to blocks that perform the same operations as in Fig. 5. Here, an example of an assumed scene for the operation in Fig. 9 will be described first with reference to Fig. 11.

[0086] Fig. 11(a) shows the scene at the time when the photographer presses the record button. In the scene shown in Fig. 11(a) (hereinafter referred to as the initial shooting scene), person A, person B, and person D (602, 603, 606) are present within the angle of view. Person A is the main subject, and three people, person A, person B, and person D, are detected as a conversation group including the main subject. Next, we will explain the scene when at least one person in the conversation group is out of the angle of view.

[0087] FIGS. 11(b) to 11(e) show examples of scenes in which a person in a conversation group moves out of the field of view. FIGS. 11(b) to 11(d) show scenes in which person B (603) moves out of the field of view 601. FIG. 11(e) shows a scene in which all members of the conversation group move out of the field of view due to the cameraman panning the imaging device 100. The horizontal "V" characters near the mouths of the people in each figure represent the vocalizations of each person, with the thickness of the lines representing the volume of their voices and the frequency of their participation in the conversation. The scenes in FIG. 11 are referred to as follows: FIG. 11(a) is the initial shooting scene, FIG. 11(b) is scene b, FIG. 11(c) is scene c, FIG. 11(d) is scene d, and FIG. 11(e) is scene e. The person 602 appearing in each figure will be referred to as person A, person 603 as person B, and person 606 as person D. The main subject of each scene is person A. In addition, angle of view 601 in each drawing indicates the imaging angle of image capture device 100, and conversation group 610 indicates the conversation group.

[0088] The assumptions for each scene in Figures 11(a) to (e) are as follows.

[0089] In scene b, person B is out of the frame of view compared to the initial shot scene, but the conversation continues as if he were in the frame of view.

[0090] In scene c, person B is out of the field of view and is not talking to the person A. In scene c, neither person A nor person D is facing person B.

[0091] Scene d shows a scene in which person B has moved into the distance compared to scene b, but the conversation continues. Note that in scene d, the voice of person B is input to the imaging device 100. Also, person A, who is within the angle of view, is facing the direction of person B, and is speaking louder.

[0092] Scene e shows a scene in which Person A, Person B, and Person D are out of the field of view compared to the initial shot scene. Note that in scene e, Person A, Person B, and Person D are continuing to have a conversation.

[0093] An example of an assumed scene in the operation of Fig. 9 has been described above with reference to Fig. 11. Hereinafter, the operation of the imaging device 100 will be described with reference to the flowchart of Fig. 9. In the description of this embodiment, steps S901 to S904 will be mainly described.

[0094] First, steps S501 to S507 are performed to detect people within the field of view, identify the main subject, detect people who are talking to the main subject, and extract voices.

[0095] In step S901, control unit 111 determines whether the number of people conversing with the main target detected in step S506 matches the number of people in the conversation group associated by feature extraction unit 801 and conversation group correction unit 802. For example, control unit 111 makes this determination by calculating the difference between the number of people in the conversation group detected in step S506 and the number of people in the current conversation group. If it is determined that the number of people has decreased, it means that there is a person in the conversation group associated by feature extraction unit 801 and conversation group correction unit 802 who is out of the field of view. If it is determined that the number of people matches, the process of step S904 is executed. If it is determined that the number of people does not match, the process of step S902 is executed.

[0096] In step S902, conversation group modification unit 802 determines whether a conversation is ongoing between a person outside the field of view and a person within the field of view. If it is determined that a conversation is not ongoing between a person outside the field of view and a person within the field of view, control unit 111 performs the processing of step S904 and subsequent steps, with the current conversation group being the result of the detection in step S506. If it is determined that a conversation is ongoing, the processing of step S903 is executed.

[0097] In step S903, control unit 111 corrects the person (subject) who is having a conversation with the main subject detected in step S506 so that the person outside the angle of view is included in a conversation group within the angle of view.

[0098] In step S904, the feature extraction unit 801 associates the sounds with the people corresponding to the sounds, based on the sounds extracted for each subject (person) extracted by the person voice extraction unit 214.

[0099] Here, an example of the determination in step S902 as to whether person B is continuing a conversation with person A and person D will be described using the above-mentioned scene.

[0100] In scene b, in steps S505 and S506 of FIG. 9, person D is identified as a person who is conversing with person A, who is the main subject. However, in step S901 of FIG. 9, it is found that person B, who belonged to the conversation group in the initially captured scene, has moved out of the angle of view. Then, in step S902 of FIG. 9, conversation group correction unit 802 determines that the person section is continuing to converse with person A and person D. Therefore, in step 903 of FIG. 9, control unit 111 adds person B to the subjects (people) who are conversing with person A, who is the main subject. That is, in scene b, the same conversation group as in the initially captured scene is maintained.

[0101] Here, an example of determining whether person B is continuing a conversation with persons A and D will be described. Conversation group correction unit 802 determines that the conversation is continuing when, based on information from feature extraction unit 801, there is no change in the volume or intonation of person B's voice and the timing of his / her speech when talking to persons A and D is synchronized. A case in which the timing of his / her speech is synchronized is, for example, when persons A and D and person B are alternately conversing (the conversation is continuing). In this case, control unit 111 adds person B to the people (subjects) conversing with person A, who is the main target. Furthermore, when image processing unit 102 can determine the direction outside the angle of view of the subject and the facial direction of the subject, conversation group correction unit 802 further determines whether the conversation is continuing based on the facial direction of person A or person D within the angle of view and the direction outside the angle of view of person B. In other words, even if it is determined that a conversation is continuing based on the volume of the voices, bathing, and timing of speech described above, if the direction of the face of person A or person D within the field of view does not match the direction in which person B has left the field of view, conversation group correction unit 802 will determine that the conversation is not continuing.

[0102] In scene c, the voice of person B is not detected in scene b. In this case, it is determined that person B is not participating in the conversation between person A and person D, and control unit 111 does not modify the subject who is conversing with person A, who is the main subject, and leaves person D as it is.

[0103] In scene d, person B has moved away from people A and D from the situation in scene b, but the conversation continues. In this scene, person B's voice has become quieter, but the timing of his speech when he is conversing with people A and D is correct. Furthermore, person B's voice has become quieter, but person A's voice, in contrast, has become louder. From this information, conversation group modification unit 802 determines that person A and person B are conversing. In response, control unit 111 adds person B as a subject conversing with person A, the main subject.

[0104] In scene e, the photographer changes the imaging direction from the direction of person A, person B, and person D toward the fireworks. That is, person A, person B, and person D continue to converse, but person A, who is the main subject, has disappeared from the angle of view. However, based on the information from the feature extraction unit 801, the acquired audio includes the voice of person A, so in this case, the control unit 111 determines that person A is the main subject. In addition, based on the information from the feature extraction unit 801, the voices of person B and person D are also continuously detected, so the control unit 111 adds person B and person D as subjects who are conversing with person A, who is the main subject. In this way, in a scene like scene e, no person is within the angle of view, but the voices of the conversation group are recorded with emphasis. Note that the control unit 111 may convert the content of the conversation into text and display it in a form such as a speech bubble, as shown in FIG. 11(f).

[0105] The operation of the imaging device 100 in the second embodiment has been described above.

[0106] The number of people conversing with the main subject detected in step S506 is based on the person detection within the field of view in step S504, so in the initial shooting scene (Figure 11(a)), there are two people, person B and person D, and in Figure 11(b) there is one person, person D.

[0107] In the second embodiment, audio extraction is performed based on the detection results of the person detection unit 222 until a predetermined time has elapsed since the start of video recording, and thereafter based on the results of the person movement detection unit 223 and information from the feature extraction unit 801.

[0108] As described above, according to the second embodiment, even if a person belonging to a conversation group moves out of the field of view, the conversation group can be corrected to an appropriate one as long as the conversation is ongoing.

[0109] 11(b), (c), and (d) in the second embodiment, the determination of conversation continuation has been explained on the assumption that the cause of a person belonging to a conversation group moving out of the angle of view is not taken into consideration, but this may also be taken into consideration. For example, if the photographer causes a person belonging to a conversation group to move out of the angle of view by operating the zoom of the lens 201, the control unit 111 may determine that the conversation is continuing without using information from the feature extraction unit 801, because the person moved out of the angle of view regardless of the photographer's intention.

[0110] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0111] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0112] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

Claims

1. a detection means for detecting a subject from a video; a determining means for determining the sound of a subject from the video; a selection means for selecting a main subject from the subjects detected from the video; a determining means for determining a subject related to the main subject selected by the selecting means, When a subject related to the main subject moves out of the angle of view of the video, the determining means determines whether the subject that moves out of the angle of view of the video continues to be related to the main subject based on the sound of the subject determined by the determining means.

1. A voice processing device comprising:

2. The determination means determines that the subject out of the field of view of the video is related to the main subject when it is determined that the voice of the main subject and the subject out of the field of view of the video are continuing a conversation, and determines that the subject out of the field of view of the video is not related to the main subject when it is determined that the voice of the main subject and the subject out of the field of view of the video are not continuing a conversation.

2. The audio processing device according to claim 1, wherein:

3. Even if it can be determined that the voice of the main subject and the subject outside the angle of view of the video are continuing to talk based on the voice of the subject determined by the determination means, if the direction of the face of the main subject and the direction of the subject outside the angle of view of the video do not match, the determination means determines that the subject outside the angle of view of the video is not continuously related to the main subject.

3. The audio processing device according to claim 1, wherein the audio processing device is a voice processing device.

4. The determination means determines that the main subject and a subject related to the main subject continue to be related to each other when it is determined that the main subject continues to have a conversation with a subject that is out of the angle of view of the video even when the main subject is out of the angle of view of the video.

4. The audio processing device according to claim 1, wherein the audio processing device is a voice processing device.

5. 5. The audio processing device according to claim 1, wherein the selection means selects a subject that is in focus in the moving image as the main subject.

6. 5. The audio processing device according to claim 1, wherein the selection unit selects a main subject from the video based on an image recorded as the main subject.

7. 5. The audio processing device according to claim 1, wherein the selection unit selects, as the main subject, a subject that appears most frequently among the subjects captured in the video.

8. 8. The audio processing device according to claim 1, wherein the determining means determines the subject related to the main subject based on the distance from the main subject.

9. 9. The audio processing device according to claim 1, wherein the determining means determines that the subject closest to the main subject is the subject related to the main subject.

10. 8. The audio processing device according to claim 1, wherein the determining means determines that an object facing the main object is an object related to the main object.

11. 8. The audio processing device according to claim 1, wherein the determining means determines an object related to the main object based on a movement of the main object.

12. an associating means for associating the subject detected by the detecting means with the sound extracted by the determining means; and image processing means, 12. The audio processing device according to claim 1, wherein the associating means associates the subject detected by the detecting means with the audio extracted by the determining means based on the audio extracted by the determining means and the movement of the subject detected by the image processing means.

13. 13. The voice processing device according to claim 12, wherein the image processing means detects the frequency of speech, the timing of speech, or the movement of the mouth of the subject.

14. 14. The audio processing device according to claim 1, wherein the determining means extracts the subject's audio based on the frequency, volume, and intonation of the audio.

15. 15. The audio processing device according to claim 1, further comprising an imaging unit for capturing the moving image.

16. 16. The audio processing device according to claim 1, wherein the determining means determines that a subject that has been detected by the detecting means has moved out of the angle of view of the video when the subject is no longer detected by the detecting means.

17. A method for controlling a voice processing device, comprising: a detection step of detecting an object from a video; a determining step of determining a subject's voice from the video; a selection step of selecting a main subject from the subjects detected from the video; a determining step of determining an object related to the main object selected in the selecting step, In the determining step, when a subject related to the main subject is out of the angle of view of the video, it is determined whether or not the subject that is out of the angle of view of the video continues to be related to the main subject, based on the sound of the subject determined in the determining step.

2. A method for controlling an audio processing device comprising:

18. A computer-readable program for causing a computer to function as each of the means of the voice processing device according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Reproduction device

    JP2010134507A

  • Video audio recorder and video audio reproducer

    JP2012138930A