Imaging device, control method, and program

The imaging device addresses the misalignment between audio and visual perception by using detection and processing means to align audio with the photographer's view, ensuring accurate reproduction of the acoustic space and conversation audio.

JP7710909B2Active Publication Date: 2025-07-22CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021112964
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-07
Publication Date
2025-07-22
Estimated Expiration
2041-07-07

AI Technical Summary

Technical Problem

Existing audio recording technologies fail to accurately reproduce the acoustic space as humans perceive it, especially in group conversations, as they focus solely on positional relationships of sound sources without considering visual cues like speaker movements and gestures.

Method used

An imaging device with detection, selection, and voice processing means to associate and process audio based on the photographer's perspective, emphasizing the main subject's voice and related conversations while suppressing others, using visual and audio cues to enhance the audio recording.

Benefits of technology

The device records audio and video aligned with the photographer's perception, ensuring the recorded audio accurately reflects the remembered conversation audio, enhancing the sense of presence and stereo effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710909000001
    Figure 0007710909000001
  • Figure 0007710909000002
    Figure 0007710909000002
  • Figure 0007710909000003
    Figure 0007710909000003
Patent Text Reader

Abstract

To record moving images and voices according to a photographer's image.SOLUTION: A voice processing apparatus has: detection means that detects subjects from moving images; selection means that selects a main subject from the subjects detected from the moving images; decision means that decides voices of the subjects from the moving images; association means that associates the subjects detected by the detection means and the voices extracted by the decision means with each other; determination means that determines a subject related to the main subject selected by the selection means; and voice processing means that performs voice processing using the voice associated with the main subject and the voice of the subject determined to be related to the main subject by the determination means, on the voices of subjects not determined to be related to the main subject by the determination means.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing device that performs audio processing on human voices.

Background Art

[0002] In video shooting with an imaging device, it is important to leave the shooting situation as the shooter imagines it, and this applies not only to video but also to audio.

[0003] Patent Document 1 discloses that by extracting the voice of a subject and individually adjusting the extracted voice signal according to the position of the subject, an acoustic space with a sense of presence and stereo effect is realized.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, when humans listen to conversations, an accurately reproduced acoustic space is not always as the human imagines. For example, even when many people are chatting, humans can naturally hear the conversations of people they are interested in or their own names. Also, it is said that humans use not only audio information but also visual information, and by visually confirming the speaker, they are said to supplement the way it sounds using information obtained from the speaker's mouth movements and gestures. That is, it is also important to record the audio recorded in the video so that it is the same as the conversation audio that remains in a person's memory (image).

[0006] However, in Patent Document 1, since the purpose is to accurately reproduce the acoustic space of the voice based on the positional relationship of a person (sound source), there is a possibility that the video may be different from the image of the photographer.

[0007] Therefore, an object of the present invention is to record a video and audio along the image of the photographer.

Means for Solving the Problems

[0008] The imaging device of the present invention includes a detection means for detecting a subject from a video, a selection means for selecting a main subject from the subject detected from the video, a determination means for determining the voice of the subject from the video, an association means for associating the subject detected by the detection means with the voice extracted by the determination means, a determination means for determining a subject related to the main subject selected by the selection means, and a voice processing for the voice associated with the main subject and the voice of the subject determined to be related to the main subject by the determination means, which is different from the voice processing for the voice of the subject not determined to be related to the main subject by the determination means. It is characterized by having a voice processing means.

Effects of the Invention

[0009] According to the present invention, it is possible to record a video and audio along the image of the photographer.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0011] Hereinafter, preferred embodiments of the present invention will be described in detail based on the accompanying drawings.

[0012] [First Embodiment] In this embodiment, the audio processing device included in the imaging device will be described with reference to FIGS. 1 to 3.

[0013] FIG. 1 is a block diagram showing the configuration of the imaging device 100 of the first embodiment.

[0014] The imaging unit 101 converts the optical image of the subject captured by the imaging optical lens into an image signal by the image sensor, and performs analog-digital conversion, image adjustment processing, etc. by the image processing unit 102 to generate image data. The imaging optical lens may be a built-in optical lens or a detachable optical lens. Also, the image sensor may be a photoelectric conversion element typified by CCD, CMOS, etc. The audio input unit 103 collects the audio around the imaging device 100 by a built-in or microphone connected via an audio terminal, and performs various audio processes on the analog-digital converted data by the audio processing unit 104 to generate audio data. The microphone may be either directional or omnidirectional. The memory 105 temporarily stores the image data obtained by the imaging unit 101 and the image processing unit 102, and the audio data obtained by the audio input unit 103 and the audio processing unit 104. The display control unit 106 causes the display unit 107 or an external display via a video terminal (not shown) to display the video related to the image data obtained by the image processing unit 102, the operation screen of the imaging device 100, the menu screen, etc. The display unit 107 has a touch panel function, and the photographer can select a menu, a subject, etc. by operating it.

[0015] The encoding processing unit 108 reads out the image data and audio data temporarily stored in the memory 105 and performs predetermined encoding to generate compressed image data, compressed audio data, etc. Also, the audio data may not be compressed. The compressed image data may be compressed by any compression method such as MPEG2, H.264 / MPEG4-AVC, etc. Also, the compressed audio data may be compressed by a compression method such as AC3(A)AC, ATRAC, ADPCM, etc. The recording / reproducing unit 109 records or reads out the compressed image data, compressed audio data or audio data, and various data generated by the encoding processing unit 108 with respect to the recording medium 110. Here, the recording medium 110 includes any type of recording medium such as a magnetic disk, an optical disk, a semiconductor memory, etc. as long as it can record image data, audio data, etc.

[0016] The control unit 111 can control each block of the imaging device 100 by transmitting control signals to each block of the imaging device 100 and the imaging unit 101, and is composed of a CPU, a memory, etc. for executing various controls. The memory 105 used by the control unit 111 is a ROM for storing various control programs, a RAM for arithmetic processing, etc., and includes an externally attached memory of the control unit 111. The operation unit 112 is composed of buttons, dials, etc., and transmits an instruction signal to the control unit 111 according to the user's operation. In the imaging device of this embodiment, it is composed of a shooting button for instructing the start and end of video recording, a zoom lever for instructing a zoom operation on an image optically or electronically, a cross key for various adjustments, a decision key, etc. The audio output unit 113 outputs the audio data reproduced by the recording and reproducing unit 109, the compressed audio data, or the audio data output by the control unit 111 to the speaker 114, the audio terminal, etc. The external output unit 115 outputs the compressed video data, the compressed audio data, the audio data, etc. reproduced by the recording and reproducing unit 109 to an external device. The data bus 116 supplies various data such as audio data and image data, and various control signals to each block of the imaging device 100.

[0017] Here, the normal operation of the imaging device 100 of this embodiment will be described.

[0018] When the user operates the operation unit 112 and an instruction to turn on the power is issued, the imaging device 100 supplies power to each block of the imaging device from a power supply unit (not shown).

[0019] When power is supplied, the control unit 111 confirms from the instruction signal from the operation unit 112 which mode the mode change switch of the operation unit 112 is in, for example, the shooting mode, the playback mode, or other modes. In the video recording mode, the image data (video data) obtained by the imaging unit 101 and the image processing unit 102 and the audio data obtained by the audio input unit 103 and the audio processing unit 104 are stored as a video file. In the playback mode, the compressed image data recorded on the recording medium 110 is reproduced by the recording and reproducing unit 109 and displayed on the display unit 107.

[0020] In the video recording mode, first, the control unit 111 transmits a control signal to each block of the imaging device 100 to shift to the shooting standby state, and causes the following operations. The imaging unit 101 converts the optical image of the subject captured by the photographing optical lens into an image signal by the image sensor, performs image adjustment processing and the like in the image processing unit 102, and generates image data. Then, the obtained image data is transmitted to the display control unit 106 and displayed on the display unit 107. The user makes shooting preparations while viewing the screen thus displayed.

[0021] The audio input unit 103 digitally converts the analog audio signals obtained by the plurality of microphones, processes the obtained plurality of digital audio signals, and generates multi-channel audio data. Then, the obtained audio data is transmitted to the audio output unit 113 and output as audio from the connected speaker 114 or an earphone (not shown). The user can also adjust the manual volume to determine the recording volume while listening to the audio output in this way.

[0022] Next, when an instruction signal for starting shooting is transmitted to the control unit 111 by the user operating the record button of the operation unit 112, the control unit 111 transmits an instruction signal for starting shooting to each block of the imaging device 100 and causes the following operations.

[0023] The imaging unit 101 converts the optical image of the subject captured by the photographing optical lens into an image signal by the image sensor, performs image adjustment processing and the like in the image processing unit 102, and generates image data. Then, the obtained image data is transmitted to the display control unit 106 and displayed on the display unit 107. Also, the obtained image data is transmitted to the memory 105.

[0024] The voice input unit 103 digitally converts the analog voice signals obtained by a plurality of microphones, processes the plurality of digital voice signals obtained by the voice processing unit 104, and generates multi-channel voice data. Then, the obtained voice data is transmitted to the memory 105. Also, when there is one microphone, the obtained analog voice signal is digitally converted to generate voice data, and the voice data is transmitted to the memory 105.

[0025] The encoding processing unit 108 reads out the image data and voice data temporarily stored in the memory 105, performs predetermined encoding, and generates compressed image data, compressed voice data, etc.

[0026] Then, the control unit 111 synthesizes these compressed image data and compressed voice data to form a data stream, and outputs it to the recording and playback unit 109. When the voice data is not compressed, the control unit 111 synthesizes the voice data stored in the memory 105 and the compressed image data to form a data stream and outputs it to the recording and playback unit 109. The recording and playback unit 109 writes the data stream as one video file to the recording medium 110 under the management of a file system such as UDF or FAT. The above operations continue during shooting.

[0027] Then, when an instruction signal to end shooting is transmitted to the control unit 111 by the user operating the recording button of the operation unit 112, the control unit 111 transmits an instruction signal to end shooting to each block of the imaging device 100 and causes the following operations to be performed.

[0028] The imaging unit 101, the image processing unit 102, the voice input unit 103, and the voice processing unit 104 each stop generating image data and voice data. The encoding processing unit 108 reads out the remaining image data and voice data stored in the memory, performs predetermined encoding, and stops operating after generating compressed image data, compressed voice data, etc. When the voice data is not compressed, of course, it stops operating after the generation of the compressed image data is completed.

[0029] Then, the control unit 111 synthesizes these last compressed image data with the compressed audio data or the audio data to form a data stream, which is output to the recording and playback unit 109. The recording and playback unit 109 writes the data stream as a single video file to the recording medium 110 under the management of a file system such as UDF or FAT. Then, when the supply of the data stream stops, the video file is completed and the recording operation is stopped. When the recording operation stops, the control unit 111 transmits a control signal to each block of the imaging device 100 to shift to the shooting standby state and returns to the shooting standby state.

[0030] Next, in the playback mode, the control unit 111 transmits a control signal to each block of the imaging device 100 to shift to the playback state and causes the following operations. The recording and playback unit 109 reads out a video file composed of the compressed image data and the compressed audio data recorded on the recording medium 110, and the read compressed image data and compressed audio data are sent to the encoding processing unit 108.

[0031] The encoding processing unit 108 decodes the compressed image data and the compressed audio data and transmits them to the display control unit 106 and the audio output unit 113, respectively. The display control unit 106 causes the decoded image data to be displayed on the display unit 107. The audio output unit 113 outputs the decoded audio data from a built-in or attached external speaker.

[0032] As described above, the imaging device 100 of this embodiment can record and play back images and audio.

[0033] In this embodiment, in the audio input unit 103 and the audio processing unit 104, when obtaining an audio signal, processing such as level adjustment processing of the audio signal obtained by the microphone is performed. This processing may be performed always after the device is started, may be performed after the shooting mode is selected, or may be performed after a mode related to audio recording is selected. Also, in a mode related to audio recording, the above processing may be performed in response to the start of audio recording. In this embodiment, the above processing is performed at the timing when moving image shooting is started.

[0034] FIG. 2 is a block diagram showing an example of the detailed configuration of the imaging unit 101, the image processing unit 102, the audio input unit 103, and the audio processing unit 104 of the imaging apparatus 100 according to the present embodiment.

[0035] The imaging unit 101 includes an optical system such as an optical lens 201 that captures an optical image of a subject, and an image sensor 202 that converts the optical image of the subject captured by the optical lens 201 into an electrical signal (image signal). Further, it has an optical lens control unit 203 having a known drive mechanism such as a position sensor and a motor for moving the optical lens 201. In the present embodiment, it is described that the optical lens 201 and the optical lens control unit 203 are built in the imaging unit 101, but these may be detachable interchangeable optical lenses. For example, when an instruction such as a zoom operation or a focus adjustment is input by the user operating the operation unit 112, the control unit 111 transmits a control signal (drive signal) for moving the optical lens to the optical lens control unit 203. The optical lens control unit 203 confirms the position of the optical lens 201 with the position sensor according to this control signal, and moves the optical lens 201 with a motor or the like.

[0036] The image processing unit 102 performs various image quality adjustment processes on the image signal converted by the image sensor 202 in the image adjustment unit 221 to form image data, and transmits it to the memory 105 via the data bus 116. Based on the image data formed here, the control unit 111 performs various adjustments such as focus adjustment and light amount adjustment.

[0037] Furthermore, in this embodiment, the image processing unit 102 has various detection functions. The person detection unit 222 extracts feature points of a person's face such as eyes, nose, and mouth from the image data formed by the image adjustment unit 221, and detects the position of the person and the size of the face in the image data. Then, by storing the information of these feature points in the memory 105, it is also possible to individually recognize the subject person based on this information. In addition, the person detection unit 222 includes a person motion detection unit 223 that detects the movement of the lips and head, and a person speech detection unit 224 that determines whether the person is speaking based on that. Further, the image processing unit 102 has a main subject selection unit 225 that selects which person among the persons detected by the person detection unit 222 will be the main subject (hereinafter also referred to as the main subject or main target) of audio processing. The main subject selection unit 225 selects the main subject based on the conditions determined by the control unit 111. The selection conditions for the main subject by the main subject selection unit 225 will be described later. Furthermore, the image processing unit 102 has a conversation group detection unit 226. The conversation group detection unit 226 detects the persons who are talking with the person selected by the main subject selection unit 225 from among the persons detected by the person detection unit 222. The detection is determined by the positional relationship between persons, the direction of the face, movements, etc. For example, the conversation group detection unit 226 determines that the subject closest to the main subject is a person (related person) who is talking with the main subject. Also, for example, the conversation group detection unit 226 determines that a subject facing the direction of the body, face, line of sight, etc. of the main subject is a person who is talking with the main subject. In addition, when the main subject is moving, the conversation group detection unit 226 determines that the subject at the tip of the moving direction of the main subject is a person who is talking with the main subject. This is because such a subject is considered to talk with the main subject in the near future.

[0038] In addition, when the conversation group detection unit 226 determines that a person who is conversing with the main subject has not conversed with the main subject for a period longer than a predetermined time, that person is considered a person who is not (unrelated to) conversing with the main subject. In other words, if the person who is conversing with the main subject is within the predetermined time, even if it is determined that the person is not conversing with the main subject, the person is determined to be a person who is conversing with the main subject.

[0039] Next, the voice input unit 103 and the voice processing unit 104 will be described. The voice input unit 103 is a microphone 211 that converts voice vibrations into electrical signals and outputs them as voice signals. In the present embodiment, the microphone 211 is a stereo system configured with two channels of left and right Lch / Rch, but it may also be a monaural system with one channel or a configuration that holds a plurality of microphones with two or more channels. The A / D conversion unit 212 is a means for converting the analog voice signal obtained by the microphone 211 into a digital voice signal.

[0040] The voice processing unit 104 is a block that performs various voice processes on the voice signal converted by the voice input unit 103. In the present embodiment, the voice processing unit 104 includes a voice extraction unit 213, a voice adjustment unit 215, and a voice synthesis unit 217. The voice extraction unit 213 can extract (determine) into a person's voice and other voices (hereinafter referred to as "non-person voices"). Furthermore, the person voice extraction unit 214 can extract a person's voice into individual voices based on the information of the person detection unit 222. For example, the person voice extraction unit 214 extracts into individual voices based on the frequency, volume, and intonation of the voice. Furthermore, in the first embodiment, the control unit 111 can associate the subject with the voice based on the voice extracted by the person voice extraction unit 214 and the action of the subject detected by the image processing unit 102. For example, the action of the subject is the frequency of speech, the timing of voice production, and the movement of the mouth.

[0041] Also, in the voice adjustment unit 215, voice processing for each frequency band such as level adjustment and equalizer can be individually performed on the voice extracted by the voice extraction unit 213. In particular, in the conversation voice adjustment unit 216, adjustment is performed based on the information of the conversation group detection unit 226, and the extracted voice is made easier to hear and emphasized, or made less audible and subdued. The details of the adjustment will be described later. Further, in the voice synthesis unit 217, the voices individually adjusted in the voice adjustment unit 215 are synthesized and returned to one voice signal again. Then, the amplitude of the synthesized voice signal is adjusted to a predetermined level by an auto level controller (hereinafter, ALC 219). With the above configuration, the voice processing unit 104 performs predetermined processing on the voice signal, forms voice data, and transmits it to the memory 105.

[0042] FIG. 3 is a block diagram showing an example of another configuration of the image processing unit 102 and the voice processing unit 104 of the imaging device 100 according to the present embodiment. The difference between FIG. 3 and FIG. 2 is that the input sources of the image data and the voice data are different. In FIG. 2, the image signal uses the signal from the imaging unit 101, and the voice signal uses the signal from the voice input unit 103. On the other hand, in FIG. 3, the input sources of the image and the voice input the data stored in the memory 105. By using the data once stored (held) in the memory 105 in this way, it becomes possible to use the proposed method not only for the processing at the time of shooting but also for the post-processing after recording. Also, in the main subject selection unit 225, it becomes possible to select the person to be subjected to voice processing from a series of moving image data.

[0043] Here, an example of the method for selecting the main subject by the main subject selection unit 225 will be described with reference to FIG. 4. In the present embodiment, the main subject will be described as a person who is considered to be the focus of the photographer. For example, in the case of FIG. 4(a), the focus mark 402 is a mark indicating the subject on which the imaging device 100 is focusing. In FIG. 4(a), since the main subject 401 and the focus mark 402 coincide, the imaging device 100 recognizes the main subject 401 as the main subject and focuses on the main subject 401. The main subject selection unit 225 determines this main subject 401 as the main subject. In this way, a person recognized as the main subject can be selected as the main subject.

[0044] Also, FIG. 4(b) shows a method using a registered face image. The registered face image 403 is an image of a subject previously registered in the memory 105. The main subject selection unit 225 selects, as the main subject, a person determined to match the face of that image.

[0045] Also, FIG. 4(c) shows a method of determining the main subject according to the intention of the photographer. For the person displayed on the display unit 107, the photographer touches the touch panel of the display unit 107 to select the subject to be the main subject. The main subject selection unit 225 determines the subject selected by the photographer as the main subject.

[0046] Also, FIG. 4(d) shows a method using recorded video data. For example, when recorded video data 404 is recorded in the memory 105, the main subject selection unit 225 determines the person 405 with the highest appearance frequency in the video data 404 as the main subject. Additionally, for example, the main subject selection unit 225 may select a person with a high focus frequency.

[0047] Note that when the main subject selection unit 225 sets, for example, the subject on which the focus is set as the main subject, even if the focus on that main subject is lost, if the focus returns to that subject within a predetermined time, the main subject is maintained. In other words, when the focus is off the main subject for a longer time than the predetermined time, the main subject selection unit 225 selects a new subject to be the main subject.

[0048] Next, the operation of the imaging device 100 of the present embodiment will be described with reference to FIGS. 5 to 7.

[0049] FIG. 5 is a flowchart showing an example of a series of recording sequences of the imaging device 100. The processing of the imaging device 100 is realized by expanding software recorded in a ROM (not shown) into the memory 105 and executed by the CPU. Also, the processing of this flowchart is started triggered by the imaging device 100 being powered on.

[0050] In step S501, the control unit 111 receives an instruction to start video recording by an operation of the operation unit 112 by the user.

[0051] In step S502, the control unit 111 connects the audio path for audio recording.

[0052] In step S503, after the audio path is established, the control unit 111 performs initial settings of signal processing including the control described in the present embodiment, and starts signal processing for video recording. Hereinafter, the recording sequence will be described focusing on it. Until the signal processing for video recording ends, the control unit 111 records the video recorded in the video.

[0053] In step S504, the person detection unit 222 of the image processing unit 102 detects a subject.

[0054] In step S505, the main subject selection unit 225 of the image processing unit 102 selects (judges) the main subject from the subjects detected in step S504.

[0055] In step S506, the conversation group detection unit 226 of the image processing unit 102 determines the person (subject) who is talking with the main subject selected in step S505.

[0056] In step S507, the audio extraction unit 213 of the audio processing unit 104 extracts the person's voice.

[0057] The voice adjustment unit 215 of the voice processing unit 104 performs adjustment processing on the voice extracted in step S507. The content of the voice adjustment processing is made different depending on whether the subject (person) of the voice extracted in step S507 belongs to the subject (person) of the conversation group that is the main subject. Details of the voice adjustment processing will be described later with reference to FIGS. 6 and 7, but will be simply explained in this flowchart.

[0058] In step S508, the voice adjustment unit 215 of the voice processing unit 104 determines whether the person of the voice extracted in step S507 belongs to the subject of the conversation group that is the main subject. If the person of the extracted voice belongs to the subject of the conversation group that is the main subject, the process of step S509 is executed. If the person of the extracted voice does not belong to the subject of the conversation group that is the main subject, the process of step S510 is executed.

[0059] In step S509, the voice adjustment unit 215 of the voice processing unit 104 adjusts the level so that the volume of the extracted voice becomes larger.

[0060] In step S510, the voice adjustment unit 215 of the voice processing unit 104 adjusts the level so that the volume of the extracted voice becomes smaller. In step S511, the voice adjustment unit 215 of the voice processing unit 104 performs adjustment processing other than volume on the extracted voice.

[0061] In step S512, the voice synthesis unit 217 of the voice processing unit 104 synthesizes the individually voice-adjusted extracted voices to generate one voice data.

[0062] In step S513, the control unit 111 determines whether to end the video recording. For example, when the control unit 111 is instructed to end the video recording by an operation of the operation unit 112 by the user, or when it is determined that the remaining capacity of the recording medium 110 is low, the control unit 111 determines to end the video recording. If it is determined not to end the video recording, the process returns to step S504, and the recording sequence process continues. If it is determined to end the video recording, the process of step S514 is executed.

[0063] Here, if it is determined not to end the video recording, the process returns to step S504. That is, during video recording, the main subject and the person conversing with the main subject are repeatedly determined. Thereby, for example, even when the subject that is the main subject disappears outside the picture angle or the focus is lost, the control unit 111 can determine another subject as the main subject. Also, even when the number of people conversing with the main subject increases or decreases, the control unit 111 can determine the person conversing with the main subject accordingly.

[0064] In step S514, the control unit 111 cuts off the audio path and ends the signal processing.

[0065] Here, the audio adjustment process will be described with reference to FIGS. 6 and 7.

[0066] FIG. 6 is a diagram showing an assumed scene of the audio adjustment process. Now, assume that four subjects (persons) of persons 602 to 605 exist within the picture angle 601, and person 602 is conversing (speaking) with person 603, and person 604 is conversing with person 605. At this time, if the main subject of the audio processing selected by the main subject selection unit 225 is person 602, then person 602 and person 603 are detected as conversation group 610 by the conversation group detection unit 226 from the image data. In this case, the voices of persons 602 and 603 are adjusted in audio so as to be emphasized as the voices to be noted, and the voices of persons 604 and 605 are adjusted in audio as unnecessary voices that are not the emphasis target.

[0067] Figs. 7(a) to 7(c) are diagrams showing voice adjustment processing. In Fig. 7, the persons 602, 603, and 604 in Fig. 6 are denoted as persons A, B, and C, respectively (person 605 is not shown).

[0068] Fig. 7(a) shows the respective voice signals of persons A to C extracted by the person voice extraction unit 214. That is, signal 701 indicates the voice signal extracted for person A, signal 702 indicates the voice signal extracted for person B, and signal 703 indicates the voice signal extracted for person C. And in each signal, a section with a large amplitude indicates the period (voiced timing) during which each person is speaking (uttering sound), and a section with a small amplitude indicates the period during which the person is not speaking (unvoiced timing). For example, when comparing signal 704 and signal 705, since persons A and B are having a conversation, the voiced timing and the unvoiced timing appear almost alternately. On the other hand, since person C is not a party to the conversation between persons A and B, in signal 706, the voiced timing and the unvoiced timing do not appear alternately as often as in signals 704 and 705.

[0069] Fig. 7(b) shows the voice correction coefficients for each person. In the present embodiment, when the correction coefficient is 1.0, it indicates that no level adjustment (gain adjustment) is performed. Also, the processing when the correction coefficient is greater than 1.0 is voice adjustment processing for emphasizing the voice to make it easier to hear (increasing the volume), and the processing when the coefficient is less than 1.0 is processing for making the voice less audible (decreasing the volume).

[0070] For example, a case where the conversation group detection unit 226 determines that person A and person B are having a conversation during period 710 will be described as an example. In this case, since person A is the main subject, the conversation voice adjustment unit 216 recognizes the voices of person A and person B respectively as targets to be emphasized, and sets the correction coefficients for their respective voices to large values (coefficient 714, coefficient 715). In the present embodiment, the correction coefficients for the voices of person A and person B are set to the same value. This is because it is assumed that the user, who is the photographer, hears both voices equally. On the other hand, the conversation voice adjustment unit 216 sets a small correction coefficient for the voice of person C who is determined not to be having a conversation with person A, making it relatively difficult for the user to hear the voice of person C (coefficient 716). In this way, the conversation voice adjustment unit 216 emphasizes the voices of the main subject person A and person B who is the conversation partner, and makes the other voices smaller. For example, the conversation voice adjustment unit 216 makes the gain and level for the voices of the main subject person A and person B who is the conversation partner larger than those for the other voices. As a result, the video and audio become moving image data that conforms to the image of the user who is the photographer.

[0071] And FIG. 7(c) shows the audio signal adjusted based on the correction coefficients of FIG. 7(b) described above. For example, when the audio adjustment by the conversation voice adjustment unit 216 is realized by gain adjustment, during period 710, the voices of person A and person B determined to be having a conversation (signal 724, signal 725) have correction coefficients greater than 1.0, so the volume becomes larger and it is easier for the user to hear. Also, the voice of person C (signal 726) for whom conversation is not determined has a correction coefficient smaller than 1.0, so the volume becomes smaller and it is difficult to hear. By synthesizing the individually adjusted extracted voices in the voice synthesis unit 217, as a result, only the conversation determined as the target of attention is generated as audio data that is easy to hear.

[0072] In this embodiment, the voice related to the main subject is emphasized (corrected to be louder), and the voice not related to the main subject is made inaudible (corrected to be quieter). However, adjustments may be applied to only one of them. That is, the correction coefficients of the main subject (person) and the subject (person) who is the conversation partner of the main subject only need to be larger than the correction coefficients of other subjects.

[0073] Also, the emphasis method by the conversation voice adjustment unit 216 is not limited to the adjustment of the overall gain as described above, and may be adjusted by frequency for each frequency band of the human voice using an equalizer or the like.

[0074] [Second Embodiment] In the first embodiment, after selecting the main subject, the people who are talking to the main subject are detected as a conversation group based on the positional relationship with the main subject and the actions of the people, and the voice of the conversation group is emphasized, or other unnecessary voices are suppressed to obtain voice data in which the conversation to be noted is easy to hear.

[0075] In the first embodiment, the method for detecting the conversation group is determined by the positional relationship, face orientation, actions, etc. of the people who are talking to the person selected by the main subject selection unit 225 among the people detected by the person detection unit 222. In this way, in the first embodiment, the detection by the conversation group detection unit 226 is performed by the people existing within the viewing angle 601 of the imaging device 100.

[0076] Now, assume that as shown in Fig. 10(a), the person A who is the main subject 602 and the people B and D (603, 606) within the viewing angle 601 are detected as a conversation group. If the person B is removed from the viewing angle due to a zoom operation or panning operation by the photographer, even if the conversation among the people A, B, and D continues, in the next detection of the conversation group, the person B will be excluded from the conversation group as shown in Fig. 10(b). As a result, even if the person B participates in the conversation, the conversation group detection unit 226 does not determine the person B as a conversation group, so there is a possibility that only the voice of the person B is not emphasized and the conversation is difficult to hear.

[0077] In the second embodiment, when at least one person in the conversation group within the shooting angle moves out of the shooting angle and it is determined that the conversation of the person who has moved out of the shooting angle continues, the conversation group is maintained in the state before moving out of the shooting angle, and the purpose is to continue to obtain easily audible voices.

[0078] Hereinafter, the second embodiment will be described in detail with reference to the accompanying drawings. Since the configuration of the imaging device 100 in FIG. 1 is the same as that in the first embodiment, the description thereof will be omitted.

[0079] FIG. 8 is a block diagram showing the detailed configuration of the imaging unit 101, the image processing unit 102, the audio input unit 103, and the audio processing unit 104 of the imaging device 100 according to the present embodiment. Blocks having the same functions as those in FIG. 2 are assigned the same numbers, and the description thereof will be omitted.

[0080] The feature extraction unit 801 associates the voice extracted by the person voice extraction unit 214 with the person corresponding to the voice. For example, the feature extraction unit 801 associates the extracted voice with the corresponding person based on the features of the voice and the actions of the subjects within the shooting angle. For example, the above-mentioned features of the voice are frequency, magnitude, and intonation. For example, the actions of the subjects are the frequency of speaking, the timing of vocalization, and the movement of the mouth. By such association, the accuracy for identifying the speaker can be improved. Thereby, the control unit 111 can identify the speaker from the voice even if the person in the conversation group moves out of the shooting angle.

[0081] The conversation group correction unit 802 determines whether the person who has moved out of the shooting angle continues to talk based on the features of the voice associated with the person obtained by the feature extraction unit 801. The control unit 111 corrects the conversation group considering the person who has moved out of the shooting angle based on this result and the detection result of the conversation group detection unit 226.

[0082] In the second embodiment, the feature extraction unit 801 and the conversation group correction unit 802 are described in a form added to the block diagram shown in FIG. 2. However, the operation content is the same even in a form where the conversation group correction unit 802 is added to the block diagram shown in FIG. 3.

[0083] Next, the operation of the imaging device 100 of the second embodiment will be described with reference to FIGS. 9 and 11.

[0084] FIG. 9 is a flowchart for explaining a series of recording operations of the imaging device 100. In FIG. 9, the same step numbers as those in FIG. 5 are assigned to the blocks that perform the same operations as those in FIG. 5. Here, first, an example of an assumed scene in the operation of FIG. 9 will be described with reference to FIG. 11.

[0085] FIG. 11(a) shows a scene at the time when the recording button is pressed by the photographer. In the scene shown in FIG. 11(a) (hereinafter referred to as the initial shooting scene), persons A, B, and D (602, 603, 606) are present within the angle of view. Assuming the main subject is person A, the conversation group including the main subject is detected as consisting of three persons: persons A, B, and D. Then, a scene where at least one person in the conversation group moves out of the angle of view will be described.

[0086] Examples of scenes where a person included in the conversation group moves out of the angle of view are shown in FIGS. 11(b) to (e). FIGS. 11(b) to (d) show scenes where person B (603) moves out of the angle of view 601. FIG. 11(e) shows a scene where all members of the conversation group move out of the angle of view due to the panning operation of the imaging device 100 by the photographer. Also, the horizontal "C" characters shown near the mouths of the persons in each figure represent the speaking states of the respective persons, and the thickness of the lines represents the volume and the degree of frequency of participation in the conversation. Also, each scene in FIG. 11 is described as follows: FIG. 11(a) is the initial shooting scene, FIG. 11(b) is scene b, FIG. 11(c) is scene c, FIG. 11(d) is scene d, and FIG. 11(e) is scene e. Also, person 602 appearing in each figure is described as person A, person 603 as person B, and person 606 as person D. Also, the main subject of each scene is person A. Also, the angle of view 601 in each figure is the shooting angle of view of the imaging device 100, and the conversation group 610 indicates the conversation group.

[0087] The assumptions for each scene in FIGS. 11(a) to (e) are as follows.

[0088] In scene b, a scene is shown where, with respect to the initial shooting scene, although person B is out of the shooting angle, the conversation continues in the same way as when person B is within the shooting angle.

[0089] In scene c, a scene is shown where, with respect to the initial shooting scene, person B is out of the shooting angle and not having a conversation. In scene c, neither person A nor person D is facing person B.

[0090] In scene d, a scene is shown where, with respect to the scene of scene b, person B is moving away to a distance but the conversation continues. In scene d, the voice of person B is input to the imaging device 100. Also, the face direction of person A within the shooting angle is facing the direction where person B is, and the vocal volume has increased.

[0091] In scene e, a scene is shown where, with respect to the initial shooting scene, persons A, B, and D are out of the shooting angle. In scene e, persons A, B, and D continue the conversation.

[0092] Above, the assumed scene examples in the operation of FIG. 9 have been described with reference to FIG. 11. Hereinafter, the operation of the imaging device 100 will be described using the flowchart of FIG. 9. In the description of this embodiment, mainly steps S901 to S904 will be described.

[0093] First, by the processes from step S501 to step S507, detection of a person within the shooting angle, identification of the main subject, detection of the person having a conversation with the main subject, and extraction of voice are carried out.

[0094] In step S901, the control unit 111 determines whether the number of people having a conversation with the main subject detected in step S506 matches the number of people in the conversation group associated by the feature extraction unit 801 and the conversation group modification unit 802. For example, the control unit 111 makes a determination by taking the difference between the number of people in the conversation group detected in step S506 and the number of people in the current conversation group. If it is determined that the number of people has decreased, among the people in the conversation group associated by the feature extraction unit 801 and the conversation group modification unit 802, there will be a person who is out of the viewing angle. If it is determined that the numbers match, the process of step S904 is executed. If it is determined that the numbers do not match, the process of step S902 is executed.

[0095] In step S902, the conversation group modification unit 802 determines whether the conversation between the person outside the viewing angle and the person inside the viewing angle continues. If it is determined that the conversation between the person outside the viewing angle and the person inside the viewing angle does not continue, the control unit 111 performs the processes after step S904 with the current conversation group as the detection result in step S506. If it is determined that the conversation continues, the process of step S903 is executed.

[0096] In step S903, the control unit 111 modifies the person (subject) having a conversation with the main subject detected in step S506 so that the person outside the viewing angle is included in the conversation group inside the viewing angle.

[0097] In step S904, the feature extraction unit 801 associates the voice with the person corresponding to the voice based on the voice extracted for each subject (person) extracted by the person voice extraction unit 214.

[0098] Here, using the above-described scene, an example of the determination in step S902 as to whether person B continues the conversation with person A and person D will be described.

[0099] In scene b, in steps S505 and S506 of FIG. 9, person D is identified as the person who is talking to person A, the main subject. However, in step S901 of FIG. 9, it can be seen that person B, who belonged to the conversation group in the initial shooting scene, has moved out of the field of view. Then, in step S902 of FIG. 9, the conversation group correction unit 802 determines that the person unit is continuing the conversation with person A and person D. Therefore, in step 903 of FIG. 9, the control unit 111 adds person B to the subject (person) who is talking to person A, the main subject. That is, in scene b, the same conversation group as in the initial shooting scene is maintained.

[0100] Here, an example of the determination of whether person B's conversation with person A and D continues will be described. The conversation group correction unit 802 determines that the conversation is continuing when, based on the information from the feature extraction unit 801, there is no change in the loudness or intonation of person B's voice and the speaking timing during the conversation with person A and D is appropriate. In this case, the control unit 111 adds person B to the person (subject) who is talking to person A, the main subject. Further, when the image processing unit 102 can determine the direction in which the subject has moved out of the field of view and the direction of the subject's face, the conversation group correction unit 802 further determines whether the conversation is continuing based on the direction of the face of person A or person D within the field of view and the direction in which person B has moved out of the field of view. That is, even if it is determined that the conversation is continuing based on the above-mentioned loudness, intonation, and speaking timing, if the direction of the face of person A or person D within the field of view does not match the direction in which person B has moved out of the field of view, the conversation group correction unit 802 determines that the conversation is not continuing.

[0101] In scene c, compared to scene b, the voice of person B is not detected. In such a case, it is determined that person B is not participating in the conversation between person A and person D, and the control unit 111 leaves the subject who is talking to person A, the main subject, as person D without making any corrections.

[0102] In scene d, from the situation of scene b, person B moves away from persons A and D, but the conversation continues. In this scene, the voice of person B becomes quieter, but the speaking timing during the conversation with persons A and D is appropriate. Also, although the voice of person B becomes quieter, the voice of person A becomes louder instead. From this information, the conversation group correction unit 802 determines that persons A and B are having a conversation. Accordingly, the control unit 111 adds person B as a subject having a conversation with person A, who is the main subject.

[0103] In scene e, it is a scene where the photographer changes the imaging direction from the direction where persons A, B, and D are to the direction of the fireworks display. That is, persons A, B, and D continue the conversation, but person A, who is the main subject, has also disappeared from the picture angle. However, from the information of the feature extraction unit 801, since the voice of person A is included in the acquired voice, in this case, the control unit 111 determines that person A is the main subject. In addition, from the information of the feature extraction unit 801, since the voices of persons B and D are also continuously detected, the control unit 111 adds persons B and D as subjects having a conversation with person A, who is the main subject. Thus, in a scene like scene e where no one is within the picture angle, the voice of the conversation group is emphasized and recorded. Note that the control unit 111 may control to convert the content of the conversation into text and display it in a form such as a speech bubble as shown in Fig. 11(f).

[0104] The operation of the imaging device 100 in the second embodiment has been described above.

[0105] Note that the number of persons having a conversation with the subject detected as the main subject in step S506 is based on the person detection in step S504 within the picture angle. Thus, in the initial shooting scene (Fig. 11(a)), there are two persons, B and D, and in Fig. 11(b), there is one person, D.

[0106] Note that the voice extraction in the second embodiment is executed based on the detection result of the person detection unit 222 until a predetermined time has elapsed since the start of video recording, and then based on the result of the person detection 223 and the information of the feature extraction unit 801.

[0107] As described above, according to the second embodiment, even when a person belonging to the conversation group is out of the shooting angle, if the conversation is continuing, it can be corrected to an appropriate conversation group.

[0108] Regarding the determination of conversation continuation in FIGS. 11(b), (c), and (d) of the second embodiment, it has been described on the premise of not considering the factor that a person belonging to the conversation group has deviated from the shooting angle, but this may also be considered. For example, when the person belonging to the conversation group has deviated from the shooting angle due to the zoom operation of the lens 201 by the photographer, since the person has deviated from the shooting angle regardless of their own intention, the control unit 111 may determine that the conversation is continuing without using the information of the feature extraction unit 801.

[0109] As described above, the preferred embodiments of the present invention have been described, but the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.

[0110] [Other Embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0111] Note that the present invention is not limited to the above-described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

Claims

1. Detection means for detecting a subject from a video, Selection means for selecting a main subject from the subjects detected from the video, Determination means for determining the voice of the subject from the video, Association means for associating the subject detected by the detection means with the voice extracted by the determination means, Judgment means for judging a subject related to the main subject selected by the selection means, Voice processing means for making the voice processing for the voice associated with the main subject and the voice of the subject judged to be related to the main subject by the judgment means different from the voice processing for the voice of the subject not judged to be related to the main subject by the judgment means A voice processing apparatus characterized by comprising the same.

2. The voice processing means is characterized in that the level adjustment for the voice associated with the main subject and the voice of the subject judged to be related to the main subject by the judgment means is made different from the level adjustment for the voice of the subject not judged to be related to the main subject by the judgment means. The voice processing apparatus according to claim 1.

3. The voice processing means is characterized in that the correction coefficient for the voice associated with the main subject and the voice of the subject judged to be related to the main subject by the judgment means is made larger than the correction coefficient for the voice of the subject not judged to be related to the main subject by the judgment means. The voice processing apparatus according to claim 1 or 2.

4. The voice processing means is characterized in that the gain for the voice associated with the main subject and the voice of the subject judged to be related to the main subject by the judgment means is made larger than the gain for the voice of the subject not judged to be related to the main subject by the judgment means. The voice processing apparatus according to claim 1 or 2.

5. The selection means is characterized in that the subject in focus in the video is selected as the main subject. The voice processing apparatus according to any one of claims 1 to 4.

6. The selection means is characterized in that the main subject is selected from the video based on the image recorded as the main subject. The voice processing apparatus according to any one of claims 1 to 4.

7. The voice processing apparatus according to any one of claims 1 to 4, wherein the selection means selects, as a main subject, the subject having the highest appearance frequency among the subjects imaged in the moving image.

8. The voice processing apparatus according to any one of claims 1 to 7, wherein the determination means determines a subject related to the main subject based on the distance from the main subject.

9. The voice processing apparatus according to any one of claims 1 to 8, wherein the determination means determines the subject closest to the main subject as the subject related to the main subject.

10. The voice processing apparatus according to any one of claims 1 to 7, wherein the determination means determines a subject facing the main subject as the subject related to the main subject.

11. The voice processing apparatus according to any one of claims 1 to 7, wherein the determination means determines a subject related to the main subject based on the action of the main subject.

12. Further comprising image processing means, The voice processing apparatus according to any one of claims 1 to 11, wherein the association means associates the subject detected by the detection means with the voice extracted by the determination means based on the voice extracted by the determination means and the action of the subject detected by the image processing means.

13. The voice processing apparatus according to claim 12, wherein the image processing means detects the frequency of speech, the timing of voice production, or the movement of the mouth of the subject.

14. The voice processing apparatus according to any one of claims 1 to 13, wherein the determination means extracts the voice of the subject based on the frequency, magnitude, and intonation of the voice.

15. The voice processing apparatus according to any one of claims 1 to 14, further comprising imaging means for imaging the moving image.

16. A control method for a voice processing apparatus, comprising: A detection step of detecting a subject from a moving image; A selection step of selecting a main subject from the subjects detected from the moving image; An extraction step of extracting the voice of the subject from the moving image; An association step of associating the subject detected in the detection step with the voice extracted in the extraction step; A determination step of determining a subject related to the main subject selected in the selection step; An audio processing step of performing audio processing on the audio associated with the main subject and the audio of the subject determined to be associated with the main subject in the determination step, with respect to the audio of the subject not determined to be associated with the main subject in the determination step A control method for an audio processing apparatus, characterized by comprising the above.

17. A computer-readable program for causing a computer to function as each means of the audio processing apparatus according to any one of Claims 1 to 15.

Citation Information

Patent Citations

  • Audio processing system

    JP2012029209A

  • Video audio recorder and video audio reproducer

    JP2012138930A

  • Customer service monitoring device, customer service monitoring system, and customer service monitoring method

    JP2016177664A

  • Information processing equipment, information processing method, and program

    JP2021033573A

  • Imaging apparatus

    JP2021082968A