Imaging system, display method, and program

The imaging system effectively separates and displays audio data from multiple sound sources, addressing the challenge of monitoring individual sound volumes in noisy environments by adjusting gain levels and providing visual feedback.

WO2025220283A1PCT designated stage Publication Date: 2025-10-23SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000244
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2025-01-08
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing methods for monitoring audio during video recording, such as using headphones or audio level meters, are inadequate for determining the volume of individual sound sources, especially in noisy environments, making it difficult to distinguish target sounds from ambient noise.

Method used

An imaging system that includes an image acquisition unit, recognition unit, audio acquisition unit, audio analysis unit, and display control unit, which separates audio data into multiple sound sources and displays their volumes in association with image data, allowing for precise monitoring of target sounds.

Benefits of technology

Enables accurate monitoring of target sound volumes, facilitating better audio recording by adjusting gain levels and providing visual feedback on sound source separation, thereby improving the quality of recorded audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000244_23102025_PF_FP_ABST
    Figure JP2025000244_23102025_PF_FP_ABST
Patent Text Reader

Abstract

An imaging system disclosed herein comprises an image acquisition unit, a recognition unit, a voice acquisition unit, a voice analysis unit, and a display control unit. The image acquisition unit acquires image data including an object of an utterance subject. The recognition unit recognizes an object included in an image of the image data. The voice acquisition unit acquires voice data related to the object of the utterance subject. The voice analysis unit separates the voice data into a plurality of pieces of voice data. The display control unit causes a display unit to display information on the voice data, emitted from the object of the utterance subject, in association with the image data on the basis of the result of the recognition by the recognition unit and the result of the separation by the voice analysis unit.
Need to check novelty before this filing date? Find Prior Art

Description

Imaging system, display method and program

[0001] The present disclosure relates to an imaging system, a display method, and a program.

[0002] When shooting video with a camera, a microphone collects surrounding sounds during recording and records the collected sounds along with the video. One way to check whether the desired audio is being properly recorded during video shooting is for the cameraperson to wear headphones and listen to the camera's monitoring sound while recording. Another method is to check the audio by looking at an audio level meter, which displays the volume during recording. An audio level meter is a visual measurement display that shows the instantaneous level of an audio signal, and generally indicates the current peak volume of the audio data in units of dB (decibels).

[0003] JP 2009-118318 A

[0004] However, monitoring audio by ear requires a high level of skill, and it becomes even more difficult when a single cameraman simultaneously monitors the video. Furthermore, the volume displayed on the audio level meter is the overall volume of the audio, and it is not possible to determine the volume of each individual sound source. For example, when recording in a noisy environment, such as a construction site, it is difficult to determine whether fluctuations in the meter level are due to ambient noise or the target sound being recorded (e.g., talking), and the volume of the target sound may not be at an appropriate level.

[0005] Therefore, the present disclosure proposes an imaging system, a display method, and a program that can grasp the volume of the voice of a speaking object included in collected voice.

[0006] In order to solve the above problems, an imaging system according to the present disclosure includes an image acquisition unit, a recognition unit, an audio acquisition unit, an audio analysis unit, and a display control unit. The image acquisition unit acquires image data including an object to be spoken to. The recognition unit recognizes an object included in the image of the image data. The audio acquisition unit acquires audio data related to the object to be spoken to. The audio analysis unit separates the audio data into multiple audio data. The display control unit displays information related to the audio data emitted from the object to be spoken to on the display unit in association with the image data, based on the recognition result by the recognition unit and the sound source separation result by the audio analysis unit.

[0007] 1 is a functional block diagram showing an example of the functional configuration of an imaging device according to an embodiment. FIG. 1 is a functional block diagram showing an example of the functional configuration of sound source separation processing in an audio analysis unit according to an embodiment. FIG. 2 is a block diagram explaining an example of sound source separation by an audio analysis unit according to an embodiment. FIG. 3 is a diagram showing an example of a display of a display unit according to an embodiment. FIG. 4 is a flowchart showing an example of a processing procedure of an imaging device according to an embodiment. FIG. 5 is a diagram explaining a gain adjustment method according to a comparative example. FIG. 6 is a diagram explaining a gain adjustment method according to an embodiment. FIG. 7 is a flowchart showing an example of a processing procedure of an imaging device according to an embodiment. FIG. 8 is a flowchart showing an example of a processing procedure of an imaging device according to a first modified example. FIG. 9 is a flowchart showing an example of a processing procedure of an imaging device according to a second modified example. FIG. 10 is a diagram showing an example of a display of a display unit according to a third modified example. FIG. 11 is a diagram showing an example of a display of a display unit according to the third modified example. FIG. 12 is a diagram showing an example of a configuration of an imaging system according to a fifth modified example. FIG. 13 is a functional block diagram showing an example of the functional configuration of a monitoring device according to the fifth modified example. FIG. 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of an imaging device.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted. In addition, the following description will be given taking as an example a case where the technology of the present disclosure is applied to an imaging device 10 such as a camera.

[0009] The present disclosure will be described in the following order: 1. Configuration of imaging device 2. Processing procedure 3. Specific example of imaging 3-1. Gain adjustment method of comparative example 3-2. Gain adjustment method of embodiment 4. Processing procedure 5. Modifications 5-1. First modification 5-1-1. Processing procedure 5-2. Second modification 5-2-1. Processing procedure 5-3. Third modification 5-4. Fourth modification 5-5. Fifth modification 5-5-1. Configuration of imaging system 5-5-2. Configuration of monitoring device 6. Hardware configuration 7. Conclusion

[0010] 1. Configuration of the Imaging Device FIG. 1 is a functional block diagram showing an example of the functional configuration of an imaging device 10 according to an embodiment. Note that FIG. 1 illustrates only components necessary for explaining the features of the embodiment, and general components are omitted. In other words, the components illustrated in FIG. 1 are functional concepts and do not necessarily have to be physically configured as illustrated. For example, the specific form of distribution and integration of each block is not limited to that illustrated, and all or part of the blocks can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0011] The imaging device 10 is a device capable of capturing images, such as a mirrorless camera, a digital single-lens reflex camera, a video camera, etc. The imaging device 10 may also be a portable information processing device having a camera function, such as a smartphone.

[0012] As shown in FIG. 1, the imaging device 10 has a display unit 20, a video recording unit 21, an optical unit 22, an imaging unit 23, an image acquisition unit 24, a subject recognition unit 25, an audio acquisition unit 26, an audio analysis unit 27, a display control unit 28, and an operation control unit 29.

[0013] The display unit 20 is, for example, a display provided in the imaging device 10 .

[0014] The imaging device 10 is provided with a card slot into which a memory card is inserted. The video recording unit 21 reads data stored in the memory card and writes data to the memory card. The memory card is a storage medium that incorporates nonvolatile semiconductor memory such as flash memory and allows data to be rewritten. The memory card retains the data written to it. Examples of memory cards include SD memory cards (Secure Digital memory cards).

[0015] The optical unit 22 is, for example, an imaging lens provided in the imaging device 10. The optical unit 22 includes a focus lens, a lens mechanism that moves the focus lens, a shutter mechanism that drives an optical shutter, and an iris mechanism that drives an aperture to adjust the amount of light passing through the lens. The optical unit 22 adjusts the focus by moving the focus lens using the lens mechanism, and collects incident light.

[0016] The imaging unit 23 is provided with an image sensor. The image sensor is arranged with its optical axis aligned with that of the optical unit 22. An image condensed by the optical unit 22 is formed on the image sensor. The imaging unit 23 captures an image using the image sensor and generates image data of the captured image. For example, the imaging unit 23 converts the light condensed on the image sensor into an electrical signal, outputs the signal, and converts and processes the electrical signal into a digital image signal to generate digital image data.

[0017] The image acquisition unit 24 controls the imaging unit 23 and acquires image data by the imaging unit 23. The image acquisition unit 24 acquires image data including the object to be spoken to by the imaging unit 23. When displaying a live view image on the display unit 20 or when capturing a video, the image acquisition unit 24 controls the imaging unit 23 to sequentially capture images at a predetermined frame rate and sequentially acquires image data of the captured images.

[0018] The object recognition unit 25 performs image recognition on the image data acquired by the image acquisition unit 24, and recognizes objects included in the image of the image data. For example, the object recognition unit 25 performs image recognition on the image data, and recognizes objects such as objects, object parts, and scenes. Any method may be used for image recognition. For example, the object recognition unit 25 recognizes the object, object parts, and scene using a recognition model that has learned the object, object parts, and scene using AI (Artificial Intelligence) technology.

[0019] The image capture device 10 includes a microphone. The audio capture unit 26 controls the microphone to capture and process ambient sounds to generate digital audio data. The audio data may be, for example, Linear PCM data.

[0020] The audio analysis unit 27 performs sound source separation processing on the audio data acquired by the audio acquisition unit 26, and generates audio data of multiple sound sources. For example, the audio analysis unit 27 generates audio data of the audio to be separated from the acquired audio data using a separation model that has learned the audio to be separated. For example, the audio analysis unit 27 performs sound source separation processing on the audio data using a processing configuration based on processing by a neural network such as a Deep Neural Network (DNN) as machine learning, and generates audio data of multiple sound sources.

[0021] 2 is a functional block diagram showing an example of the functional configuration of sound source separation processing in the audio analysis unit 27 according to the embodiment. In FIG. 2, an example of sound source separation using a neural network as a separation model will be described.

[0022] The audio analysis unit 27 has, as processing units, an encoder unit 30, a sequence modeling unit 31, and a decoder unit 32.

[0023] The input signal is a time waveform of the audio data or a spectrogram obtained by performing a short-time Fourier transform on the time waveform. The encoder unit 30 projects the input signal into a space that is easy for the sequence modeling unit 31 to process.

[0024] The sequence modeling unit 31 is responsible for modeling the temporal change in the output of the encoder unit 30. The decoder unit 32 is responsible for projecting the output of the sequence modeling unit 31 into the same space as the space of the input signal. Note that the encoder unit 30 and the decoder unit 32 may be omitted as necessary.

[0025] The input signal and the output of each processing unit of the speech analysis unit 27 may be used not only by the immediately succeeding processing unit, but also in skip connection processing for subsequent processing units as needed. The encoder unit 30, sequence modeling unit 31, and decoder unit 32 are realized by a combination of fully connected layers, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and the like. Low-latency processing is achieved by not using inputs that are temporally future than the outputs when calculating the outputs of each processing unit. That is, the speech analysis unit 27 is configured so that when each processing unit calculates its output in sequence, all data required for the calculation is input to each processing unit. This allows the speech analysis unit 27 to achieve low-latency processing without waiting for data. For example, this is achieved by adopting causal convolution for CNNs and unidirectional RNNs for RNNs.

[0026] Here, an example of a separation model training method will be described. In training the separation model, first audio data containing only the audio of the sound source to be separated and second audio data containing the audio of the first audio data and audio of another sound source are prepared. Then, in training the separation model, the first audio data is used as correct answer data, and sound source separation processing of the second audio data is performed using the separation model, and training is performed so that the correct answer data is separated to generate a separation model. A separation model is prepared in advance for each sound source to be separated.

[0027] In training the separation model, various parameters of the separation model are determined by collecting sufficient audio data of the sound source to be separated and other audio data, and performing machine learning using this as training data. For example, by performing machine learning using audio data containing only wind noise and audio data other than wind noise, it is possible to create a separation model that separates only wind noise.

[0028] The audio analysis unit 27 separates the audio data into audio data of multiple audio sources by serially or in parallel performing the audio source separation process shown in Fig. 2. Fig. 3 is a block diagram illustrating an example of audio source separation by the audio analysis unit 27 according to the embodiment. In the example of Fig. 3, first, voice separation is performed on the input audio to obtain voice and non-voice sounds. Then, the non-voice sounds are separated into wind and music sounds to obtain wind and music sounds. The audio analysis unit 27 outputs audio data for each separation target. If the input audio does not contain the sound to be separated, the audio data becomes silent data.

[0029] To perform voice separation, a separation model is created by training a large amount of human voice data and non-voice data in advance to separate the sound sources, and the data is output as voice and non-voice sound data. Preparing multiple separation models for the sounds to be separated in advance makes it possible to separate the target sound sources. For example, when performing the sound source separation shown in FIG. 3 , the sound analysis unit 27 includes a first separation model that separates voice-related sound data and non-voice sound data based on the voice data. The sound analysis unit 27 also includes a second separation model that separates wind-related sound data based on the non-voice sound data. The sound analysis unit 27 also includes a third separation model that separates music-related sound data based on the non-voice sound data. The sound analysis unit 27 uses the first separation model to separate voice from non-voice sounds. The sound analysis unit 27 then uses the second separation model to separate wind sounds from non-voice sounds. The sound analysis unit 27 also uses the third separation model to separate music sounds from non-voice sounds.

[0030]

[0031] Table 1 shows an example of sound sources to be separated.

[0032] The more sound sources there are to be separated, the more accurately the user can be informed of the sounds from the recorded sound sources. However, the more sound sources there are to be separated, the longer the processing time and the lower the separation accuracy may be.

[0033] The audio analysis unit 27 measures the volume of each piece of audio data obtained by sound source separation. Methods for measuring the volume include calculating a peak value, calculating an RMS (Root Mean Square) value, and calculating a loudness value.

[0034] The audio analysis unit 27 adjusts the volume gain of each piece of separated audio data. For example, the audio analysis unit 27 performs an adjustment called Auto Gain Control (AGC) on the separated audio data to generate audio data with adjusted volume. AGC is a process that keeps the input signal at a constant level by increasing the sensitivity when the input signal is weak and decreasing the sensitivity when the input signal is strong.

[0035] The moving image recording unit 21 records the image data sequentially acquired by the image acquisition unit 24 and the audio acquired by the audio acquisition unit 26 as moving image data (e.g., MPEG-4, etc.) in non-volatile memory (e.g., an SD card). For example, the audio analysis unit 27 adds together each piece of audio data whose volume has been adjusted by AGC, to generate audio data whose volume balance has been adjusted. The moving image recording unit 21 records moving image data consisting of the image data sequentially acquired by the image acquisition unit 24 and the audio data whose volume balance has been adjusted.

[0036] The display control unit 28 controls the display unit 20 to display images acquired by the image acquisition unit 24. When displaying a live view image or capturing a video, the display control unit 28 controls the display unit 20 to sequentially display images acquired sequentially at a predetermined frame rate by the image acquisition unit 24. Furthermore, the display control unit 28 controls the display unit 20 to display information regarding audio data emitted from a target object in association with image data, based on the recognition result by the object recognition unit 25 and the sound source separation result by the audio analysis unit 27. For example, the display control unit 28 displays information indicating the volume measured by the audio analysis unit 27 on the display unit 20. Methods of displaying the information include presenting the volume of the sound contained in the current moment, presenting the transition of the volume of the sound over the past few seconds, notifying that the volume of the sound has exceeded a certain threshold, and presenting a section in which the volume of the sound has exceeded a certain threshold. Furthermore, the display control unit 28 generates an OSD (On Screen Display) image that notifies the user of camera information such as the imaging conditions, and displays it on the display unit 20.

[0037] The operation control unit 29 receives operations from the user and controls the entire device according to the state.

[0038] Fig. 4 is a diagram showing an example of a display on the display unit 20 according to the embodiment. Fig. 4 shows a case where a moving image is displayed on the display unit 20, such as when a live view image is displayed or when a moving image is captured.

[0039] People, birds, the sun, and clouds are displayed on the display unit 20. The display unit 20 also displays display areas 40a to 40c.

[0040] Display area 40a displays a time series graph showing the transition of volume for each piece of audio data separated by sound source. For example, display area 40a displays a time series graph of the level of a target sound (e.g., a human voice) with a solid line 41a, and a time series graph of the level of non-target sounds, such as environmental sound (ambient) or noise, with a dashed line 41b. The time series graph displays the change in volume from the present time to the past few seconds.

[0041] The display area 40b displays the volume of a specific sound source at the current moment. For example, the display area 40b displays the volume of a specific noise contained in the sound at the current moment, in this case, the volume of wind noise.

[0042] The display area 40c displays the current volume for each source-separated audio data. For example, the display area 40c displays the level of non-target sounds, i.e., environmental sounds (ambient sounds) and noise, on channel 1 (CH1), and the level of target sounds (e.g., human voices) on channel 2 (CH2).

[0043] The photographer can grasp the volume of the voice of the speaking object contained in the collected audio by looking at the display unit 20. For example, the photographer can grasp the volume of the target sound (e.g., a person's voice) contained in the collected audio by looking at the display areas 40a and 40c. Furthermore, the photographer can confirm the level of a specific noise sound by looking at the display areas 40a to 40c, and can take measures according to the noise, such as changing the microphone settings, installing a windshield, or changing the shooting location.

[0044] 5 is a flowchart showing an example of a processing procedure of the imaging device 10 according to the embodiment. Fig. 5 shows a processing procedure for displaying a live view image.

[0045] The audio acquisition unit 26 controls the microphone to acquire surrounding audio and generates audio data in a format such as Linear PCM data (step S10).

[0046] The audio analysis unit 27 performs a sound source separation process on the audio data acquired by the audio acquisition unit 26 to generate audio data for each sound source (step S11). If the input of the audio analysis unit 27 is Linear PCM data, the output is also Linear PCM data.

[0047] The audio analysis unit 27 measures the volume of each piece of audio data obtained by the sound source separation (step S12).

[0048] The display control unit 28 controls the display unit 20 to display the images acquired by the image acquisition unit 24 (step S13). For example, the display control unit 28 controls the display unit 20 to sequentially display images acquired sequentially at a predetermined frame rate by the image acquisition unit 24. The display control unit 28 also controls the display unit 20 to display information about audio data emitted from the target object in association with image data, based on the recognition result by the object recognition unit 25 and the sound source separation result by the audio analysis unit 27. For example, the display control unit 28 displays information indicating the volume measured by the audio analysis unit 27 on the display unit 20. For example, the display control unit 28 measures the level of the audio data for each separated sound source, converts it into a format (e.g., a peak value, an RMS value, or a loudness value) that is easy for the photographer to view, and displays it.

[0049] The display control unit 28 determines whether to end the processing (step S14). For example, if the live view image is to be ended, the display control unit 28 determines that the processing is to be ended. If the processing is not to be ended (step S14: No), the display control unit 28 proceeds to step S10. On the other hand, if the processing is to be ended (step S14: Yes), the processing is ended.

[0050] 3. Specific Examples of Imaging Next, a gain adjustment method according to the present disclosure and a gain adjustment method according to a comparative example will be described with reference to FIGS. 6 and 7. FIG.

[0051] <3-1. Gain Adjustment Method of Comparative Example> Here, an example of conventional AGC processing will be described as a gain adjustment method of comparative example. FIG. 6 is a diagram illustrating a gain adjustment method according to the comparative example. Conventionally, in an imaging device, audio data is generated by performing AGC on input sound (input signal) input from a microphone. By performing AGC on the input sound, the sound is adjusted to be louder when the overall volume is low and to be quieter when the overall volume is high. This ensures that the recorded sound is at a constant level, making it easier for the human ear to hear. Furthermore, since peaks are suppressed when recording, it is possible to avoid sound distortion caused by exceeding the maximum value.

[0052] 6, AGC is performed on the input sound so that the volume does not exceed a predetermined upper limit level 50a and lower limit level 50b. The AGC output sound has a small amplitude waveform as a result of the gain being adjusted small so that the volume does not exceed the upper limit level 50a and lower limit level 50b.

[0053] However, if AGC is applied at a time when there is a high level of noise, such as wind noise, the gain will be adjusted in response to the higher level of noise, resulting in a problem where the human voice that you want to record as the target sound will be recorded at a low volume only at that time. Even if you separate such sounds into sound sources, the level of the separated human voice will vary, varying from loud to soft depending on the timing.

[0054] Figure 6 shows a case where sound source separation is performed on the output sound, and the sound is separated into voice and non-voice sounds (for example, wind noise). In Figure 6, when a high level of noise other than voice is present, the AGC adjusts the amplitude of the waveform to a small value, temporarily reducing the volume of the voice contained in the output sound. As a result, it becomes difficult to identify the human voice contained in the output sound when a high level of noise is present.

[0055] 3-2. Gain Adjustment Method of the Embodiment Next, a gain adjustment method of the embodiment will be described. Fig. 7 is a diagram illustrating the gain adjustment method of the embodiment.

[0056] The audio analysis unit 27 adjusts the volume gain of each of the separated audio data. For example, the audio analysis unit 27 performs sound source separation processing on the audio data of the input sound to separate it into audio data for each sound source. The audio analysis unit 27 performs AGC on each of the separated audio data to generate audio data. Then, the audio analysis unit 27 performs AGC and adds up the audio data to generate audio data for the output sound.

[0057] In Fig. 7, audio data of an input sound (input signal) is separated into audio data of multiple sound sources. Fig. 7 shows a case where sound source separation is performed on the input sound to separate it into voice and non-voice sounds (for example, wind noise). In Fig. 7, AGC is performed on the separated sounds so that the volume does not exceed an upper limit level 50a or a lower limit level 50b.

[0058] In Figure 7, when a high level of noise other than voice is present, the AGC reduces the amplitude of the waveform of the non-voice sound. However, by separating the voice sound and applying AGC, it is possible to prevent the volume of the voice contained in the output sound from being reduced due to the influence of noise. As a result, the human voice contained in the output sound can be kept easy to understand.

[0059] 4. Processing Procedures> Fig. 8 is a flowchart showing an example of a processing procedure of the imaging device 10 according to the embodiment. Fig. 8 shows a processing procedure when capturing a video while displaying a live view image. The processing shown in Fig. 8 is partially similar to the processing shown in Fig. 5 , and therefore the same parts are assigned the same reference numerals and their descriptions are omitted, and the following mainly describes the differences.

[0060] The audio analysis unit 27 adjusts the volume gain of each of the separated audio data (step S20). For example, the audio analysis unit 27 performs AGC on the separated audio data to generate audio data with adjusted volume. The audio analysis unit 27 adds together the audio data whose volume has been adjusted by performing AGC, and generates audio data with adjusted volume balance.

[0061] The moving image recording unit 21 records the image data sequentially acquired by the image acquisition unit 24 and the audio acquired by the audio acquisition unit 26 as moving image data in a non-volatile memory (step S21). For example, the moving image recording unit 21 records moving image data consisting of the image data sequentially acquired by the image acquisition unit 24 and audio data with adjusted volume balance.

[0062] In the processing procedure shown in FIG. 8, steps S20 and S21 are performed in parallel with steps S12 and S13, but steps S12 and S13 and steps S20 and S21 may also be performed in series.

[0063] 5. Modifications Incidentally, the above-described embodiment of the present disclosure can be modified in several ways.

[0064] 5-1. First Modification The audio analysis unit 27 may determine a sound source to be separated based on either the recognition result by the object recognition unit 25 or the type of microphone used to collect sound, and may separate audio data of the determined sound to be separated from the acquired audio data. For example, the audio analysis unit 27 may determine a sound source to be a target sound, ambient sound, or noise based on the object recognition result by the object recognition unit 25. Ambient sound refers to environmental sound generated by the surrounding environment, etc. For example, if the object recognition unit 25 recognizes a person, the audio analysis unit 27 may separate a "human voice" as the target sound. The audio analysis unit 27 may establish a correspondence between the recognition target and a sound source to be a target sound, ambient sound, or noise, and separate audio data of the sound source corresponding to the recognition target recognized by the object recognition unit 25 in the correspondence. The correspondence may be determined one-to-one or one-to-many, with the recognition target corresponding to the sound source to be separated.

[0065]

[0066] Table 2 shows an example of the correspondence between the recognition target and the sound source to be separated.

[0067] Using the correspondence relationship in Table 2, the sound source to be separated can be determined based on the recognition results by the object recognition unit 25 and the type of microphone used for sound collection. For example, if the correspondence relationship is defined as shown in Table 2, when the sound analysis unit 27 recognizes a person, it separates "human voice" as the target sound. When it recognizes a beach, it separates "sounds of the ocean," "birds singing," and "surrounding voices" as ambient sounds. Furthermore, when collecting a person's voice using a pin microphone (lavalier microphone), pin microphones are often attached to clothing. Depending on the person's movement and how the microphone is attached, the rustling noise of clothing against the clothing can be picked up as noise. Therefore, when a wireless microphone is connected to the imaging device 10, the sound analysis unit 27 assumes that a pin microphone is used on the wireless handset microphone side and separates "clothes rustling" as noise. A correspondence relationship may also be defined between the location information of the imaging device 10 and a sound source. For example, if the imaging device 10 is located in a city, the sound source to be separated may be determined as "car sounds."

[0068] In this way, the image capture device 10 according to the first modification can separate only the appropriate sound sources by determining the sound source to be separated based on either the recognition result by the object recognition unit 25 or the type of microphone used to collect sound. The sound source separation process takes longer as the number of separations increases. The image capture device 10 according to the first modification can suppress the separation of unnecessary sound sources, ensuring real-time performance and improving efficiency.

[0069] 9 is a flowchart showing an example of a processing procedure of the imaging device 10 according to the first modified example. Fig. 9 shows a processing procedure for determining the target sound and ambience to be separated.

[0070] The subject recognition unit 25 performs image recognition on the image data and determines whether a main subject is present in the image of the image data (step S40). The main subject is the object that is the lens focus (in-focus) among the recognized objects contained in the image recognized by image recognition. For example, if a person and an animal are recognized as objects in the image and the focus is on the person, the main subject is the person. The main subject may be determined automatically from the recognition state, such as by autofocus, or may be determined by the user selecting a location on the screen where the user wants to focus.

[0071] If a main subject is present (step S40: Yes), the audio analysis unit 27 determines the sound source corresponding to the recognized main subject as the target sound (step S41) using the correspondence relationship in Table 2. On the other hand, if a main subject is not present (step S40: No), the audio analysis unit 27 determines that there is no target sound.

[0072] The object recognition unit 25 determines as ambient any sound source other than noise that is associated with the recognition state and microphone type of a subject other than the main subject in the correspondence relationship in Table 2 (step S42). Sound sources that are considered noise are determined in advance, for example, as shown in Table 3. The object recognition unit 25 determines as ambient any sound source other than the sound sources registered in Table 3 that is associated with the recognition state and microphone type of a subject other than the main subject.

[0073]

[0074] Table 3 shows examples of sound sources that are used as noise.

[0075] The noise sources are sounds that humans find unpleasant or disturbing, such as wind noise, rustling of clothes, and household appliance noise (such as the running sound of an air conditioner or refrigerator), so they can be uniquely determined regardless of the situation.

[0076] <5-2. Second Modification> Before performing sound source separation, the audio analysis unit 27 may determine the sound sources to be the target sound, ambient sound, and noise based on the recognition result by the object recognition unit 25 or the type of microphone used to collect sound. Furthermore, the audio analysis unit 27 may further perform gain adjustment for each of the target sound, ambient sound, and noise sound sources for each piece of audio data whose volume has been adjusted by AGC.

[0077] <5-2-1. Processing Procedure> Fig. 10 is a flowchart showing an example of the processing procedure of the imaging device 10 according to the second modification. Fig. 10 shows the processing procedure when capturing a video while displaying a live view image. The processing shown in Fig. 10 is partially similar to the processing shown in Figs. 5 and 8, and therefore the same parts are assigned the same reference numerals and their descriptions are omitted, and the following mainly describes the differences.

[0078] After step S10, the sound analysis unit 27 determines the sound source to be the target sound, ambient sound, or noise based on the recognition result by the object recognition unit 25 and the type of microphone used to collect the sound (step S50), and proceeds to step S11.

[0079] After step S20, the audio analysis unit 27 multiplies each piece of audio data whose volume has been adjusted by AGC by a coefficient for each of the target sound, ambient sound, and noise sound sources to generate audio data multiplied by the coefficients (step S51), and then proceeds to step S21. The coefficients are set so that the ambient sound is smaller than the target sound, and the noise is smaller than the ambient sound. For example, the audio analysis unit 27 performs gain adjustment for each piece of audio data whose volume has been adjusted by AGC using the coefficients shown in Table 4 for each of the target sound, ambient sound, and noise sound sources to generate audio data. The audio analysis unit 27 adds together the pieces of audio data whose gains have been adjusted to generate audio data whose volume balance has been adjusted.

[0080]

[0081] Table 4 shows examples of coefficients for performing gain adjustment for the target sound, ambient sound, and noise sound sources. Note that other coefficients are also registered in Table 4. The audio analysis unit 27 may perform gain adjustment by multiplying audio data of sound sources other than the target sound, ambient sound, and noise by a coefficient registered as other.

[0082] The audio analysis unit 27 desensitizes each sound source using the coefficients in Table 4 when determining that the sound source is a target sound, ambient sound, or noise. Sound sources such as noise can be significantly desensitized, while sound sources that should be left as part of the atmosphere or surrounding environment, such as ambient sound, can be desensitized slightly. This coefficient may also be set by the user. A configuration may also be adopted in which each sound source has its own gain coefficient. A coefficient may also be used to increase the sensitivity of the target sound source with a gain to maintain a certain level or higher. In combination with the detailed recognition state, different gain coefficients may be used depending on whether the person is facing forward or backward (for example, a stronger gain may be used when facing backward because the voice entering the microphone is weaker).

[0083] In this way, the imaging device 10 of the second variant adjusts the gain for ambient and noise sound sources other than the target sound and records them with reduced sensitivity, thereby lowering the level of other sounds, making it easier to hear only the target sound in the recorded sound, and improving functionality.

[0084] 5-3. Third Modification The display unit 20 may accept designation of sound sources as target sound, ambient sound, and noise. The audio analysis unit 27 may separate audio data of the sound sources designated as target sound, ambient sound, and noise.

[0085] 11A and 11B are diagrams showing an example of a display on the display unit 20 according to the third modified example. Fig. 11A shows an example of a category designation area 42 displayed on the display unit 20. Fig. 11B shows an example of a display area 40a displayed on the display unit 20.

[0086] The category designation area 42 shown in FIG. 11A is a window that displays the sound sources currently being separated by category: target sound, ambient, and noise. The category designation area 42 has areas 42a to 42c for target sound, ambient, and noise. Areas 42a to 42c display icons 43 indicating the sound sources that have been determined to be target sound, ambient, or noise. By looking at the category designation area 42, the user can see at a glance which sound sources have been detected / separated. The current gain and volume of each individual icon 43 may also be displayed. The user can also move sound source categories by dragging and dropping, and can manually add new sound sources.

[0087] The display area 40a shown in FIG. 11B is a window that displays a time series graph of the sound level for each category of target sound, ambient sound, and noise. In the display area 40a, the target sound, ambient sound, and noise levels are displayed using solid lines, dashed lines, or different colors. In FIG. 11B, the time series graph of the target sound is displayed using line 44a, the time series graph of ambient sound is displayed using line 44b, and the time series graph of noise is displayed using line 44c. When a specific sound source is selected from the window of the category designation area 42, the display control unit 28 may display only the time series graph of the sound of the selected sound source in the display area 40a. The display area 40a displays an upper limit level 50a and a lower limit level 50b of the target sound. The audio analysis unit 27 adjusts the gain so that the volume of the target sound does not exceed the upper limit level 50a and the lower limit level 50b. The upper limit level 50a and the lower limit level 50b may be freely adjustable by the user. Furthermore, upper and lower limit bars for ambient and noise levels as well as the target sound may be displayed so that the upper and lower limit levels of the ambient and noise levels can be adjusted.Also, upper and lower limit bars for each sound source may be displayed so that the upper and lower limit levels can be adjusted for each sound source.

[0088] In this way, the imaging device 10 according to the third variant allows the user to easily select the sound source to be used as target sound, ambient sound, or noise while shooting, and provides a user interface suitable for collecting sounds that are easy to hear, thereby improving usability.

[0089] 5-4. Fourth Modification The operation control unit 29 may control imaging based on the sound source separated by the audio analysis unit 27. For example, the operation control unit 29 may control imaging in accordance with a predetermined rule based on the separated sound source. Adjustments are made to generate audio data. For example, the operation control unit 29 may control imaging in accordance with the rule table of Table 5.

[0090]

[0091] Table 5 shows an example of a rule table for executing imaging control.

[0092] As shown in Table 5, the operation control unit 29 performs an action according to a rule-based condition when it detects it. For example, the operation control unit 29 starts recording when the separated sound source "animal cries (or footsteps)" exceeds a certain level. When photographing wild animals, animals do not approach if there is a person present. For this reason, in conventional methods of photographing wild animals, a camera is often installed and equipped with an infrared sensor or the like to automatically start photographing when an animal approaches.

[0093] In contrast, the imaging device 10 according to the fourth modification starts imaging by using a specific separated sound, and therefore can be realized by the imaging device 10 alone, without the need for an additional device such as an infrared sensor. Note that the display control unit 28 may display a caution on the display unit 20 urging the user to wear a wind jammer when strong winds are detected. This allows the user and the imaging device 10 to cooperate to collect sound better.

[0094] In this way, the imaging device 10 according to the fourth modification can use the separated sound sources to capture images appropriate to the current situation, thereby improving usability.

[0095] <5-5. Fifth Modification> In the above-described embodiment, an example has been described in which the technology of the present disclosure is applied to the imaging device 10. However, this is not limiting. For example, the technology of the present disclosure may be implemented in a distributed manner between the imaging device 10 and another device. In other words, the technology of the present disclosure may be realized in a system configuration including multiple devices.

[0096] 12 is a diagram showing an example of the configuration of an imaging system 1 according to Modification 5. The imaging system 1 includes an imaging device 10 and a monitoring device 11.

[0097] The monitoring device 11 is, for example, an information processing device such as a smartphone or a tablet. The monitoring device 11 receives video and audio from the imaging device 10 and transmits control signals. The imaging device 10 and the monitoring device 11 are connected to each other so that they can communicate with each other via a wireless or wired network. Examples of communication methods between the imaging device 10 and the monitoring device 11 include a wired connection using a USB (Universal Serial Bus) cable, a LAN (Local Area Network) cable, or the like, and a wireless connection using Wi-Fi, Bluetooth (registered trademark), or the like.

[0098] <5-5-2. Configuration of Monitoring Device> FIG. 13 is a functional block diagram showing an example of the functional configuration of a monitoring device 11 according to the fifth modified example.

[0099] As shown in FIG. 13, the monitoring device 11 includes a display unit 50, a receiving unit 51, a transmitting unit 52, a voice analyzing unit 53, and a display control unit 54.

[0100] The display unit 50 is a display provided on the monitoring device 11. The receiving unit 51 receives video, audio, and related metadata from the imaging device 10. The receiving unit 51 receives the recognition result by the object recognition unit 25 and the type of microphone collecting sound via metadata. The transmitting unit 52 transmits control signals to the imaging device 10 to externally control the operation of the imaging device 10. The operation control unit 55 accepts operations from the user and controls the entire device depending on the status.

[0101] The audio analysis unit 53 performs sound source separation processing on the audio data received by the receiving unit 51 to generate audio data for multiple sound sources. For example, the audio analysis unit 53 generates audio data for the audio to be separated from the acquired audio data using a separation model that has learned the audio to be separated. The audio analysis unit 53 measures the volume of each piece of audio data obtained by the sound source separation. Note that, similar to the audio analysis unit 27, the audio analysis unit 53 may perform gain adjustment on the separated audio data, generate audio data by adding together the gain-adjusted audio data, and transmit the generated audio data to the imaging device 10 for recording.

[0102] The display control unit 54 controls the display unit 20 to display the video received via the metadata. Furthermore, the display control unit 54 controls the display unit 20 to display information about audio data emitted from the object to be spoken, in association with an image, based on the recognition result by the object recognition unit 25 and the sound source separation result by the audio analysis unit 53. For example, the display control unit 54 displays information indicating the volume measured by the audio analysis unit 27 on the display unit 20. Furthermore, the display control unit 54 generates an OSD image that notifies the user of camera information such as the imaging conditions, and displays it on the display unit 20.

[0103] In this way, the imaging system 1 according to the fifth modification makes it possible to check moving images on the monitoring device 11 connected to the imaging device 10, and to grasp the volume of the speech of the object being spoken about in the collected audio on a large screen even in a location far from the imaging device 10. Furthermore, the imaging system 1 according to the fifth modification uses the resources of the monitoring device 11 for analysis, thereby enabling more advanced processing than the imaging device 10 and improving functionality.

[0104] Among the processes described in the above embodiments of the present disclosure, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the process procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the illustrated information.

[0105] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0106] The above-described embodiments of the present disclosure can be combined as appropriate within the scope of the present disclosure without causing any contradiction in the processing content. The order of the steps shown in the sequence diagrams or flowcharts of the present embodiments can be changed as appropriate.

[0107] 6. Hardware Configuration The imaging device 10 and monitoring device 11 according to the above-described embodiments of the present disclosure are realized by a computer 1000 configured as shown in FIG. 14 , for example. The imaging device 10 will be described as an example. FIG. 14 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the imaging device 10. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM 1300, a secondary storage device 1400, a communication interface 1500, an input / output interface 1600, a display unit 1700, a camera unit 1800, a microphone 1900, and a speaker 2000. The components of the computer 1000 are connected via a bus 1050.

[0108] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the secondary storage device 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the secondary storage device 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0109] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0110] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the imaging device 10 according to an embodiment of the present disclosure, which are examples of program data 1450.

[0111] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550. For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0112] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a microphone or a touch panel via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display or a speaker via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of the media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0113] For example, when the computer 1000 functions as the imaging device 10 according to an embodiment of the present disclosure, the CPU 1100 of the computer 1000 executes programs loaded on the RAM 1200 to function as an image acquisition unit 24, an object recognition unit 25, an audio acquisition unit 26, an audio analysis unit 27, a display control unit 28, and an operation control unit 29. The secondary storage device 1400 stores the programs according to the present disclosure and data on the separation model. The CPU 1100 reads and executes program data 1450 from the secondary storage device 1400, but as another example, the CPU 1100 may acquire these programs from another device via an external network 1550.

[0114] 7. Conclusion As described above, according to an embodiment of the present disclosure, the imaging device 10 (corresponding to an example of an "imaging system") includes the image acquisition unit 24, the object recognition unit 25 (recognition unit), the audio acquisition unit 26, the audio analysis unit 27, and the display control unit 28. The image acquisition unit 24 acquires image data including a target object. The object recognition unit 25 recognizes an object included in the image of the image data. The audio acquisition unit 26 acquires audio data related to the target object. The audio analysis unit 27 separates the audio data into multiple audio data. The display control unit 28 associates information related to the audio data emitted from the target object with the image data and displays it on the display unit 20 based on the recognition result by the object recognition unit 25 and the sound source separation result by the audio analysis unit 27. This allows the imaging device 10 to grasp the volume of the audio of the target object included in the collected audio.

[0115] Moreover, according to an embodiment of the present disclosure, the imaging system 1 (corresponding to an example of an "imaging system") includes an image acquisition unit 24, an object recognition unit 25 (recognition unit), an audio acquisition unit 26, an audio analysis unit 53, and a display control unit 54. The image acquisition unit 24 acquires image data including a target object. The object recognition unit 25 recognizes an object included in the image of the image data. The audio acquisition unit 26 acquires audio data related to the target object. The audio analysis unit 53 separates the audio data into multiple audio data. The display control unit 54 associates information related to the audio data emitted from the target object with the image data and displays it on the display unit 50 based on the recognition result by the object recognition unit 25 and the sound source separation result by the audio analysis unit 53. This allows the imaging system 1 to grasp the volume of the audio of the target object included in the collected audio.

[0116] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.

[0117] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0118] The present technology may also be configured as follows: (1) An imaging system comprising: an image acquisition unit that acquires image data including a target object; a recognition unit that recognizes an object included in an image of the image data; an audio acquisition unit that acquires audio data related to the target object; an audio analysis unit that separates the audio data into multiple audio data; and a display control unit that causes a display unit to display information related to the audio data emitted from the target object in association with the image data based on a recognition result by the recognition unit and a sound source separation result by the audio analysis unit. (2) The imaging system described in (1), wherein the audio analysis unit comprises a first separation model that separates audio data related to voice and audio data other than voice based on the audio data. (3) The imaging system described in (2), wherein the audio analysis unit further comprises a second separation model that separates audio data related to wind based on the audio data other than voice. (4) The imaging system described in (3), wherein the audio analysis unit further comprises a third separation model that separates audio data related to music based on the audio data other than voice. (5) The imaging system according to (1), wherein the image acquisition unit sequentially acquires image data at a predetermined frame rate, the audio acquisition unit acquires audio data while the image acquisition unit sequentially acquires the image data, and the display control unit outputs information related to the plurality of audio data separated from the sound sources, synchronized with the sequentially acquired image data. (6) The imaging system according to (5), wherein the display control unit displays a transition in volume of each of the plurality of audio data separated from the sound sources on the display unit. (7) The imaging system according to (5), wherein the audio analysis unit adjusts a gain of volume for each of the separated audio data, and further includes a recording unit that records moving image data sequentially acquired by the image acquisition unit and audio data whose volume balance has been adjusted.(8) The imaging system according to (1), wherein the audio analysis unit determines a sound source to be separated based on the recognition result by the recognition unit or the type of microphone used to collect sound, and performs sound source separation on the audio data of the determined sound source from the audio data acquired by the audio acquisition unit. (9) The imaging system according to (8), wherein the audio analysis unit determines a sound source to be determined as target sound, ambient sound, or noise based on the object recognition result by the recognition unit. (10) The imaging system according to (6), wherein the display unit displays a volume for each of the target sound, ambient sound, and noise, and accepts designation of the sound source to be determined as target sound, ambient sound, or noise, and the audio analysis unit performs sound source separation on the audio data of the sound source designated on the display unit as target sound, ambient sound, or noise. (11) The imaging system according to (1), further comprising an operation control unit that controls imaging based on the sound source separated by the audio analysis unit. (12) A display method comprising: acquiring image data including an object to be spoken to; recognizing an object included in an image of the image data; acquiring audio data related to the object to be spoken to; separating the audio data into a plurality of pieces of audio data; and controlling a display unit to display information related to the audio data emitted from the object to be spoken to in association with the image data, based on the object recognition result and the sound source separation result. (13) A program causing a computer to execute the following processes: acquiring image data including an object to be spoken to; recognizing an object included in an image of the image data; acquiring audio data related to the object to be spoken to; separating the audio data into a plurality of pieces of audio data, and controlling a display unit to display information related to the audio data emitted from the object to be spoken to in association with the image data, based on the object recognition result and the sound source separation result.

[0119] REFERENCE SIGNS LIST 1 Imaging system 10 Imaging device 11 Monitoring device 20, 50 Display unit 21 Video recording unit 22 Optical unit 23 Imaging unit 24 Image acquisition unit 25 Object recognition unit 26 Audio acquisition unit 27, 53 Audio analysis unit 28, 54 Display control unit 29, 55 Operation control unit 30 Sequence encoder unit 31 Modeling unit 32 Decoder unit 40a, 40b, 40c Display area 42 Category designation area 42a, 42b, 42c Area 51 Receiving unit 52 Transmitting unit

Claims

1. An imaging system comprising: an image acquisition unit that acquires image data including an object to be spoken to; a recognition unit that recognizes an object included in an image of the image data; an audio acquisition unit that acquires audio data related to the object to be spoken to; an audio analysis unit that separates the audio data into multiple audio data; and a display control unit that displays information about the audio data emitted from the object to be spoken to on a display unit in association with the image data based on the recognition result by the recognition unit and the sound source separation result by the audio analysis unit.

2. The imaging system according to claim 1, wherein the audio analysis unit includes a first separation model that separates audio data related to voice and audio data other than voice based on the audio data.

3. The imaging system according to claim 2, wherein the audio analysis unit further comprises a second separation model that separates audio data relating to wind based on the audio data other than the voice.

4. The imaging system according to claim 3, wherein the audio analysis unit further comprises a third separation model for separating audio data relating to music based on the audio data other than the voice.

5. The imaging system of claim 1, wherein the image acquisition unit sequentially acquires image data at a predetermined frame rate, the audio acquisition unit acquires audio data while the image acquisition unit sequentially acquires the image data, and the display control unit outputs information related to the multiple audio data separated from the sound sources to the display unit in synchronization with the image data being sequentially acquired.

6. The imaging system according to claim 5, wherein the display control unit causes the display unit to display a transition in the volume of each of the plurality of audio data separated from each other.

7. The imaging system according to claim 5, further comprising a recording unit that adjusts the volume gain of each of the separated audio data, and records moving image data consisting of the image data sequentially acquired by the image acquisition unit and the audio data whose volume balance has been adjusted.

8. The imaging system of claim 1, wherein the audio analysis unit determines the sound source to be separated based on either the recognition result by the recognition unit or the type of microphone used to collect sound, and separates the audio data of the determined sound source from the audio data acquired by the audio acquisition unit.

9. The imaging system according to claim 8, wherein the sound analysis unit determines the sound source to be the target sound, ambient sound, or noise based on the object recognition result by the recognition unit.

10. The imaging system according to claim 6, wherein the display unit displays the volume for each of the target sound, ambient sound, and noise, and accepts designation of the sound source to be designated as the target sound, ambient sound, or noise, and the audio analysis unit separates the audio data of the sound source designated as the target sound, ambient sound, or noise on the display unit.

11. The imaging system according to claim 1, further comprising an operation control unit that controls imaging based on the sound source separated by the sound analysis unit.

12. A display method in which a computer executes the following process: acquires image data including an object to be spoken to; recognizes the object included in the image of the image data; acquires audio data related to the object to be spoken to; separates the audio data into multiple pieces of audio data; and, based on the object recognition result and the audio source separation result, controls the display unit to display information related to the audio data emitted from the object to be spoken to in association with the image data.

13. A program that causes a computer to perform the following processes: acquire image data including an object to be spoken to; recognize the object included in the image of the image data; acquire audio data related to the object to be spoken to; separate the audio data into multiple pieces of audio data; and, based on the object recognition results and the audio source separation results, control the display unit to display information related to the audio data emitted from the object to be spoken to in association with the image data.

Citation Information

Patent Citations

  • Processing method and device, camera equipment and electronic system

    CN113709378A

  • Sound monitoring apparatus

    JP2009118318A

  • Sound collection device, sound collection device control method

    JP2020003724A

  • Acoustic scene reconstruction device, acoustic scene reconstruction method, and program

    JP2020030376A

  • Photographing device

    JP2020156076A