Imaging apparatus, control method, and program

The imaging apparatus addresses the issue of misaligned sound direction in stereophonic technologies by using correction means within the imaging apparatus to align the sound data with the shooting direction, effectively reducing user discomfort during playback.

JP2025082063APending Publication Date: 2025-05-28CANON KK

Patent Information

Application Number
JP2023195292
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-28

AI Technical Summary

Technical Problem

Existing stereophonic technologies, such as ambisonics, struggle to correct deviations between the shooting direction of an imaging device and the sound collection direction of a sound collection device, leading to misalignment of sound direction with respect to video during playback, causing user discomfort.

Method used

An imaging apparatus equipped with connection means for a sound collection device, sound processing means for generating stereophonic sound data, and correction means to align the sound data with the shooting direction, using attitude detection and communication units to calculate and correct the difference between the shooting and sound collection directions.

Benefits of technology

The solution effectively corrects deviations between the shooting and sound collection directions, ensuring that the sound direction aligns with the video, thereby reducing user discomfort during playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025082063000001_ABST
    Figure 2025082063000001_ABST
Patent Text Reader

Abstract

To correct displacement between a photographing direction and a sound collecting direction.SOLUTION: An imaging apparatus has: connection means that connects to a sound collecting device that collects voice data in different directions; voice processing means that generates three-dimensional acoustic data on the basis of the voice data acquired from the sound collecting device; and correction means that corrects the three-dimensional acoustic data on the basis of the difference between the photographing direction of the imaging apparatus and a sound collecting direction of the sound collecting device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for correcting stereophonic data.

Background Art

[0002] Conventionally, stereophonic technologies such as ambisonics have been known. Ambisonics is a technology that enables sound collection in all sound fields of 360° in a three-dimensional space by four microphone pieces of an ambisonics microphone, and is often used as the sound for stereoscopic video. Ambisonics can convert an A-format audio signal collected by an ambisonics microphone into B-format four-channel audio data, and generate an audio signal with an arbitrary directivity based on the B-format audio data. Patent Document 1 describes a method of generating stereophonic data by ambisonics during video shooting and correcting the stereophonic data based on the posture of the imaging device during video shooting.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Patent Document 1 assumes that the imaging unit and the sound collection unit are integrally configured and the shooting direction and the sound collection direction coincide, and cannot correct the deviation between the shooting direction and the sound collection direction. For this reason, there is a possibility that the direction of the sound with respect to the video does not match during playback, giving the user a sense of discomfort.

[0005] The present invention has been made in view of the above problems, and an object thereof is to realize a technique capable of correcting the deviation between the shooting direction and the sound collection direction so as to match the direction of the sound with respect to the video.

Means for Solving the Problems

[0006] In order to solve the above problems and achieve the object, an imaging apparatus according to the present invention includes connection means for connecting to a sound collection device that collects sound data in different directions, sound processing means for generating stereophonic sound data based on the sound data acquired from the sound collection device, and correction means for correcting the stereophonic sound data based on a difference between a shooting direction of the imaging apparatus and a sound collection direction of the sound collection device.

Advantages of the Invention

[0007] According to the present invention, it is possible to correct a deviation between a shooting direction and a sound collection direction so as to align the direction of sound with respect to an image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0010] [Embodiment 1] FIG. 1 is a block diagram showing the configuration of the imaging device 100 according to Embodiment 1.

[0011] The lens unit 300 is detachable from the imaging device 100. The lens unit 300 includes an optical lens and a drive mechanism for driving the optical lens. By driving the optical lens, the lens unit 300 performs focusing, zooming, shake correction, etc. of a subject during imaging. The subject includes a moving object and a stationary object. A plurality of types of lens units 300 are prepared, and the user can replace and use a desired lens according to the application. The lens unit 300 constitutes a binocular lens having two optical axes so that a stereoscopic video such as a 180° video or a 360° video can be imaged by the imaging device 100. The stereoscopic video utilizes binocular parallax, which is the displacement of the images on the retinas of a person's left and right eyes, and stereoscopic vision is realized by the human visual system recognizing the depth of the video due to the displacement of the images. The lens unit 300 can form a shooting range (angle of view) so as to cover the entire omnidirectional range at maximum.

[0012] The imaging unit 101 includes an imaging element that converts the optical image of the subject transmitted through the lens unit 300 into an analog image signal by the imaging element. The imaging unit 101 converts the analog image signal obtained by the imaging element into a digital signal and generates image data by performing various image processes. The imaging element is a photoelectric conversion element such as a CCD or a CMOS. The lens control unit 102 controls the lens unit 300 as necessary based on information obtained from the image data generated by the imaging unit 101 and information obtained from the first control unit 110 described later.

[0013] The first audio processing unit 103 receives audio data from the sound collection device 200 connected to the imaging device 100 by the first communication unit 112 described later. The first audio processing unit 103 includes a B-format encoder 1031 and an audio correction unit 1032, and performs first audio processing on the audio data received from the sound collection device 200. The B-format encoder 1031 performs a process of converting the audio data in the A-format of ambisonics received from the sound collection device 200 into the B-format. The audio correction unit 1032 performs a process of correcting the B-format audio data based on the difference between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 described later. Details of the first audio processing will be described later. In this embodiment, an example in which the imaging device 100 and the sound collection device 200 are connected by a wireless communication method will be described, but the present invention is not limited to this, and they may be wired-connected by an audio cable or the like capable of inputting and outputting stereo audio.

[0014] The memory 104 temporarily stores the image data obtained by the imaging unit 101 and the audio data obtained by the first audio processing unit 103.

[0015] The display control unit 105 generates display data for displaying the image data obtained by the imaging unit 101, the operations and settings of the imaging device 100, and the GUI, etc., and outputs it to the display unit 106. The display unit 106 includes a display device such as a liquid crystal or an organic EL that displays a video and a GUI based on the display data.

[0016] The symbolization processing unit 107 reads out the image data and audio data stored in the memory 104, performs predetermined symbolization processing, and generates compressed and symbolized image data and audio data. Note that the audio data may not be symbolized. The compression and symbolization method of the image data may be any method such as MPEG2 or H.264 / MPEG4-AVC. Also, the compression and symbolization method of the audio data may be any method such as AC3, AAC, ATRAC, or ADPCM.

[0017] The recording unit 108 controls the writing of data to the recording medium 109 and the reading of data from the recording medium 109. The recording unit 108 controls the writing and reading of the image data and audio data compressed and symbolized by the symbolization processing unit 107 or the non-compressed and non-symbolized audio data to and from the recording medium 109.

[0018] The recording medium 109 is an auxiliary storage device such as a magnetic disk, an optical disk, or a semiconductor memory that can record image data and audio data.

[0019] The first control unit 110 includes a processor such as a CPU that performs arithmetic processing and control processing related to the imaging device 100, a non-volatile memory that stores programs executed by the processor and data referred to by the programs, and a volatile memory into which the programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM. The non-volatile memory and the volatile memory may be in the form of an internal memory built into the imaging device 100 or in the form of an external memory externally connected to the imaging device 100. Note that the program of this embodiment includes a program for executing the flowcharts described later in FIGS. 5, 7, and 9.

[0020] The operation unit 111 includes operation members such as push buttons, rotary dials, and touch panels integrally configured with the display unit 106 that receive user operations. The operation unit 111 sends state information corresponding to the operation state of the operation unit 111 to the first control unit 110. The operation members included in the operation unit 111 include a power button for turning on and off the power of the imaging device 100, a mode dial for setting the operation mode of the imaging device 100, a shooting button for instructing the start or end of the shooting operation of the imaging device 100, a volume for adjusting the volume of the reproduced sound, and the like. The operation mode of the imaging device 100 can be switched to any of a still image shooting mode, a moving image shooting mode, and a playback mode for playing back still images, moving images, and sounds.

[0021] The first communication unit 112 is communicably connected to an external device such as the sound collecting device 200 by a wireless communication method such as Wi-Fi (registered trademark) or Bluetooth (registered trademark) or a wired communication method such as USB, and performs transmission and reception of control information, voice data, and the like.

[0022] The first attitude detection unit 113 detects the attitude information of the imaging device 100 and sends it to the first control unit 110. The first control unit 110 detects the direction of the optical axis (shooting direction) of the imaging device 100 based on the attitude information obtained from the first attitude detection unit 113. The first attitude detection unit 113 includes an acceleration sensor, an angular velocity sensor, a geomagnetic sensor, and the like.

[0023] The system bus 114 includes an address bus, a data bus, and a control bus that connect the respective components of the imaging device 100 so that data can be exchanged.

[0024] Here, the basic operation of the imaging device 100 of the present embodiment will be described.

[0025] When the power button included in the operation unit 111 of the imaging device 100 of the present embodiment is turned on, power is supplied from a power supply unit (not shown) to each component of the imaging device 100, and each component of the imaging device 100 becomes operable.

[0026] The first control unit 110 determines the operation mode of the imaging device 100 based on the state information corresponding to the operation state of the mode dial included in the operation unit 111.

[0027] In the video shooting mode, the first control unit 110 sends control information for causing each component of the imaging device 100 to shift to the shooting standby state. The operations in the shooting standby state are as follows.

[0028] The imaging unit 101 captures the optical image of the subject captured by the lens unit 300 with the imaging element to generate image data. The image data generated by the imaging unit 101 is sent to the display control unit 105. The display control unit 105 generates display data based on the image data and displays it on the display unit 106. The user prepares for shooting while viewing the video displayed on the display unit 106.

[0029] In the shooting standby state, when the user operates the shooting button included in the operation unit 111, an instruction to start the recording / recording operation is notified to the first control unit 110. When receiving the instruction to start the recording / recording operation, the first control unit 110 sends control information for shifting each component of the imaging device 100 to the video shooting mode. When the recording starts, the first control unit 110 stores the image data obtained by the imaging unit 101 and the audio data obtained by the first audio processing unit 103 as one file. The recording / recording operation of the imaging device 100 in the video shooting mode is as follows.

[0030] The imaging unit 101 captures the optical image of the subject captured by the lens unit 300 with the imaging element to generate image data. The image data generated by the imaging unit 101 is sent to the display control unit 105 and stored in the memory 104. The display control unit 105 generates display data based on the image data and displays it on the display unit 106.

[0031] The first audio processing unit 103 executes first audio processing on the audio data received from the sound collection device 200 by the first communication unit 112. The first audio processing unit 103 executes first audio processing on the plurality of audio data obtained by the plurality of microphones provided in the sound collection device 200 to generate multi-channel audio data. The generated audio data is stored in the memory 104.

[0032] The encoding processing unit 107 reads out the image data and audio data stored in the memory 104, performs predetermined encoding processing, and generates compressed and encoded image data and audio data.

[0033] The first control unit 110 synthesizes the compressed and encoded image data and audio data by the encoding processing unit 107 to form a data stream, and sends it to the recording unit 108. When the audio data is not compressed and encoded, the first control unit 110 synthesizes the audio data stored in the memory 104 and the compressed and encoded image data to form a data stream, and sends it to the recording unit 108.

[0034] The recording unit 108 writes the data stream as one video file to the recording medium 109 so that it can be managed according to a file system such as UDF (Universal Disk Format) or FAT (File Allocation Tables).

[0035] During recording and recording, the above operations are continued.

[0036] During recording and recording, when the user operates the shooting button included in the operation unit 111, an instruction to end the recording and recording operation is notified to the first control unit 110. When receiving the instruction to end the recording and recording operation, the first control unit 110 sends control information for shifting each component of the imaging device 100 to the shooting standby state.

[0037] The first control unit 110 stops the generation of image data by the imaging unit 101 and the audio processing of the audio data obtained by the first audio processing unit 103.

[0038] The encoding processing unit 107 reads the remaining image data and audio data stored in the memory 104, performs predetermined encoding processing, and stops the recording / recording operation when compressed and encoded image data and compressed and encoded audio data are generated. Similarly, when the generation of compressed and encoded image data ends even if the audio data is not compressed and encoded, the recording / recording operation is stopped.

[0039] The first control unit 110 synthesizes the compressed and encoded image data by the encoding processing unit 107 with the compressed and encoded audio data or the uncompressed audio data to form a data stream, and outputs it to the recording unit 108. Then, after the data stream is written to the recording medium 109 as one video file by the recording unit 108, the first control unit 110 stops the recording / recording operation.

[0040] When the first control unit 110 stops the recording / recording operation, each component of the imaging device 100 returns to the shooting standby state.

[0041] The video file generated by the imaging device 100 can be reproduced as a stereoscopic video by a reproducing device such as a head-mounted display. Also, audio is synthesized into the reproduced video, and audio corresponding to the orientation of the head of the user wearing the head-mounted display as the reproducing device is output, thereby providing a video with audio having a sense of presence.

[0042] The above is the configuration of the imaging device 100 and the basic operation during recording.

[0043] Next, with reference to FIGS. 2 to 5, the recording / recording operation of the imaging device 100 and the sound collection operation of the sound collection device 200 according to Embodiment 1 will be described.

[0044] FIG. 2 is a block diagram illustrating the configuration of the sound collection device 200 according to Embodiment 1.

[0045] The microphone unit 201 includes a plurality of microphone pieces. In this embodiment, it includes four microphone pieces, namely a first microphone piece 201a, a second microphone piece 201b, a third microphone piece 201c, and a fourth microphone piece 201d. The microphone unit 201 can generate stereophonic data (3D audio data) in the ambisonics method from the audio data simultaneously picked up by the first microphone piece 201a, the second microphone piece 201b, the third microphone piece 201c, and the fourth microphone piece 201d. The ambisonics method converts the audio data in A format picked up by four cardioid microphone pieces into audio data in B format with four channels, namely an omnidirectional component (W component) and bidirectional components (X, Y, Z components) for front-back, left-right, and up-down directions. Then, using spherical harmonic functions, audio data with an arbitrary directivity and having more than four channels can be generated from the audio data in B format.

[0046] Fig. 3(a) illustrates the external configuration of the microphone unit 201. As shown in Fig. 3(a), the four microphone pieces 201a to 201d of the microphone unit 201 are provided so as to face the four vertices of a cube. The three-dimensional space in Fig. 3(a) is defined by the directions of FRONT, REAR, LEFT, RIGHT, UP, and DOWN shown in Fig. 3(b). In this case, the first microphone piece 201a faces the front upper left (FLU), the second microphone piece 201b faces the front lower right (FRD), the third microphone piece 201c faces the back lower left (BLD), and the fourth microphone piece 201d faces the back upper right (BRU). Since the four microphone pieces 201a to 201d are microphones with cardioid directivity, they can pick up the audio signals in the four directions they face, and the audio signals in the four directions are called the A format of ambisonics.

[0047] The second audio processing unit 202 performs various signal processes on the analog audio signal acquired by the microphone unit 201. The second audio processing unit 202 amplifies the analog audio signal by the signal amplification unit 2021 and converts it into a digital audio signal by the A / D converter 2022.

[0048] The second control unit 203 includes a processor such as a CPU that performs arithmetic processing and control processing related to the sound collection device 200, a non-volatile memory that stores programs executed by the processor, data referred to by the programs, etc., and a volatile memory into which the programs and reference data stored in the non-volatile memory are loaded. The non-volatile memory is an EEROM or a flash memory. The volatile memory is a DRAM. The non-volatile memory and the volatile memory may be in the form of an internal memory built into the sound collection device 200 or in the form of an external memory externally connected to the sound collection device 200. Note that the program of the present embodiment includes a program for executing the flowcharts described later with reference to FIGS. 5, 7, and 9.

[0049] The second communication unit 204 is communicably connected to an external device such as the imaging device 100 by a wireless communication method such as Wi-Fi (registered trademark) or Bluetooth (registered trademark) or a wired communication method such as USB, and transmits and receives control information, voice data, etc.

[0050] The second attitude detection unit 205 detects the attitude information of the sound collection device 200 and sends it to the second control unit 203. The second control unit 203 detects the orientation (sound collection direction) of the sound collection device 200 based on the attitude information obtained from the second attitude detection unit 205. The second attitude detection unit 205 includes an acceleration sensor, an angular velocity sensor, a geomagnetic sensor, etc.

[0051] Next, with reference to FIG. 4, the relationship between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 in Embodiment 1 will be described.

[0052] The imaging device 100 and the sound collection device 200 can be operated and moved independently of each other. In this embodiment, for example, it is assumed that a user holds the imaging device 100 in one hand and the sound collection device 200 in the other hand, and performs recording and sound recording while walking. The imaging device 100 and the sound collection device 200 are connected by a wireless communication method, and can communicate with each other such as control information. Also, the imaging device 100 can receive audio data from the sound collection device 200. The shooting direction of the imaging device 100 is the front direction of the lens unit 300 of the imaging device 100. The sound collection direction of the sound collection device 200 is the front direction of the microphone unit 201 of the sound collection device 200, which is the FRONT direction shown in FIG. 3(b). Also, since the imaging device 100 and the sound collection device 200 can be operated and moved independently of each other, depending on the situation of recording and sound recording, there are cases where the shooting direction and the sound collection direction are not in the same direction (Φ = 0) as shown in FIG. 4.

[0053] Next, with reference to FIG. 5, the control processing of the imaging device 100 and the sound collection device 200 according to Embodiment 1 will be described.

[0054] FIG. 5(a) is a flowchart illustrating the control processing of the imaging device 100 according to Embodiment 1. FIG. 5(b) is a flowchart illustrating the control processing of the sound collection device 200 according to Embodiment 1.

[0055] The processing in FIG. 5(a) is realized by the first control unit 110 loading a program stored in the non-volatile memory into the volatile memory and executing it. Also, the processing in FIG. 5(b) is realized by the second control unit 203 loading a program stored in the non-volatile memory into the volatile memory and executing it. The same applies to FIGS. 7 and 9 described later. Also, the processing in FIGS. 5(a) and 5(b) is started in a state where the imaging device 100 is communicably connected to the sound collection device 200. Also, in FIG. 5(a), the recording process in the recording / recording operation of the imaging device 100 will be described, but it is assumed that the recording process is also executed simultaneously. The same applies to FIGS. 7(a) and 9(a) described later.

[0056] In step S101, in response to the recording / recording start instruction being notified from the operation unit 111, the first control unit 110 starts the recording process.

[0057] In step S102, the first control unit 110 starts detecting the posture of the imaging device 100 by the first posture detection unit 113. For the posture detection, for example, using a three-axis geomagnetic sensor, information such as the posture information of the imaging device 100 with respect to the azimuth, for example, information that the shooting direction of the imaging device 100 is facing northward, can be obtained.

[0058] In step S103, the first control unit 110 transmits an instruction to start the sound collection operation to the sound collection device 200 through the first communication unit 112.

[0059] In step S104, the second control unit 203 waits until it receives an instruction to start the sound collection operation from the imaging device 100. When it receives the instruction to start the sound collection operation, the process proceeds to step S105.

[0060] In step S105, the second control unit 203 starts detecting the posture of the sound collection device 200 by the second posture detection unit 205. For the posture detection, for example, using a three-axis geomagnetic sensor, information such as the posture information of the sound collection device 200 with respect to the azimuth, for example, information that the sound collection direction of the sound collection device 200 is facing northward, can be obtained.

[0061] In step S106, the second control unit 203 starts the sound collection operation by the sound collection device 200. The sound collection device 200 acquires sound at arbitrary regular intervals and executes the second audio processing. The sound collection operation continues until it receives an instruction to stop the sound collection operation from the imaging device 100.

[0062] In step S107, the second control unit 203 transmits the posture information of the sound collection device 200 obtained in step S105 and the audio data obtained in step S106 as Ambisonics A-format data to the imaging device 100 through the second communication unit 204. The posture information shall include time information synchronized with the section where the sound was collected.

[0063] In step S108, the first control unit 110 receives, via the first communication unit 112, the audio data acquired by the sound collection device 200 from the sound collection device 200 and the attitude information of the sound collection device 200.

[0064] In step S109, the first control unit 110 converts the audio data in the A format of ambisonics received in step S108 into the B format. The B format is audio data including the omnidirectional component and the bidirectional component in the acoustic technology based on the ambisonics method, and is a data format composed of the omnidirectional signal W in all directions, the front-back directional signal X, the left-right directional signal Y, and the up-down directional signal Z. The conversion from the A format to the B format is performed according to the following formula 1. (Formula 1) W = FLU + FRD + BLD + BRU X = FLU + FRD - BLD - BRU Y = FLU - FRD + BLD - BRU Z = FLU - FRD - BLD + BRU W: Omnidirectional signal (B format) X: Front-back bidirectional signal (B format) Y: Left-right bidirectional signal (B format) Z: Up-down bidirectional signal (B format) FLU: Audio signal in the upper left front acquired by the first microphone piece (A format) FRD: Audio signal in the lower right front acquired by the second microphone piece (A format) BLD: Audio signal in the lower left back acquired by the third microphone piece (A format) BRU: Audio signal in the upper right back acquired by the fourth microphone piece (A format) In step S110, the first control unit 110 calculates the difference between the posture of the imaging device 100 and the posture of the sound collection device 200 in the sound collection section based on the posture information of the imaging device 100 acquired by the first posture detection unit 113 and the posture information of the sound collection device 200 received in step S108. Then, the first control unit 110 calculates the difference between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 from the difference between the postures of the imaging device 100 and the sound collection device 200. In the present embodiment, since the first posture detection unit 113 and the second posture detection unit 205 perform posture detection using three-axis geomagnetic sensors, the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 can be relatively compared in the same coordinate system. In the example of FIG. 4, the shooting direction when the X direction is north is detected by the first posture detection unit 113, and the sound collection direction is detected by the second posture detection unit 205. Based on the information obtained from these posture detection units, the difference Φ between the shooting direction and the sound collection direction can be calculated. The example of FIG. 4 illustrates that there is a difference Φ between the shooting direction and the sound collection direction on the XY plane.

[0065] In step S111, the first control unit 110 corrects the B-format audio data generated in step S109 based on the difference Φ between the shooting direction and the sound collection direction calculated in step S110 by the audio correction unit 1032 of the first audio processing unit 103.

[0066] The correction of the audio data is performed by coordinate conversion in the XY plane according to the following formula 2. (Formula 2) TIFF2025082063000002.tif51126 Formula 3 is a correction formula in the XY plane. Similarly, the correction in the XZ plane is performed by formula 3, and the correction in the YZ plane is performed by formula 4. (Formula 3) TIFF2025082063000003.tif51127(Formula 4) TIFF2025082063000004.tif51127 In step S112, the first control unit 110 synthesizes the B-format audio data corrected in step S111 with the image data generated by the simultaneously executed recording process to form a data stream, and records it on the recording medium 109.

[0067] Note that the processes from step S101 to S111 are executed before the recording operation starts.

[0068] In step S113, the first control unit 110 determines whether it has received a stop instruction notification for the recording / recording operation from the operation unit 111. If the first control unit 110 has received a stop instruction notification for the recording / recording operation from the operation unit 111, the process proceeds to step S114. If the first control unit 110 has not received a stop instruction notification for the recording / recording operation from the operation unit 111, the process returns to step S108 to continue the recording / recording operation.

[0069] In step S114, the first control unit 110 transmits a stop instruction for the sound collection operation to the sound collection device 200 through the first communication unit 112.

[0070] In step S115, the second control unit 203 determines whether it has received a stop instruction for the sound collection operation from the imaging device 100. If the second control unit 203 has received a stop instruction for the sound collection operation from the imaging device 100, the process proceeds to step S116. If the second control unit 203 has not received a stop instruction for the sound collection operation from the imaging device 100, the process returns to step S106 to continue the sound collection operation.

[0071] In step S116, the second control unit 203 stops the sound collection operation. The second control unit 203 sends control information for stopping the sound collection operation to each component of the sound collection device 200.

[0072] In step S117, the first control unit 110 executes a stop process for the recording / recording operation. The first control unit 110 transmits control information for stopping the recording / recording operation to each component of the imaging device 100.

[0073] According to the above-described Embodiment 1, when the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 do not match, based on the difference between the shooting direction and the sound collection direction, the sound data acquired by the sound collection device 200 is corrected so that the shooting direction and the sound collection direction match. Although it is desirable to correct the sound data so that the shooting direction and the sound collection direction match, the present invention is not limited to this, and the difference between the shooting direction and the sound collection direction may be equal to or less than a predetermined threshold value, and correction may be performed so that the shooting direction and the sound collection direction approach each other. Thereby, the deviation between the shooting direction and the sound collection direction can be corrected, and the direction of the sound with respect to the moving image can be aligned. Therefore, it is possible to reduce the sense of discomfort experienced by the user due to the misalignment of the direction of the sound with respect to the moving image during moving image playback.

[0074] [Embodiment 2] Next, Embodiment 2 will be described with reference to FIGS. 6 and 7.

[0075] Note that the configurations and basic operations of the imaging device 100 and the sound collection device 200 in Embodiment 2 are the same as those in FIGS. 1 and 2 of Embodiment 1.

[0076] FIG. 6 is a diagram for explaining the relationship between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 in Embodiment 2.

[0077] In Embodiment 2, the first communication unit 112 of the imaging device 100 and the second communication unit 204 of the sound collection device 200 can detect the direction in which the other party is located with respect to their own position by wireless communication. The direction detection is realized, for example, by a direction detection function compliant with Bluetooth 5.1 (registered trademark), but is not limited to this, and may be realized by other methods.

[0078] FIG. 7(a) is a flowchart illustrating the control process of the imaging device 100 in Embodiment 2. FIG. 7(b) is a flowchart illustrating the control process of the sound collection device 200 in Embodiment 1.

[0079] Step S201 is the same process as step S101 in FIG. 5(a).

[0080] In step S202, the first control unit 110 detects the attitude of the imaging device 100 by the first attitude detection unit 113. Also, the direction in which the sound collection device 200 is located with respect to the position of the imaging device 100 is detected by the direction detection function of the first communication unit 112. Thereafter, the detection of the attitude of the imaging device 100 by the first attitude detection unit 113 and the detection of the direction in which the sound collection device 200 is located by the direction detection function of the first communication unit 112 are repeatedly executed. The attitude of the imaging device 100 is detected by, for example, a three-axis acceleration sensor. Since the mounting state of the first communication unit 112 in the imaging device 100 is known, as shown in FIG. 6(a), the angle θ formed by the direction in which the sound collection device 200 is located with respect to the position of the imaging device 100 and the imaging direction of the imaging device 100 is uniquely determined.

[0081] Steps S203 and S204 are the same processes as steps S103 and S104 in FIG. 5(a).

[0082] In step S205, the second control unit 203 detects the attitude of the sound collection device 200 by the second attitude detection unit 205. Also, the direction in which the imaging device 100 is located with respect to the position of the sound collection device 200 is detected by the direction detection function of the second communication unit 204. Thereafter, the detection of the attitude of the sound collection device 200 by the second attitude detection unit 205 and the detection of the direction in which the imaging device 100 is located by the direction detection function of the second communication unit 204 are repeatedly executed. The attitude of the sound collection device 200 is detected by, for example, a three-axis acceleration sensor. Since the mounting state of the second communication unit 204 in the sound collection device 200 is known, as shown in FIG. 6(a), the angle ψ formed by the direction in which the imaging device 100 is located with respect to the position of the sound collection device 200 and the sound collection direction of the sound collection device 200 is uniquely determined.

[0083] Step S206 is the same process as step S106 in FIG. 5(a).

[0084] In step S207, the second control unit 203 transmits, via the second communication unit 204, the direction information of the imaging device 100 with respect to the sound collection device 200, the sound collection direction information of the sound collection device 200, and the audio data in A format acquired in step S206 to the imaging device 100. The direction information of the imaging device 100 and the sound collection direction information of the sound collection device 200 include time information synchronized with the sound collection section.

[0085] In step S208, the first control unit 110 receives, via the first communication unit 112, from the sound collection device 200, the direction information of the imaging device 100 with respect to the sound collection device 200, the attitude information of the sound collection device 200, and the audio data in A format.

[0086] Step S209 is the same process as step S109 in FIG. 5(a).

[0087] In step S210, the first control unit 110 calculates the difference Φ between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200.

[0088] Here, with reference to FIG. 6(b), a method for calculating the difference θ between the shooting direction and the sound collection direction will be described.

[0089] As shown in FIG. 6(b), the difference Φ between the shooting direction and the sound collection direction is calculated by the following formula 5. (Formula 5) Φ = θ - (180 - ψ) = θ + ψ - 180 In step S211, the first control unit 110 corrects the audio data in B format converted in step S209 based on the difference Φ between the shooting direction and the sound collection direction calculated in step S210 by the audio correction unit 1032 of the first audio processing unit 103. The correction method is as described in step S111 of FIG. 5(a).

[0090] Steps S212 to S217 are the same processes as steps S112 to S117 in FIG. 5(a). Also, the processes of steps S201 to S211 are executed before the recording operation is started.

[0091] According to the above-described Embodiment 2, the imaging device 100 and the sound collection device 200 can detect the direction in which the other is located with respect to their own positions by wireless communication. Then, based on the direction information of the sound collection device 200 with respect to the imaging device 100 and the direction information of the imaging device 100 with respect to the sound collection device 200, the difference between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 is calculated. When the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 do not match, based on the difference between the shooting direction and the sound collection direction, the audio data acquired by the sound collection device 200 is corrected so that the shooting direction and the sound collection direction match. Although it is desirable to correct the audio data so that the shooting direction and the sound collection direction match, the present invention is not limited to this, and the difference between the shooting direction and the sound collection direction may be corrected so that it is equal to or less than a predetermined threshold value and the shooting direction and the sound collection direction approach each other. Thereby, the deviation between the shooting direction and the sound collection direction can be corrected, and the direction of the sound with respect to the moving image can be adjusted, so that it is possible to reduce the discomfort felt by the user due to the mismatch between the direction of the sound with respect to the moving image during moving image playback.

[0092] In the above-described Embodiments 1 and 2, in the imaging device 100, the process of converting the audio data in A format acquired by the sound collection device 200 into B format and the correction process of the audio data in B format were executed. However, the conversion process and the correction process of the audio data may be executed in the sound collection device 200, and the processed data may be transmitted to the imaging device 100.

[0093] [Embodiment 3] Next, Embodiment 3 will be described with reference to FIGS. 8 and 9.

[0094] The configurations and basic operations of the imaging device 100 and the sound collection device 200 in Embodiment 3 are the same as those in FIGS. 1 and 2 of Embodiment 1.

[0095] FIG. 8 is a diagram illustrating the external configuration of the sound collection device 200 according to Embodiment 3.

[0096] In Embodiment 3, the sound collection device 200 includes a plurality of detection marks 301, 302, and 303 for detecting the sound collection direction. The detection marks 301, 302, and 303 are provided on the surface of the housing of the sound collection device 200 and have sizes, shapes, and colors that can be identified by the image data captured by the imaging device 100. The detection marks 301 to 303 each have at least one of different sizes, shapes, and colors. Also, the detection marks are not limited to the sizes, shapes, and colors illustrated in FIG. 8. Also, the number of detection marks of the sound collection device 200 is not limited to three, and may be three or less, or three or more.

[0097] FIG. 9(a) is a flowchart illustrating the control process of the imaging device 100 according to Embodiment 3. FIG. 9(b) is a flowchart illustrating the control process of the sound collection device 200 according to Embodiment 1.

[0098] Step S301 is the same process as step S101 in FIG. 5(a).

[0099] In step S302, the first control unit 110 starts detecting the sound collection direction of the sound collection device 200 based on the image captured by the imaging unit 101. The detection marks 301, 302, and 303 are used for detecting the sound collection direction. The imaging device 100 identifies the detection marks 301, 302, and 303 from the image data obtained by imaging the sound collection device 200, and detects the sound collection direction of the sound collection device 200 based on the positions of the detection marks 301 to 303 of the sound collection device 200. The sound collection direction with respect to the positions of the detection marks of the sound collection device 200 may be stored in advance in the non-volatile memory of the first control unit 110 or the like, or may be acquired from the sound collection device 200 via the first communication unit 112. The detection of the sound collection direction is realized by a known object detection method such as R-CNN, YOLO, or SSD.

[0100] Steps S303 and S304 are the same processes as steps S103 and S104 in FIG. 5(a).

[0101] In step S305, the second control unit 203 starts detecting the posture of the sound collection device 200 by the second posture detection unit 205. The posture of the sound collection device 200 is detected by, for example, a three-axis acceleration sensor or a three-axis gyro sensor.

[0102] Step S306 is the same process as step S106 in Fig. 5(a).

[0103] In step S307, the second control unit 203 transmits the posture information of the sound collection device 200 acquired in step S305 and the audio data in A format acquired in step S306 to the imaging device 100 by the second communication unit 204. The posture information of the sound collection device 200 includes time information synchronized with the sound collection section and the displacement amount of the posture of the sound collection device 200 in the sound collection section.

[0104] Steps S308 and S309 are the same processes as steps S108 and S109 in Fig. 5(a).

[0105] In step S310, the first control unit 110 calculates the difference Φ between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 based on the sound collection direction of the sound collection device 200 detected in step S302 and the posture information of the sound collection device 200 received in step S308. The shooting direction of the imaging device 100 is, for example, the front direction of the lens unit 300 attached to the imaging device 100. Let the sound collection direction of the sound collection device 200 detected in step S302 be Φ0, and the posture information of the sound collection device 200 received in step S308, that is, the displacement amount of the posture of the sound collection device 200 be ΔΦ. Then, the difference Φ between the shooting direction and the sound collection direction is calculated by the following formula 6. (Formula 6) Φ = Φ0 + ΔΦ In step S311, the first control unit 110 corrects the audio data in B format converted in step S309 based on the difference Φ between the shooting direction and the sound collection direction calculated in step S310 by the audio correction unit 1032 of the first audio processing unit 103. The correction method is as described in step S111 of Fig. 5(a).

[0106] Steps S312 to S317 are the same processes as steps S112 to S117 in FIG. 5(a). Also, the processes from step S301 to S312 are executed before the recording operation starts.

[0107] According to the above-described Embodiment 3, the imaging device 100 identifies the detection mark 301, the detection mark 302, and the detection mark 303 from the image data obtained by imaging the sound collection device 200. Then, the imaging device 100 detects the sound collection direction of the sound collection device 200 based on the positions of the detection marks 301 to 303 of the sound collection device 200, and calculates the difference between the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200. When the shooting direction of the imaging device 100 and the sound collection direction of the sound collection device 200 do not match, based on the difference between the shooting direction and the sound collection direction, the audio data acquired by the sound collection device 200 is corrected so that the shooting direction and the sound collection direction match. Although it is desirable to correct the audio data so that the shooting direction and the sound collection direction match, the present invention is not limited to this. For example, the correction may be performed so that the difference between the shooting direction and the sound collection direction becomes equal to or less than a predetermined threshold value, and the shooting direction and the sound collection direction approach each other. Thereby, the deviation between the shooting direction and the sound collection direction can be corrected, and the direction of the sound with respect to the moving image can be aligned. Therefore, it is possible to reduce the discomfort felt by the user due to the misalignment of the direction of the sound with respect to the moving image during moving image playback. In Embodiment 3, the sound collection direction of the sound collection device 200 is detected from the image obtained by imaging the sound collection device 200 at the start of recording, and the change in the sound collection direction of the sound collection device 200 is detected based on the displacement amount of the posture of the sound collection device 200 during recording. However, the present invention is not limited to this. For example, the change in the sound collection direction of the sound collection device 200 during recording may be calculated from the image obtained by imaging the sound collection device 200. Further, in the process of imaging the sound collection device 200, the execution permission of the imaging process may be switched according to whether the sound collection device 200 is within the shooting range (field angle) of the imaging device 100.

[0108] [Other Embodiments] The present invention can also be realized by supplying a program that implements one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that implements one or more functions.

[0109] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the claims are appended to disclose the scope of the invention.

[0110] The disclosure of this specification includes the following imaging device, control method, and program. [Configuration 1] An imaging device, comprising: connection means for connecting to a sound collection device that collects audio data in different directions; audio processing means for generating stereophonic audio data based on the audio data acquired from the sound collection device; correction means for correcting the stereophonic audio data based on a difference between a shooting direction of the imaging device and a sound collection direction of the sound collection device. The imaging device is characterized by having the above. [Configuration 2] detection means for detecting attitude information of the imaging device; communication means for acquiring the audio data and attitude information of the sound collection device from the sound collection device, The correction means corrects based on the attitude information of the imaging device and the attitude information of the sound collection device. The imaging device according to Configuration 1 is characterized by this. [Configuration 3] The shooting direction of the imaging device is obtained from the attitude information of the imaging device, and the sound collection direction of the sound collection device is obtained from the attitude information of the sound collection device. The imaging device according to Configuration 2 is characterized by this. [Configuration 4] detection means for detecting a direction of the sound collection device with respect to the imaging device; Communication means for obtaining the voice data from the sound collection device and information regarding the direction of the imaging device with respect to the sound collection device, The correction means performs the correction based on the direction of the sound collection device with respect to the imaging device and the information regarding the direction of the imaging device with respect to the sound collection device. The imaging device according to Configuration 1. [Configuration 5] The difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device is obtained from the direction of the sound collection device with respect to the imaging device and the direction of the imaging device with respect to the sound collection device. The imaging device according to Configuration 4. [Configuration 6] The imaging device according to Configuration 1, further comprising detection means for detecting the sound collection direction of the sound collection device from an image of the sound collection device. [Configuration 7] The detection means identifies a plurality of detection marks provided on the sound collection device from the image data of the sound collection device, and detects the sound collection direction of the sound collection device based on the positions of the plurality of detection marks. The imaging device according to Configuration 6. [Configuration 8] The plurality of detection marks are different in at least one of size, shape, and color. The imaging device according to Configuration 7. [Configuration 9] Communication means for obtaining the voice data from the sound collection device and the attitude information of the sound collection device, The detection means obtains the difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device based on the sound collection direction of the sound collection device and the attitude information of the sound collection device. The imaging device according to Configuration 7 or 8. [Configuration 10] The correction means corrects the stereophonic sound data so that the shooting direction of the imaging device and the sound collection direction of the sound collection device coincide. The imaging device according to any one of Configurations 1 to 9. [Configuration 11] The correction means corrects the stereophonic data so that the difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device is equal to or less than a predetermined threshold value. The imaging device according to any one of configurations 1 to 10. [Configuration 12] The audio data acquired by the sound collection device is audio data in the A format of Ambisonics. The audio processing means converts the audio data in the A format into audio data in the B format. The correction means corrects the audio data in the B format. The imaging device according to any one of configurations 1 to 11. [Configuration 13] Imaging means; Generating means for generating a moving image file by synthesizing the image data captured by the imaging means and the audio data generated by the audio processing means. The imaging device according to any one of configurations 1 to 12. [Configuration 14] The connection means connects to the sound collection device by wireless communication. The imaging device according to any one of configurations 1 to 13. [Configuration 15] A control method for an imaging device, comprising: Connecting to a sound collection device that collects audio data in different directions; Generating stereophonic data based on the audio data acquired from the sound collection device; Correcting the stereophonic data based on the difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device. A control method characterized by comprising. [Configuration 16] A program for causing a computer to function as the imaging device according to any one of configurations 1 to 14.

Explanation of Signs

[0111] 100... Imaging device, 101... Imaging unit, 103... First audio processing unit, 110... First control unit, 112... First communication unit, 200... Sound collection device, 201... Sound collection unit, 202... Second audio processing unit, 203... Second control unit, 204... Second communication unit

Claims

1. An imaging device, connection means for connecting to a sound collection device that collects audio data in different directions; audio processing means for generating stereophonic audio data based on the audio data acquired from the sound collection device; and correction means for correcting the stereophonic audio data based on a difference between a shooting direction of the imaging device and a sound collection direction of the sound collection device. The imaging device is characterized by having these components.

2. detection means for detecting attitude information of the imaging device; and communication means for acquiring the audio data and the attitude information of the sound collection device from the sound collection device. The correction means performs the correction based on the attitude information of the imaging device and the attitude information of the sound collection device. The imaging device according to claim 1 is characterized by this.

3. The shooting direction of the imaging device is obtained from the attitude information of the imaging device, and the sound collection direction of the sound collection device is obtained from the attitude information of the sound collection device. The imaging device according to claim 2 is characterized by this.

4. detection means for detecting a direction of the sound collection device with respect to the imaging device; and communication means for acquiring the audio data and information regarding the direction of the imaging device with respect to the sound collection device from the sound collection device. The correction means performs the correction based on the direction of the sound collection device with respect to the imaging device and the information regarding the direction of the imaging device with respect to the sound collection device. The imaging device according to claim 1 is characterized by this.

5. The difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device is obtained from the direction of the sound collection device with respect to the imaging device and the information regarding the direction of the imaging device with respect to the sound collection device. The imaging device according to claim 4 is characterized by this.

6. The imaging device according to claim 1, characterized by having detection means for detecting the sound collection direction of the sound collection device from an image of the sound collection device taken by the imaging device.

7. The detection means identifies a plurality of detection marks provided on the sound collection device from image data of the sound collection device taken by the imaging device, and detects the sound collection direction of the sound collection device based on the positions of the plurality of detection marks. The imaging device according to claim 6 is characterized by this.

8. The plurality of detection marks are characterized in that at least one of size, shape, and color is different. The imaging device according to claim 7 is characterized by this.

9. and communication means for acquiring the audio data and the attitude information of the sound collection device from the sound collection device. The imaging device according to claim 7, wherein the detection means obtains a difference between the sound collection direction of the sound collection device and the shooting direction of the imaging device based on the posture information of the sound collection device and the sound collection direction of the sound collection device.

10. The imaging device according to claim 1, wherein the correction means corrects the stereophonic sound data so that the shooting direction of the imaging device coincides with the sound collection direction of the sound collection device.

11. The imaging device according to claim 1, wherein the correction means corrects the stereophonic sound data so that the difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device is equal to or less than a predetermined threshold value.

12. The audio data acquired by the sound collection device is audio data in the A format of Ambisonics, the audio processing means converts the audio data in the A format into audio data in the B format, The imaging device according to claim 1, wherein the correction means corrects the audio data in the B format.

13. imaging means; generating means for generating a moving image file by synthesizing the image data captured by the imaging means and the audio data generated by the audio processing means, the imaging device according to claim 1.

14. The imaging device according to claim 1, wherein the connection means is connected to the sound collection device by wireless communication.

15. A control method for an imaging device, comprising: connecting to a sound collection device that collects audio data in different directions; generating stereophonic sound data based on the audio data acquired from the sound collection device; correcting the stereophonic sound data based on a difference between the shooting direction of the imaging device and the sound collection direction of the sound collection device, the control method being characterized by comprising the steps of:

16. A program for causing a computer to function as the imaging device according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Acoustic recording device, acoustic system, acoustic recording method, program, and data structure

    JP2018152846A

Cited By

  • Optical component and method for producing optical component

    WO2025263593A1