Image pickup method and image pickup device
By obtaining the first digit sound data corresponding to the image data and creating the second sound data that is one digit larger than the original sound data, the problem of low quality of sound data in the prior art is solved, and the high-quality correspondence effect between sound and image data is achieved.
Patent Information
- Application Number
- CN202380079760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-10-05
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to improve the quality of sound data, which makes the corresponding effect of sound and image data poor.
By obtaining the sound data of the first digit corresponding to the image data and creating a second sound data that is one digit larger than the original sound data based on the sound data, the quality of the sound data is improved.
The quality of sound data is improved, so that the corresponding effect between sound and image data is significantly improved.
Smart Images

Figure CN120226375A_ABST
Abstract
Description
Technical Field
[0001] The technology of the present invention relates to a camera method and a camera device. Background Art
[0002] Japanese Patent Application Laid-Open No. 2012-073435 discloses a sound signal conversion device that samples an input analog sound signal of an L channel and an R channel at a sampling frequency of 192 kHz and a quantization bit number of 24 bits by an A / D conversion device to generate a digital signal. A signal processing device is connected to the output side of the A / D conversion device. The signal processing device performs processing of downsampling the frequency to 1 / 4 (48 kHz) and processing of converting the downsampled signal into a floating-point format with a quantization bit number of 32 bits.
[0003] Japanese Patent Application Laid-Open No. 2002-246913 discloses a data processing device that converts input data from a fixed-point form to a floating-point form by a conversion unit. Summary of the Invention
[0004] Technical Problem to be Solved by the Invention
[0005] An object of an embodiment of the technology of the present invention is to provide a camera method and a camera device capable of improving the quality of sound data.
[0006] Means for Solving the Technical Problem
[0007] To achieve the above object, the camera method of the present invention includes: an acquisition step of acquiring first sound data of a first number of bits, the first sound data of the first number of bits being generated based on a sound signal output from a sound collection element and corresponding to image data obtained by photographing a subject; and a first creation step of creating second sound data of a second number of bits larger than the first number of bits based on the first sound data.
[0008] Preferably, the image data is moving image data.
[0009] Preferably, it further includes: a second creation step of creating third sound data of a third number of bits smaller than the second number of bits based on the second sound data; and a third creation step of creating a first image file including the image data and the third sound data.
[0010] Preferably, in the second creation step, the third sound data is created by extracting data within a volume range having a width of the third number of bits based on the second sound data.
[0011] Preferably, it further includes a setting step of setting a volume range for the second sound data at a predetermined time before executing the second creation step.
[0012] Preferably, in the setting process, the volume range is set according to the characteristics of the image data or the imaging conditions.
[0013] Preferably, in the setting process, the volume range is set according to the characteristics of the main subject imaged in the frames constituting the image data.
[0014] Preferably, in the setting process, the main subject is selected from among a plurality of subjects imaged in the frame according to the size or category of the subject in the frame, the position of the subject within the viewing angle, the focus position of the imaging lens that captures the subject, the input information of the user, or the line-of-sight information of the user.
[0015] Preferably, in the first creation process, the second audio data is created by synthesizing the first audio data generated by performing a first gain process on the audio signal and the first audio data generated by performing a second gain process with a gain amount different from that of the first gain process on the audio signal.
[0016] Preferably, the second audio data is in floating-point format.
[0017] Preferably, there are a first mode for creating a first image file and a second mode for creating a second image file including image data and second audio data, and there is also a process of recommending the execution of the second mode or a process of executing the second mode according to the volume information of the audio signal output from the sound collection element or the volume information determined according to the characteristics of the image data.
[0018] Preferably, there is also a gain adjustment process for performing a gain process on the audio signal. In the gain adjustment process, the gain amount of the gain process is changed according to the volume information of the sound collected by the sound collection element, the volume information determined according to the characteristics of the subject, or the priority information of the recorded sound source.
[0019] The imaging device of the present invention includes: an imaging element that captures a subject to generate image data; and a processor that acquires first audio data of a first number of bits and creates second audio data of a second number of bits larger than the first number of bits according to the first audio data. The first audio data of the first number of bits is generated according to the audio signal output from the sound collection element and corresponds to the image data.
[0020] Preferably, it has: a housing that houses the imaging element and the processor; and a connection part for connecting an external microphone including a sound collection element to the housing.
[0021] The audio signal can be converted into a digital signal inside the external microphone and then sent to the inside of the housing via the connection part. Description of the Drawings
[0022] Figure 1 It is a diagram showing an example of the structure of the imaging device.
[0023] Figure 2 This is a diagram showing an example of the configuration of a sound signal processing circuit.
[0024] Figure 3 FIG. 1 is a diagram conceptually showing audio signal processing by an audio signal processing circuit.
[0025] Figure 4 This is a diagram showing an example of the functional structure of a processor.
[0026] Figure 5 This is a diagram conceptually showing the synthesis process and the data format conversion process.
[0027] Figure 6 This is a diagram conceptually showing the feature extraction process and the volume range setting process.
[0028] Figure 7 This is a diagram conceptually showing the data extraction process.
[0029] Figure 8 This is a flowchart showing an example of the operation of the imaging device.
[0030] Figure 9 This is a flowchart showing an example of a mode setting process according to the third modification.
[0031] Figure 10 This is a flowchart showing an example of a gain adjustment process according to the fourth modification.
[0032] Figure 11 It is a diagram showing the structure of an imaging device according to a fifth variation. DETAILED DESCRIPTION
[0033] An example of an embodiment according to the technology of the present invention will be described with reference to the drawings.
[0034] First, the words and phrases used in the following description are explained.
[0035] In the following description, "AF" is the abbreviation of "Auto Focus". "MF" is the abbreviation of "Manual Focus". "IC" is the abbreviation of "Integrated Circuit". "CPU" is the abbreviation of "Central Processing Unit". "RAM" is the abbreviation of "Random Access Memory". "CMOS" is the abbreviation of "Complementary Metal Oxide Semiconductor".
[0036] "FPGA" is an abbreviation for "Field Programmable Gate Array", "PLD" is an abbreviation for "Programmable Logic Device", "ASIC" is an abbreviation for "Application Specific Integrated Circuit", "OVF" is an abbreviation for "Optical View Finder", "EVF" is an abbreviation for "Electronic View Finder", and "ADC" is an abbreviation for "Analog to Digital Converter", and "LPCM" is an abbreviation for "Linear Pulse Code Modulation".
[0037] As an embodiment of the imaging device, a lens interchangeable digital camera is exemplified to describe the technology of the present invention. In addition, the technology of the present invention is not limited to lens interchangeable type, and can also be applied to a lens-integrated digital camera.
[0038] Figure 1 An example of the structure of the imaging device 10 is shown. The imaging device 10 is a lens interchangeable digital camera. The imaging device 10 is composed of a housing 11 and an imaging lens 12 that is interchangeably mounted on the housing 11 and includes a focusing lens 31. The imaging lens 12 is mounted on the front surface side of the housing 11 via a bayonet 11A.
[0039] In addition, an external microphone 13 can be detachably mounted on the housing 11. The external microphone 13 is mounted on the housing 11 via a connection portion 11B provided on the upper surface of the housing 11. The external microphone 13 is a shotgun microphone, a zoom microphone, etc. The connection portion 11B is, for example, a hot shoe.
[0040] An operation portion 16 including a dial, a release button, etc. is provided on the housing 11. As operation modes of the imaging device 10, for example, a still image shooting mode, a moving image shooting mode, and an image display mode are included. The operation portion 16 is operated by the user when setting the operation mode. In addition, the operation portion 16 is operated by the user when starting to execute still image shooting or moving image shooting.
[0041] Further, the operation unit 16 is operated by the user when selecting a focusing mode. The focusing modes include an AF mode and an MF mode. The AF mode is a mode in which focusing control is performed by setting a subject area selected by the user or a subject area automatically detected by the imaging device 10 as a focus detection area (hereinafter referred to as an AF area). The MF mode is a mode in which the user manually performs focusing control by operating a focusing ring (not shown).
[0042] Further, a viewfinder 14 is provided on the housing 11. For example, the viewfinder 14 is a hybrid viewfinder (registered trademark). The hybrid viewfinder is, for example, a viewfinder that selectively uses an optical viewfinder (hereinafter referred to as an "OVF") and an electronic viewfinder (hereinafter referred to as an "EVF"). The user can observe an optical image or an instant preview image of the subject projected through the viewfinder 14 via a viewfinder eyepiece portion (not shown).
[0043] Further, a display 15 is provided on the back side of the housing 11. An image based on the image data PD obtained by shooting and various menu screens are displayed on the display 15. The user can also observe the instant preview image projected through the display 15 instead of observing the viewfinder 14.
[0044] The housing 11 and the imaging lens 12 are electrically connected via electrical contacts 11C provided in the bayonet 11A.
[0045] The imaging lens 12 includes a focusing lens 31, an aperture 32, and a lens drive control unit 33. The lens drive control unit 33 is electrically connected to a processor 25 housed in the housing 11 via the electrical contacts 11C.
[0046] The lens drive control unit 33 drives the focusing lens 31 and the aperture 32 according to a control signal sent from the processor 25. In order to adjust the position of the focusing lens 31, the lens drive control unit 33 performs drive control of the focusing lens 31 according to a focusing control signal sent from the processor 25.
[0047] The aperture 32 has an aperture with a variable opening diameter. In order to adjust the amount of incident light on the imaging sensor 20, the lens drive control unit 33 performs drive control of the aperture 32 according to an aperture adjustment signal sent from the processor 25.
[0048] Further, an imaging sensor 20, an image processing circuit 21, a built-in microphone 22, a sound signal processing circuit 23, a processor 25, and a storage device 26 are provided inside the housing 11. The imaging sensor 20, the image processing circuit 21, the built-in microphone 22, the sound signal processing circuit 23, the storage device 26, and the display 15 are controlled in operation by the processor 25.
[0049] The processor 25 is constituted by a CPU, for example. A RAM 25A serving as a primary storage memory is connected to the processor 25. The storage device 26 is constituted by a non-volatile memory such as a flash memory, for example. The processor 25 executes various processes according to a program 27 stored in the storage device 26. In addition, the processor 25 may be constituted by an aggregate of a plurality of IC chips. And, for example, an image file 28 generated as a result of the imaging operation performed by the imaging device 10 is stored in the storage device 26.
[0050] The imaging sensor 20 is a CMOS image sensor, for example. Light (subject image) that has passed through the imaging lens 12 is incident on a light-receiving surface 20A of the imaging sensor 20. A plurality of pixels that generate a imaging signal by performing photoelectric conversion are formed on the light-receiving surface 20A. The imaging sensor 20 generates and outputs image data PD by performing photoelectric conversion on the light incident on each pixel. In addition, the imaging sensor 20 is an example of an "imaging element" related to the technology of the present invention.
[0051] The image processing circuit 21 performs image processing including white balance correction, gamma correction processing, etc. on the image data PD output from the imaging sensor 20.
[0052] The built-in microphone 22 is a stereo microphone having a pair of sound collection elements 22A and 22B. The sound collection elements 22A and 22B are a sound sensor for the left channel (hereinafter referred to as the L channel) and a sound sensor for the right channel (hereinafter referred to as the R channel). The sound collection elements 22A and 22B are sound sensors such as electrostatic type, piezoelectric type, and electrodynamic type, and output the collected sound as a sound signal. The sound signal processing circuit 23 performs sound signal processing including gain processing, A / D conversion processing, etc. on the sound signals output from the sound collection elements 22A and 22B.
[0053] The external microphone 13 includes a sound collection element 41, an amplifier 42, and a microphone control unit 43. In the present embodiment, the external microphone 13 is a monaural microphone having one sound collection element 41. The sound collection element 41 is a sound sensor such as electrostatic type, piezoelectric type, and electrodynamic type, and outputs the collected sound as a sound signal. The amplifier 42 performs gain processing on the sound signal output from the sound collection element 41. The microphone control unit 43 controls the gain amount of the gain processing based on the amplifier 42.
[0054] And, the microphone control unit 43 supplies the sound signal that has been subjected to gain processing by the amplifier 42 to the sound signal processing circuit 23 in the housing 11 via the connection unit 11B. A monaural and analog sound signal AS is supplied from the external microphone 13 to the sound signal processing circuit 23. In addition, the operation of the microphone control unit 43 is controlled by the processor 25.
[0055] Figure 2An example of the structure of the sound signal processing circuit 23 is shown. The sound signal processing circuit 23 includes a first preamplifier 51A, a first ADC 52A, a second preamplifier 51B, and a second ADC 52B.
[0056] The first preamplifier 51A and the first ADC 52A are processing units for the L channel that perform gain processing and A / D conversion processing on the sound signal output from the sound collection element 22A included in the built-in microphone 22. The second preamplifier 51B and the second ADC 52B are processing units for the R channel that perform gain processing and A / D conversion processing on the sound signal output from the sound collection element 22B included in the built-in microphone 22.
[0057] The gain amount G1 of the first preamplifier 51A is controlled by the processor 25. The gain amount G2 of the second preamplifier 51B is controlled by the processor 25. When performing gain processing on the sound signal output from the built-in microphone 22, the gain amount G1 and the gain amount G2 are set to the same value by the processor 25. The first ADC 52A and the second ADC 52B convert the analog sound signal into a digital signal in the form of 24-bit LPCM by sampling with, for example, a quantization bit number of 24 bits. Additionally, it is an example of the pulse code modulation form.
[0058] When the external microphone 13 is connected to the connection part 11B, the sound signal AS output from the external microphone 13 is input to the sound signal processing circuit 23. In this case, the operation of the built-in microphone 22 is invalidated, and the sound signal is not input from the built-in microphone 22 to the sound signal processing circuit 23.
[0059] The sound signal AS output from the external microphone 13 is input to the first preamplifier 51A and the second preamplifier 51B. The first preamplifier 51A performs gain processing on the sound signal AS with the gain amount G1. The second preamplifier 51B performs gain processing on the sound signal AS with the gain amount G2. When performing gain processing on the sound signal AS output from the external microphone 13, the gain amount G1 and the gain amount G2 are set to different values by the processor 25. Hereinafter, the gain processing performed by the first preamplifier 51A is referred to as the first gain processing, and the gain processing performed by the second preamplifier 51B is referred to as the second gain processing.
[0060] The first ADC 52A converts the sound signal AS, which has undergone first gain processing by the first preamplifier 51A, into a digital signal. The second ADC 52B converts the sound signal AS, which has undergone second gain processing by the second preamplifier 51B, into a digital signal. Hereinafter, the sound signal AS digitized by the first ADC 52A is referred to as the first sound data AS1H, and the sound signal AS digitized by the second ADC 52B is referred to as the first sound data AS1L. The first sound data AS1H and AS1L are output from the sound signal processing circuit 23 to the processor 25.
[0061] Figure 3 Conceptually shows the sound signal processing based on the sound signal processing circuit 23. The sound signal AS output from the external microphone 13 is input to the L-channel processing unit and the R-channel processing unit. The sound signal AS input to the L-channel processing unit is converted into a digital signal after being subjected to first gain processing with a gain amount G1, and thus is output from the sound signal processing circuit 23 as the first sound data AS1H. The sound signal AS input to the R-channel processing unit is converted into a digital signal after being subjected to second gain processing with a gain amount G2, and thus is output from the sound signal processing circuit 23 as the first sound data AS1L. In the present embodiment, the number of bits (hereinafter referred to as the first number of bits) of the first sound data AS1H and AS1L is 24 bits.
[0062] For example, the gain amount G1 is set to +48 dB, and the gain amount G2 is set to -48 dB. 48 dB corresponds to a volume width of 8 bits. Therefore, as Figure 3 shown, there is a deviation of 16 bits between the high-gain first sound data AS1H and the low-gain first sound data AS1L. In other words, there is an overlap of 8 bits between the first sound data AS1H and the first sound data AS1L.
[0063] Figure 4 Shows an example of the functional structure of the processor 25. The processor 25 realizes various functional units by executing processing according to the program 27 stored in the storage device 26. Figure 4 The various functional units shown are realized in the dynamic image shooting mode. As Figure 4 shown, for example, the processor 25 realizes the main control unit 60, the synthesis processing unit 61, the data format conversion unit 62, the volume range setting unit 63, the data extraction unit 64, the video file creation unit 65, and the feature extraction unit 66.
[0064] The main control unit 60 uniformly controls each part of the imaging device 10. The main control unit 60 controls the operation of the imaging device 10 according to the instruction signal input from the operation unit 16. The main control unit 60 controls the imaging sensor 20 to cause the imaging sensor 20 to perform an imaging operation. The imaging sensor 20 outputs image data PD generated by taking a picture via the imaging lens 12. In the moving image shooting mode, the imaging sensor 20 outputs the image data PD in one frame period. The image data PD output from the imaging sensor 20 is input to the processor 25 after being subjected to image processing by the image processing circuit 21. In the case of the moving image shooting mode, the image data PD is moving image data composed of a plurality of frames.
[0065] Moreover, in the moving image shooting mode, when the external microphone 13 is connected to the connection part 11B, the main control unit 60 controls the external microphone 13 to perform a sound collection operation. The external microphone 13 outputs the sound signal AS to the sound signal processing circuit 23 via the connection part 11B during the imaging operation of the imaging sensor 20. The sound signal processing circuit 23 outputs the first sound data AS1H, AS1L of the first number of bits by performing the above-mentioned sound signal processing. That is, the first sound data AS1H, AS1L are the sound data corresponding to the image data PD obtained by the imaging sensor 20 shooting the subject.
[0066] The synthesis processing unit 61 acquires the first sound data AS1H, AS1L output from the sound signal processing circuit 23, and synthesizes the first sound data AS1H, AS1L, thereby creating the second sound data AS2 of the second number of bits larger than the first number of bits. The second sound data AS2 is digital data in the LPCM format.
[0067] The data format conversion unit 62 converts the data format of the second sound data AS2 into the floating point format. Hereinafter, the second sound data AS2 converted into the floating point format is referred to as the second sound data AS2F.
[0068] The volume range setting unit 63 sets a volume range VR having a width of the third number of bits smaller than the second number of bits for the dynamic range of the second sound data AS2F. In the present embodiment, the volume range setting unit 63 sets the volume range VR according to the feature extracted by the feature extraction unit 66. The feature extraction unit 66 calculates, for example, the luminance value of the image based on the image data PD, and supplies the calculated luminance value as a feature of the image data PD to the volume range setting unit 63.
[0069] The data extraction unit 64 extracts the data of the volume range VR set by the volume range setting unit 63 from the second sound data AS2F, thereby creating the third sound data AS3 of the third number of bits. The third sound data AS3 is digital data in the LPCM format.
[0070] The video file creation unit 65 creates a video file 28 that includes the video data PD output from the image processing circuit 21 and the third audio data AS3 corresponding to the video data PD, and stores it in the storage device 26. The video file 28 corresponds to the "first video file" related to the technology of the present invention.
[0071] Figure 5 Conceptually shows the synthesis process based on the synthesis processing unit 61 and the data format conversion process based on the data format conversion unit 62. The synthesis processing unit 61 synthesizes the first audio data AS1H and the first audio data AS1L by mixing the overlapping portions of the 8-bit amounts of the first audio data AS1H and the first audio data AS1L. The number of bits of the second audio data AS2 generated by this synthesis process (i.e., the second number of bits) is 40 bits. Thus, the second audio data AS2 with an expanded dynamic range of volume is obtained by synthesizing the first audio data AS1H and the first audio data AS1L with different synthesis gain amounts.
[0072] The data format conversion unit 62 converts the 40-bit fixed-point form second audio data AS2 into 32-bit floating-point form (so-called 32-bit floating-point number) second audio data AS2F. The 32-bit floating-point number consists of 1-bit sign, 8-bit exponent part, and 23-bit mantissa part. The conversion from the fixed-point form to the floating-point form can use a known method. A wide range of numerical expressions can be performed in the floating-point form.
[0073] Figure 6 Conceptually shows the feature extraction process based on the feature extraction unit 66 and the volume range setting process based on the volume range setting unit 63. The feature extraction unit 66 calculates the luminance value according to the video data PD at a specified time, and supplies the calculated luminance value as the feature of the video data PD to the volume range setting unit 63. The volume range setting unit 63 sets the volume range VR at a specified time according to the luminance value supplied from the feature extraction unit 66 at a specified time. For example, the number of bits representing the width of the volume range VR (i.e., the third number of bits) is 24 bits.
[0074] In Figure 6 the example shown, the larger the luminance value, the more the volume range setting unit 63 sets the volume range VR on the high-volume side, and the smaller the luminance value, the more the volume range setting unit 63 sets the volume range VR on the low-volume side. For example, the feature extraction unit 66 calculates the luminance value at a 1-frame period, and the volume range setting unit 63 sets the volume range VR at a 1-frame period. In addition, the time interval at which the feature extraction unit 66 calculates the luminance value and the time interval at which the volume range setting unit 63 sets the volume range VR may also be non-constant.
[0075] It is speculated that when the estimated brightness value is large, it is in a noisy environment during the day and the volume level of the sound with the subject as the sound source is relatively high. Therefore, the volume range VR is set on the high-volume side. On the other hand, it is speculated that when the estimated brightness value is large, it is in a quiet environment at night and the volume level of the sound with the subject as the sound source is relatively high. Therefore, the volume range VR is set on the low-volume side. In this way, by setting the volume range VR according to the volume level of the sound with the subject as the sound source, the sound with the subject as the sound source can be appropriately extracted.
[0076] Figure 7 Conceptually shows the data extraction process based on the data extraction unit 64. The data extraction unit 64 extracts the data of the volume range VR according to the second sound data AS2F, thereby creating the third sound data AS3 in 24-bit fixed-point format. Specifically, the data extraction unit 64 selects the values of the sign and exponent parts of the 32-bit floating-point number according to the volume range VR, thereby creating the 24-bit third sound data AS3 represented by the mantissa part.
[0077] Generally, the sound data included in the moving image file is 24-bit digital data in LPCM format. Therefore, in this embodiment, the 24-bit third sound data AS3 is created.
[0078] Figure 8 Is a flowchart showing an example of the operation of the imaging device 10. Figure 8 Represents the operation when the moving image shooting mode is selected as the operation mode and the external microphone 13 is connected to the connection part 11B.
[0079] First, the main control unit 60 determines whether there is an instruction to start moving image shooting by the user (step S10). If it is determined that there is a start instruction (step S10: Yes), the shooting process (step S11) and the acquisition process (step S12) are executed simultaneously. In the shooting process, the imaging sensor 20 shoots the subject to generate image data PD. In the acquisition process, the sound collecting element 41 of the external microphone 13 collects sound to obtain the first sound data AS1H and AS1L of the first number corresponding to the image data PD. In this embodiment, the first sound data AS1H and AS1L are generated by performing gain processing on the sound signal output from the sound collecting element 41 with different gain amounts in the sound signal processing circuit 23 and then performing A / D conversion.
[0080] After the acquisition process, the first creation process (step S13) is executed. In the first creation process, the synthesis processing unit 61 creates the second sound data AS2 of the second number larger than the first number according to the first sound data AS1H and AS1L. And in the first creation process, the second sound data AS2 is converted into the floating-point form second sound data AS2F by the data form conversion unit 62.
[0081] After the imaging process, a setting process (step S14) is performed. In the setting process, the volume range setting unit 63 sets a volume range VR having a width of the third digit for the second sound data AS2F. Specifically, the volume range VR is set according to the feature (brightness value in the present embodiment) of the image data PD extracted by the feature extraction unit 66.
[0082] After the first creation process and the setting process, a second creation process (step S15) is performed. In the second creation process, the data extraction unit 64 creates third sound data AS3 of the third digit based on the second sound data AS2F. Specifically, the third sound data AS3 is created by extracting data of the volume range VR according to the second sound data AS2F.
[0083] After the second creation process, the main control unit 60 determines whether there is an end instruction for the moving image shooting by the user (step S16). When it is determined that there is no end instruction (step S16: No), the process returns to steps S11 and S12. The steps S11 to S15 are repeatedly executed until it is determined that there is an end instruction in step S16.
[0084] When it is determined that there is an end instruction (step S16: Yes), a third creation process (step S17) is performed. In the third creation process, the image file creation unit 65 creates an image file 28 including the image data PD and the third sound data AS3, and stores it in the storage device 26. Thus, the operation of the imaging device 10 ends.
[0085] In addition, in Figure 8 the example shown, the second creation process is executed during the moving image shooting, but the second creation process may also be executed after the moving image shooting ends. In this case, the second sound data AS2F created during the moving image shooting and the set value of the volume range VR are stored in the RAM 25A. Then, after the moving image shooting ends, the data extraction unit 64 may read the second sound data AS2F and the set value of the volume range VR from the RAM 25A to create the third sound data AS3.
[0086] As described above, the imaging method of the present invention includes: an acquisition process of acquiring first sound data of the first digit, the first sound data of the first digit being generated based on a sound signal output from a sound collecting element and corresponding to image data obtained by photographing a subject; and a first creation process of creating second sound data of a second digit larger than the first digit based on the first sound data. Thereby, the quality of the sound data corresponding to the image data can be improved.
[0087] Hereinafter, various modification examples of the above-described embodiment will be described.
[0088] [First Modification Example]
[0089] In the above-described embodiment, the feature extraction unit 66 extracts the luminance value as a feature of the image data PD, but it is also possible to extract the features of the main subject imaged in the frames constituting the image data PD. The main subject refers to the subject determined by the user or the main control unit 60 to have a high importance among the multiple subjects imaged in the frame. In this modification example, the volume range setting unit 63 sets the volume range VR according to the features of the main subject.
[0090] The features of the main subject extracted by the feature extraction unit 66 are, for example, the category of the main subject. In this case, the volume range setting unit 63 sets the volume range VR according to the category of the main subject. For example, when the category of the main subject is a category where the sound emitted is considered to be a high volume, such as an "airplane", the volume range setting unit 63 sets the volume range VR on the high volume side. On the other hand, when the category of the main subject is a category where the sound emitted is considered to be a low volume, such as a "person", the volume range setting unit 63 sets the volume range VR on the low volume side.
[0091] As a process for the feature extraction unit 66 to select the main subject from the multiple subjects imaged in the frame, various processes can be applied. As an example, the feature extraction unit 66 selects the main subject according to the sizes of the multiple subjects imaged in the frame. In this case, the feature extraction unit 66 selects the subject with the largest size among the multiple subjects as the main subject.
[0092] Moreover, the feature extraction unit 66 can also select the main subject according to the categories of the multiple subjects imaged in the frame. In this case, for example, the feature extraction unit 66 determines the category of each subject and selects the subject that matches the category set by the user using the operation unit 16 as the main subject. For example, when the portrait photography mode is set, the feature extraction unit 66 selects a person as the main subject from the multiple subjects imaged in the frame.
[0093] Moreover, the feature extraction unit 66 can also select the main subject according to the positions of the multiple subjects imaged in the frame within the viewing angle. In this case, for example, the feature extraction unit 66 calculates the positions of each subject within the viewing angle and selects the subject located at the center of the viewing angle as the main subject.
[0094] Moreover, the feature extraction unit 66 can also select the main subject according to the focus position of the imaging lens 12. In this case, for example, the feature extraction unit 66 obtains information related to the focus position from the main control unit 60 and selects the subject closest to the focus position from the multiple subjects imaged in the frame as the main subject.
[0095] Moreover, the feature extraction unit 66 can also select the main subject according to the input information of the user. In this case, for example, the feature extraction unit 66 selects the subject located in the subject area selected by the user using the operation unit 16 as the main subject. In addition, the feature extraction unit 66 can also select the subject located in the AF area automatically detected by the imaging device 10 as the main subject.
[0096] Moreover, the feature extraction unit 66 can also select the main subject according to the gaze information of the user. In this case, for example, the imaging device 10 has a function of detecting the gaze of the user looking into the viewfinder 14. The feature extraction unit 66 acquires the gaze information of the user and selects the subject located at the position of the gaze within the viewing angle as the main subject.
[0097] Moreover, the feature extraction unit 66 can also identify the shooting scene based on the image data PD and use the identified shooting scene as a feature of the image data PD. The shooting scene can be determined using a machine learning model. For example, when the shooting scene is a scene where the sound is loud, such as "sports", the volume range setting unit 63 sets the volume range VR on the high volume side. On the other hand, when the shooting scene is a scene where the sound is low, such as "night view", the volume range setting unit 63 sets the volume range VR on the low volume side. In addition, the feature extraction unit 66 can also use the shooting scene set by the user using the operation unit 16 as a feature of the image data PD.
[0098] [Second Modification Example]
[0099] In the above-described embodiment, the volume range setting unit 63 sets the volume range VR according to the features of the image data PD extracted by the feature extraction unit 66, but the volume range VR can also be set according to the shooting conditions of the subject. In this case, for example, the volume range setting unit 63 sets the volume range VR according to shooting conditions such as the exposure value set by the main control unit 60. For example, when the exposure value is small, the volume range setting unit 63 sets the volume range VR on the high volume side. On the other hand, when the exposure value is large, the volume range setting unit 63 sets the volume range VR on the low volume side.
[0100] Also, when the imaging device 10 has an optical zoom function or an electronic zoom function, the volume range setting unit 63 may also set the volume range VR according to the zoom ratio as an imaging condition. For example, when the zoom ratio is small (i.e., wide-angle), it is considered that the main subject is nearby and the sound emitted by the main subject is at a high volume. Therefore, the volume range setting unit 63 sets the volume range VR on the high-volume side. On the other hand, when the zoom ratio is large (i.e., telephoto), it is considered that the main subject is far away and the sound emitted by the main subject is at a low volume. Therefore, the volume range setting unit 63 sets the volume range VR on the low-volume side.
[0101] [Third Modification Example]
[0102] In the above-described embodiment, the image file creation unit 65 created the image file 28 including the image data PD and the third sound data AS3. However, instead of the third sound data AS3, an image file 28 including the second sound data AS2 or the second sound data AS2F and the image data PD may be created. Hereinafter, the image file 28 including the image data PD and the third sound data AS3 is referred to as the first image file 28A, and the image file 28 including the second sound data AS2 or the second sound data AS2F and the image data PD is referred to as the second image file 28B. In this modification example, the image file 28 including the second sound data AS2F and the image data PD is used as the second image file 28B.
[0103] In this modification example, the imaging device 10 has a first mode for creating the first image file 28A and a second mode for creating the second image file 28B. Since the dynamic range of the second sound data AS2F is large, when the volume width of the sound signal AS output from the external microphone 13 exceeds the dynamic range of the third sound data AS3, it is preferable to execute the second mode instead of the first mode. In this modification example, the second mode is executed or the execution of the second mode is recommended according to the volume information of the sound signal AS.
[0104] Figure 9 is a flowchart showing an example of the mode setting process according to the third modification example. In this modification example, the main control unit 60 executes the mode setting process at the imaging preparation stage before the imaging operation.
[0105] In the mode setting process, first, the main control unit 60 sets the creation mode to the first mode (step S21). Next, the main control unit 60 acquires the analog sound signal AS from the external microphone 13 (step S22). The main control unit 60 measures the volume width as an example of the volume information according to the acquired sound signal AS (step S23).
[0106] The main control unit 60 determines whether the measured volume width is equal to or greater than a specified value (step S24). When the measured volume width is equal to or greater than the specified value (step S24: Yes), the main control unit 60 displays a message recommending the execution of the second mode on the display 15 (step S25). If the user wishes to change the mode according to the message displayed on the display 15, the user can use the operation unit 16 to change the creation mode to the second mode.
[0107] The main control unit 60 determines whether there is a mode change operation performed by the user (step S26). When there is a mode change operation (step S26: Yes), the main control unit 60 changes the creation mode to the second mode (step S27).
[0108] When the volume width is less than the specified value (step S24: No) or when there is no mode change operation (step S26: No), the main control unit 60 ends the process without changing the creation mode from the first mode. Thus, the mode setting process ends.
[0109] When the creation mode is the first mode, the imaging operation of the imaging device 10 is the same as in the above-described embodiment, and in the Figure 8 third creation process (step S17) shown, the first image file 28A is created. On the other hand, when the creation mode is the second mode, in the third creation process, the second image file 28B is created. In addition, when the creation mode is the second mode, since it is not necessary to create the third audio data AS3, the setting process (step S14) and the second creation process (step S15) shown can be omitted. Figure 8
[0110] In addition, in this modification, when the volume width is equal to or greater than the specified value, the process of recommending the execution of the second mode is executed, and when there is a mode change operation, the creation mode is changed to the second mode. Instead, when the volume width is equal to or greater than the specified value, the process of recommending the execution of the second mode may not be executed and the creation mode may be changed to the second mode (i.e., the second mode is executed).
[0111] [Fourth Modification Example]
[0112] In the above-described embodiment, the amplifier 42 performs gain processing on the sound signal output from the sound collection element 41 in the external microphone 13, but the gain amount of the gain processing based on the amplifier 42 can be further changed according to the volume information of the sound signal AS output from the external microphone 13.
[0113] Figure 10 This is a flowchart showing an example of the gain adjustment process according to the fourth modification. In this modification, the main control unit 60 executes the gain adjustment process during the camera preparation stage before the imaging operation.
[0114] In the gain adjustment process, first, the main control unit 60 acquires the analog sound signal AS from the external microphone 13 (step S31). Next, the main control unit 60 measures the volume width, which is an example of volume information, based on the acquired sound signal AS (step S32). Then, the main control unit 60 changes the gain amount of the amplifier 42 via the microphone control unit 43 according to the measured volume width. Thus, the volume of the sound signal AS input to the sound signal processing circuit 23 can be set to an appropriate width in advance.
[0115] In addition, the main control unit 60 may also change the gain amount of the amplifier 42 according to the volume information determined by the characteristics of the subject or the priority information of the recorded sound source. For example, when the main control unit 60 determines that the volume emitted by the subject is small based on the characteristics of the subject, it increases the gain amount. And, for example, the main control unit 60 determines the subject to be prioritized based on the focus position of the imaging lens 12, etc., and estimates the sound of the subject with a high priority, thereby setting the optimal gain amount.
[0116] [Fifth Modification]
[0117] In the above-described embodiment, the sound signal AS input from the external microphone 13 to the sound signal processing circuit 23 is an analog signal, which is converted into a digital signal inside the sound signal processing circuit 23. Instead, the sound signal AS may be converted into a digital signal inside the external microphone 13 and then the digital sound signal AS may be transmitted to the sound signal processing circuit 23 via the connection part 11B.
[0118] Figure 11 This shows the structure of the imaging device 10 according to the fifth modification. In this modification, an ADC 44 is provided inside the external microphone 13. The ADC 44 generates a digital sound signal AS by performing A / D conversion on the sound signal that has been gain-processed by the amplifier 42. The microphone control unit 43 transmits the digital sound signal AS to the sound signal processing circuit 23 via the connection part 11B. The sound signal AS is, for example, a digital signal in the form of 24-bit LPCM. In this modification, it is not necessary to perform A / D conversion on the sound signal AS inside the sound signal processing circuit 23.
[0119] <<Other Modifications>>
[0120] In the above-described embodiment, the third sound data AS3 is monaural sound data without directivity information, but the third sound data AS3 may be set to stereophonic sound data with directivity information. For example, the directivity information may be generated using a sound signal obtained by the built-in microphone 22 as a stereo microphone.
[0121] In addition, the technology of the present invention is not limited to digital cameras, and can also be applied to electronic devices such as smartphones and tablet terminals having a photographing function.
[0122] In the above-described embodiment, as the hardware structure of the control unit exemplified by the processor 25, various processors shown below can be used. The above various processors include a CPU which is a common processor that functions as software (program), and in addition, processors such as an FPGA which can change the circuit structure after manufacturing. The FPGA includes a dedicated circuit and the like, and the dedicated circuit is a processor such as a PLD or an ASIC which has a circuit structure designed specifically for executing a specific process.
[0123] The control unit may be constituted by one of these various processors, or may be constituted by a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, a plurality of control units may be constituted by one processor.
[0124] There are multiple examples of constituting a plurality of control units by one processor. The first example is as follows: represented by computers such as a client and a server, one processor is constituted by a combination of one or more CPUs and software, and this processor functions as a plurality of control units. The second example is as follows: represented by a System On Chip (SOC) or the like, a processor that realizes the functions of an entire system including a plurality of control units using one IC chip is used. Thus, as the hardware structure, the control unit can be constituted by using one or more of the above various processors.
[0125] Furthermore, as the hardware structure of these various processors, more specifically, a circuit formed by combining circuit elements such as semiconductor elements can be used.
[0126] As long as there is no contradiction, the above-described embodiment and each modification example can be appropriately combined.
[0127] The description and illustration shown above are detailed descriptions of parts related to the technology of the present invention, and are only examples of the technology of the present invention. For example, the descriptions related to the above structures, functions, actions, and effects are descriptions related to examples of the structures, functions, actions, and effects of parts related to the technology of the present invention. Therefore, of course, it is also possible to delete unnecessary parts, add new elements, or make replacements to the description and illustration shown above without departing from the gist of the technology of the present invention. In addition, in order to avoid trouble and facilitate understanding of the parts related to the technology of the present invention, descriptions related to common technical knowledge that do not require special explanation when implementing the technology of the present invention are omitted in the description and illustration shown above.
[0128] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each document, patent application, and technical standard were specifically and individually incorporated by reference.
[0129] From the above description, the following technology can be grasped.
[0130] [Supplementary Note Item 1]
[0131] A photographing method, comprising:
[0132] An acquisition step of acquiring first sound data of a first digit, the first sound data of the first digit being generated based on a sound signal output from a sound collection element and corresponding to image data obtained by photographing a subject; and
[0133] A first creation step of creating second sound data of a second digit larger than the first digit based on the first sound data.
[0134] [Supplementary Note Item 2]
[0135] According to the photographing method described in Supplementary Note Item 1, wherein
[0136] The image data is moving image data.
[0137] [Supplementary Note Item 3]
[0138] According to the photographing method described in Supplementary Note Item 1 or 2, it further comprises:
[0139] A second creation step of creating third sound data of a third digit smaller than the second digit based on the second sound data; and
[0140] A third creation step of creating a first image file including the image data and the third sound data.
[0141] [Supplementary Note Item 4]
[0142] The imaging method according to appended item 3, wherein,
[0143] In the second creation process, the third sound data is created by extracting data of a volume range having the width of the third digit according to the second sound data.
[0144] [Appended item 5]
[0145] The imaging method according to appended item 4, further comprising a setting process,
[0146] In the setting process, before executing the second creation process, the volume range is set for the second sound data at a prescribed time.
[0147] [Appended item 6]
[0148] The imaging method according to appended item 5, wherein,
[0149] In the setting process, the volume range is set according to the characteristics of the image data or the imaging conditions.
[0150] [Appended item 7]
[0151] The imaging method according to appended item 5, wherein,
[0152] In the setting process, the volume range is set according to the characteristics of the main subject reflected in the frame constituting the image data.
[0153] [Appended item 8]
[0154] The imaging method according to appended item 7, wherein,
[0155] In the setting process, the main subject is selected from a plurality of subjects reflected in the frame according to the size or category of the subject in the frame, the position of the subject within the viewing angle, the focus position of the imaging lens for photographing the subject, the input information of the user, or the line-of-sight information of the user.
[0156] [Appended item 9]
[0157] The imaging method according to any one of appended items 1 to 8, wherein,
[0158] In the first creation process, the second sound data is created by synthesizing the first sound data generated by performing a first gain process on the sound signal and the first sound data generated by performing a second gain process having a gain amount different from the first gain process on the sound signal.
[0159] [Appended item 10]
[0160] The imaging method according to any one of Supplementary Notes 1 to 9, wherein,
[0161] The second sound data is in floating-point form.
[0162] [Supplementary Note 11]
[0163] The imaging method according to any one of Supplementary Notes 3 to 8 has a first mode for creating the first image file and a second mode for creating a second image file including the image data and the second sound data,
[0164] The imaging method further includes a step of recommending the execution of the second mode or a step of executing the second mode based on the volume information of the sound signal output from the sound collecting element or the volume information determined based on the characteristics of the image data.
[0165] [Supplementary Note 12]
[0166] The imaging method according to any one of Supplementary Notes 1 to 11 further includes a gain adjustment step of performing gain processing on the sound signal,
[0167] In the gain adjustment step, the gain amount of the gain processing is changed based on the volume information of the sound collected by the sound collecting element, the volume information determined based on the characteristics of the subject, or the priority information of the recorded sound source.
Claims
1. A video recording method, comprising: An acquisition step of acquiring first audio data of a first number of digits, the first audio data of the first number of digits being generated based on a sound signal output from a sound collection element and corresponding to video data obtained by photographing a subject; And A first creation step of creating second audio data of a second number of digits greater than the first number of digits based on the first audio data.
2. The video recording method according to claim 1, wherein The video data is moving image data.
3. The video recording method according to claim 1, further comprising: A second creation step of creating third audio data of a third number of digits smaller than the second number of digits based on the second audio data; And A third creation step of creating a first video file including the video data and the third audio data.
4. The video recording method according to claim 3, wherein In the second creation step, the third audio data is created by extracting data within a volume range having a width of the third number of digits based on the second audio data.
5. The video recording method according to claim 4, further comprising a setting step The setting step is to set the volume range for the second audio data at a predetermined time before executing the second creation step.
6. The video recording method according to claim 5, wherein In the setting step, the volume range is set according to the characteristics of the video data or the shooting conditions.
7. The video recording method according to claim 5, wherein In the setting step, the volume range is set according to the characteristics of the main subject reflected in the frames constituting the video data.
8. The video recording method according to claim 7, wherein In the setting step, the main subject is selected from a plurality of subjects reflected in the frame according to the size or category of the subject in the frame, the position of the subject within the viewing angle, the focus position of the imaging lens for photographing the subject, the input information of the user, or the line-of-sight information of the user.
9. The video recording method according to claim 1, wherein In the first creation step, the second audio data is created by synthesizing the first audio data generated by performing a first gain process on the sound signal and the first audio data generated by performing a second gain process with a gain amount different from the first gain process on the sound signal.
10. The video recording method according to claim 1, wherein The second audio data is in floating-point form.
11. The video recording method according to claim 3 has a first mode of creating the first video file and a second mode of creating a second video file including the video data and the second audio data, The video recording method further includes a step of recommending the execution of the second mode or a step of executing the second mode according to the volume information of the sound signal output from the sound collection element or the volume information determined according to the characteristics of the video data.
12. The video recording method according to claim 1, further comprising a gain adjustment step of performing a gain process on the sound signal, In the gain adjustment process, the gain amount of the gain processing is changed based on the volume information of the sound collected by the sound collection element, the volume information determined according to the characteristics of the subject, or the priority information of the sound source being recorded.
13. An imaging device, comprising: an imaging element that captures a subject to generate image data; and a processor that acquires first sound data of a first number of digits and creates second sound data of a second number of digits larger than the first number of digits based on the first sound data, the first sound data of the first number of digits being generated based on a sound signal output from a sound collection element and corresponding to the image data.
14. The imaging device according to claim 13, having: a housing that houses the imaging element and the processor; and a connection portion for connecting an external microphone including the sound collection element to the housing.
15. The imaging device according to claim 14, wherein the sound signal is converted into a digital signal inside the external microphone and then sent to the inside of the housing via the connection portion.
Citation Information
Patent Citations
Data processing device, data processing method and digital audio mixer
JP2002246913A
Voice signal converter
JP2012073435A