Creation method and creation device

By acquiring sound data in floating point form and creating accompanying information, and editing and processing, the problem of low sound data quality in the prior art is solved, and higher quality sound data processing is achieved.

CN120226340APending Publication Date: 2025-06-27FUJIFILM CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380079758.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-22
Filing Date
2023-10-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to improve the quality of sound data, especially during sampling and processing, which can easily lead to sound distortion and deterioration.

Method used

By acquiring the first sound data in the floating point form and creating accompanying information, including information related to the sound collecting device or the sound source, the editing process is performed to generate the second sound data with a lower number of digits.

Benefits of technology

Improve the quality of sound data, reduce sound distortion and degradation, and enhance the setting and processing capabilities of the volume range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226340A_ABST
    Figure CN120226340A_ABST
Patent Text Reader

Abstract

This creation method comprises: an acquisition step for acquiring first sound data in a floating point form on the basis of sound emitted from a sound source and collected by a sound collection device; and an attached information creation step for creating attached information attached to the first sound data and including device information, which is information relating to the sound collection device, or sound source information, which is information relating to the sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present invention relates to a creation method and a creation device. Background Art

[0002] Japanese Patent Laid-Open No. 2012-073435 discloses a sound signal conversion device that samples input analog sound signals of L channel and R channel at a sampling frequency of 192 kHz and a quantization bit number of 24 Bit using an A / D conversion device to generate digital signals. A signal processing device is connected to the output side of the A / D conversion device. The signal processing device performs processing of downsampling the frequency to 1 / 4 (48 kHz) and processing of converting the downsampled signal into a floating-point format with a quantization bit number of 32 Bit.

[0003] Japanese Patent Laid-Open No. 2002-246913 discloses a data processing device that converts input data from a fixed-point form to a floating-point form by a conversion unit. Summary of the Invention

[0004] Technical Problem to be Solved by the Invention

[0005] An object of an embodiment of the technology of the present invention is to provide a creation method and a creation device capable of improving the quality of sound data.

[0006] Means for Solving the Technical Problem

[0007] To achieve the above object, the creation method of the present invention includes: an acquisition step of acquiring first sound data in a floating-point form based on sound emitted from a sound source and collected by a sound collection device; and an additional information creation step of creating additional information attached to the first sound data and including device information as information related to the sound collection device or sound source information as information related to the sound source.

[0008] Preferably, the first sound data is used to create second sound data having a smaller bit number than the bit number of the first sound data.

[0009] Preferably, the additional information includes device information. In this case, preferably, the device information is information related to gain processing used for the sound collected by the sound collection device or information related to the performance of the sound collection device.

[0010] Preferably, the additional information includes sound source information. In this case, preferably, the sound source information has a corresponding relationship with the time information of the first sound data.

[0011] Preferably, it further includes a photographing step in which image data corresponding to the first sound data is created by a photographing device.

[0012] Preferably, the sound source is a subject included in the image data.

[0013] Preferably, the sound source is the main subject selected from among a plurality of subjects included in the image data.

[0014] Preferably, the sound source information is the type of driving sound accompanying the driving of the imaging device.

[0015] Preferably, it further includes a first file creation step in which a first file containing first sound data and attached information is created.

[0016] Preferably, it further includes an editing step in which second sound data having a smaller number of bits than the first sound data is created by editing the first sound data according to the attached information.

[0017] In the editing step, degradation information that is information related to the sound of the second sound data degraded due to editing may be created, and a second file containing the second sound data and the degradation information may be created. Preferably, the second file includes attached information.

[0018] The creation device of the present invention includes a processor that executes: an acquisition step of acquiring first sound data in floating point form based on the sound emitted from the sound source and collected by the sound collection device; and an attached information creation step of creating attached information attached to the first sound data and including device information that is information related to the sound collection device or sound source information that is information related to the sound source. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a diagram showing an example of the structure of the imaging device according to the first embodiment.

[0020] Figure 2 It is a diagram showing an example of the structure of the sound signal processing circuit.

[0021] Figure 3 It is a diagram conceptually showing the sound signal processing.

[0022] Figure 4 It is a diagram showing an example of the functional structure of the processor.

[0023] Figure 5 It is a diagram conceptually showing the synthesis process and the data format conversion process.

[0024] Figure 6 It is a diagram conceptually showing the editing process.

[0025] Figure 7 It is a flowchart showing an example of the operation of the imaging device.

[0026] Figure 8This is a diagram showing an example of the functional structure of the processor according to the second embodiment.

[0027] Figure 9 This is a diagram showing an example of the relationship between sound source information and the first sound data.

[0028] Figure 10 This is a flowchart showing an example of the operation of the imaging device according to the second embodiment.

[0029] Figure 11 This is a diagram showing an example of the functional structure of the processor according to the third embodiment.

[0030] Figure 12 This is a diagram showing an example of the relationship between the sound source information and the first sound data according to the third embodiment.

[0031] Figure 13 This is a diagram showing an example of the sound data file created by the editorial department. Detailed Embodiment

[0032] An example of an embodiment related to the technology of the present invention will be described with reference to the accompanying drawings.

[0033] First, the terms used in the following description will be explained.

[0034] In the following description, "AF" is an abbreviation for "Auto Focus", "MF" is an abbreviation for "Manual Focus", "IC" is an abbreviation for "Integrated Circuit", "CPU" is an abbreviation for "Central Processing Unit", "RAM" is an abbreviation for "Random Access Memory", and "CMOS" is an abbreviation for "Complementary Metal Oxide Semiconductor".

[0035] "FPGA" is an abbreviation for "Field Programmable Gate Array", "PLD" is an abbreviation for "Programmable Logic Device", "ASIC" is an abbreviation for "Application Specific Integrated Circuit", "OVF" is an abbreviation for "Optical View Finder", "EVF" is an abbreviation for "Electronic View Finder", and "ADC" is an abbreviation for "Analog to Digital Converter", and "LPCM" is an abbreviation for "Linear Pulse Code Modulation".

[0036] As an embodiment of the imaging device, a lens interchangeable digital camera will be exemplified to describe the technology of the present invention. In addition, the technology of the present invention is not limited to lens interchangeable type, and can also be applied to a lens integrated type digital camera.

[0037] [First Embodiment]

[0038] Figure 1 An example of the structure of the imaging device 10 according to the first embodiment is shown. The imaging device 10 is a lens interchangeable digital camera. The imaging device 10 includes a housing 11 and an imaging lens 12 that is detachably mounted on the housing 11 and includes a focusing lens 31. The imaging lens 12 is mounted on the front surface side of the housing 11 via a bayonet 11A. In addition, the imaging device 10 is an example of a "creation device" according to the technology of the present invention.

[0039] In addition, an external microphone 13 can be detachably mounted on the housing 11. The external microphone 13 is mounted on the housing 11 via a connection portion 11B provided on the upper surface of the housing 11. The external microphone 13 is a shotgun microphone, a zoom microphone, etc. The connection portion 11B is, for example, a hot shoe. In addition, the external microphone 13 is an example of a "sound collecting device" according to the technology of the present invention.

[0040] An operation portion 16 including a dial, a release button, etc. is provided on the housing 11. As operation modes of the imaging device 10, for example, a still image shooting mode, a moving image shooting mode, and an image display mode are included. The operation portion 16 is operated by a user when setting an operation mode. In addition, the operation portion 16 is operated by a user when starting to execute still image shooting or moving image shooting.

[0041] Further, the operation unit 16 is operated by the user when selecting a focusing mode. The focusing modes include an AF mode and an MF mode. The AF mode is a mode in which focusing control is performed by setting a subject area selected by the user or a subject area automatically detected by the imaging device 10 as a focus detection area (hereinafter referred to as an AF area). The MF mode is a mode in which the user manually performs focusing control by operating a focusing ring (not shown).

[0042] Further, a viewfinder 14 is provided on the housing 11. For example, the viewfinder 14 is a hybrid viewfinder (registered trademark). The hybrid viewfinder is, for example, a viewfinder that selectively uses an optical viewfinder (hereinafter referred to as "OVF") and an electronic viewfinder (hereinafter referred to as "EVF"). The user can observe an optical image or an instant preview image of the subject projected through the viewfinder 14 via a viewfinder eyepiece portion (not shown).

[0043] Further, a display 15 is provided on the back side of the housing 11. An image based on the image data PD obtained by shooting and various menu screens are displayed on the display 15. The user can also observe the instant preview image projected through the display 15 instead of observing the viewfinder 14.

[0044] Further, a speaker 17 is provided on the housing 11. The speaker 17 outputs sound according to second sound data AS2 included in a moving image file 28 described later.

[0045] The housing 11 and the imaging lens 12 are electrically connected via electrical contacts 11C provided on the bayonet 11A.

[0046] The imaging lens 12 includes a focusing lens 31, an aperture 32, and a lens drive control unit 33. The lens drive control unit 33 is electrically connected to a processor 25 accommodated in the housing 11 via the electrical contacts 11C.

[0047] The lens drive control unit 33 drives the focusing lens 31 and the aperture 32 according to a control signal sent from the processor 25. In order to adjust the position of the focusing lens 31, the lens drive control unit 33 performs drive control of the focusing lens 31 according to a focusing control signal sent from the processor 25.

[0048] The aperture 32 has an opening with a variable opening diameter. In order to adjust the amount of incident light incident on the imaging sensor 20, the lens drive control unit 33 performs drive control of the aperture 32 according to an aperture adjustment signal sent from the processor 25.

[0049] In addition, an imaging sensor 20, an image processing circuit 21, a built-in microphone 22, a sound signal processing circuit 23, a processor 25, and a storage device 26 are provided inside the housing 11. The imaging sensor 20, the image processing circuit 21, the built-in microphone 22, the sound signal processing circuit 23, the storage device 26, the display 15, and the speaker 17 are controlled by the processor 25 to operate.

[0050] The processor 25 is constituted by, for example, a CPU. A RAM 25A serving as a primary storage memory is connected to the processor 25. The storage device 26 is constituted by, for example, a non-volatile memory such as a flash memory. The processor 25 executes various processes according to a program 27 stored in the storage device 26. In addition, the processor 25 may be constituted by an aggregate of a plurality of IC chips.

[0051] The imaging sensor 20 is, for example, a CMOS image sensor. Light (subject image) after passing through the imaging lens 12 is incident on a light receiving surface 20A of the imaging sensor 20. A plurality of pixels that generate a captured image signal by performing photoelectric conversion are formed on the light receiving surface 20A. The imaging sensor 20 generates and outputs image data PD by performing photoelectric conversion on the light incident on each pixel.

[0052] The image processing circuit 21 performs image processing including white balance correction, gamma correction processing, etc. on the image data PD output from the imaging sensor 20.

[0053] The built-in microphone 22 is a stereo microphone having a pair of sound collecting elements 22A, 22B. The sound collecting elements 22A, 22B are a sound sensor for the left channel (hereinafter referred to as the L channel) and a sound sensor for the right channel (hereinafter referred to as the R channel). The sound collecting elements 22A, 22B are sound sensors such as electrostatic type, piezoelectric type, and dynamic type, and output the collected sound as sound signals AL, AR. The sound signal processing circuit 23 performs sound signal processing including gain processing, A / D conversion processing, etc. on the sound signals AL, AR output from the sound collecting elements 22A, 22B.

[0054] The external microphone 13 includes a sound collecting element 41, an amplifier 42, and a microphone control unit 43. In the present embodiment, the external microphone 13 is a monaural microphone having one sound collecting element 41. The sound collecting element 41 is a sound sensor such as electrostatic type, piezoelectric type, and dynamic type, and outputs the collected sound as a sound signal. The amplifier 42 performs gain processing on the sound signal output from the sound collecting element 41. The microphone control unit 43 controls the gain amount of the gain processing based on the amplifier 42.

[0055] Further, the microphone control unit 43 supplies the sound signal that has been gain-processed by the amplifier 42 to the sound signal processing circuit 23 within the housing 11 via the connection unit 11B. A monaural and analog sound signal AS is supplied from the external microphone 13 to the sound signal processing circuit 23. Additionally, the operation of the microphone control unit 43 is controlled by the processor 25.

[0056] Moreover, device information 13A, which is information related to the external microphone 13, is stored in the storage device 26. In the present embodiment, it is information related to the gain processing used for the sound collected by the device information 13A. Specifically, the device information 13A is the gain amount based on the gain processing of the amplifier 42 (i.e., the sensitivity of the external microphone 13).

[0057] Figure 2 An example of the structure of the sound signal processing circuit 23 is shown. The sound signal processing circuit 23 includes a first preamplifier 51A, a first ADC 52A, a second preamplifier 51B, and a second ADC 52B.

[0058] The first preamplifier 51A and the first ADC 52A are processing units for the L channel that perform gain processing and A / D conversion processing on the sound signal AL output from the sound collection element 22A included in the built-in microphone 22. The second preamplifier 51B and the second ADC 52B are processing units for the R channel that perform gain processing and A / D conversion processing on the sound signal AR output from the sound collection element 22B included in the built-in microphone 22.

[0059] The gain amount G1 of the first preamplifier 51A is controlled by the processor 25. The gain amount G2 of the second preamplifier 51B is controlled by the processor 25. When performing gain processing on the sound signals AL and AR output from the built-in microphone 22, the gain amount G1 and the gain amount G2 are set to the same value by the processor 25. The first ADC 52A and the second ADC 52B convert the analog sound signal into a digital signal in the 24-bit LPCM format, for example, by sampling with a quantization bit number of 24 bits. Additionally, the LPCM format is an example of the "pulse code modulation format" related to the technology of the present invention.

[0060] The sound signal AS output from the external microphone 13 is input to the first preamplifier 51A and the second preamplifier 51B. The first preamplifier 51A performs a gain process on the sound signal AS with a gain amount G1. The second preamplifier 51B performs a gain process on the sound signal AS with a gain amount G2. When performing a gain process on the sound signal AS output from the external microphone 13, the gain amount G1 and the gain amount G2 are set to different values by the processor 25. Hereinafter, the gain process performed by the first preamplifier 51A is referred to as the first gain process, and the gain process performed by the second preamplifier 51B is referred to as the second gain process.

[0061] The first ADC 52A converts the sound signal AS that has undergone the first gain process by the first preamplifier 51A into a digital signal. The second ADC 52B converts the sound signal AS that has undergone the second gain process by the second preamplifier 51B into a digital signal. Hereinafter, the sound signal AS digitized by the first ADC 52A is referred to as the modulated sound data ASH, and the sound signal AS digitized by the second ADC 52B is referred to as the modulated sound data ASL. The modulated sound data ASH and ASL are output from the sound signal processing circuit 23 to the processor 25.

[0062] Figure 3 Conceptually shows the sound signal processing of the sound signal AS based on the sound signal processing circuit 23. The sound signal AS output from the external microphone 13 is input to the L-channel processing unit and the R-channel processing unit. The sound signal AS input to the L-channel processing unit is converted into a digital signal after undergoing the first gain process with the gain amount G1, and thus is output from the sound signal processing circuit 23 as the modulated sound data ASH. The sound signal AS input to the R-channel processing unit is converted into a digital signal after undergoing the second gain process with the gain amount G2, and thus is output from the sound signal processing circuit 23 as the modulated sound data ASL. In the present embodiment, the number of bits of the modulated sound data ASH and ASL is 24 bits.

[0063] For example, the gain amount G1 is set to +48 dB, and the gain amount G2 is set to -48 dB. Since 48 dB corresponds to an 8-bit volume width, as Figure 3 shown, there is a 16-bit deviation between the high-gain modulated sound data ASH and the low-gain modulated sound data ASL. In other words, there is an 8-bit overlap between the modulated sound data ASH and the modulated sound data ASL.

[0064] Figure 4 Shows an example of the functional structure of the processor 25. The processor 25 realizes various functional units by executing processes according to the program 27 stored in the storage device 26. Figure 4 The various functional units shown are realized in the dynamic image shooting mode. As Figure 4As shown, for example, the processor 25 implements the main control unit 60, the synthesis processing unit 61, the data format conversion unit 62, the additional information creation unit 63, the audio data file creation unit 64, the editing unit 65, and the file creation unit 66. The editing unit 65 includes a volume range setting unit 65A and a data extraction unit 65B.

[0065] The main control unit 60 uniformly controls each part of the imaging device 10. The main control unit 60 controls the operation of the imaging device 10 according to the instruction signal input from the operation unit 16. The main control unit 60 controls the imaging sensor 20 to perform an imaging operation. The imaging sensor 20 outputs the image data PD generated by shooting through the imaging lens 12. In the moving image shooting mode, the imaging sensor 20 outputs the image data PD in one frame period. The image data PD output from the imaging sensor 20 is input to the processor 25 after being subjected to image processing by the image processing circuit 21. In the case of the moving image shooting mode, the image data PD is data composed of a plurality of frames.

[0066] Moreover, in the moving image shooting mode, when the external microphone 13 is connected to the connection unit 11B, the main control unit 60 controls the external microphone 13 to perform a sound collection operation. The external microphone 13 outputs the sound signal AS to the sound signal processing circuit 23 via the connection unit 11B during the imaging operation of the imaging sensor 20. The sound signal processing circuit 23 outputs the modulated sound data ASH, ASL by performing the above-mentioned sound signal processing. The modulated sound data ASH, ASL are the sound data corresponding to the image data PD obtained by shooting the subject with the imaging sensor 20.

[0067] The synthesis processing unit 61 acquires the modulated sound data ASH, ASL output from the sound signal processing circuit 23, and synthesizes the modulated sound data ASH, ASL to create the first sound data AS1 of the first number of bits. The first sound data AS1 is digital data in the LPCM format.

[0068] The data format conversion unit 62 converts the data format of the first sound data AS1 into the floating point format. Hereinafter, the first sound data AS1 converted into the floating point format is referred to as the first sound data AS1F. The first sound data AS1F is used to create the second sound data AS2 having a smaller number of bits than the first sound data AS1F.

[0069] The additional information creation unit 63 reads out the device information 13A from the storage device 26 and creates the additional information SI including the device information 13A. The additional information creation unit 63 supplies the created additional information SI to the audio data file creation unit 64. The additional information SI is so-called meta information.

[0070] The sound data file creation unit 64 creates a sound data file 67 that includes the first sound data AS1F created by the data format conversion unit 62 and the attached information SI created by the attached information creation unit 63. The sound data file creation unit 64 records the created sound data file 67 in the storage device 26. Additionally, the sound data file 67 corresponds to the "first file" related to the technology of the present invention.

[0071] The editing unit 65 creates a second sound data AS2 having a second number of bits smaller than the first number of bits by editing the first sound data AS1F included in the sound data file 67 recorded in the storage device 26 according to the attached information SI. For example, the second number of bits is 24 bits.

[0072] Specifically, the volume range setting unit 65A sets a volume range VR having a width of the second number of bits for the dynamic range of the first sound data AS1F. In the present embodiment, the volume range setting unit 65A sets the volume range VR according to the attached information SI. The data extraction unit 65B extracts the data of the volume range VR set by the volume range setting unit 65A from the first sound data AS1F, thereby creating the second sound data AS2.

[0073] The file creation unit 66 creates a moving image file 28 that includes the video data PD output from the image processing circuit 21 and the second sound data AS2 output from the data extraction unit 65B, and stores it in the storage device 26.

[0074] Figure 5 Conceptually shows the synthesis process based on the synthesis processing unit 61 and the data format conversion process based on the data format conversion unit 62. The synthesis processing unit 61 synthesizes the modulated sound data ASH and the modulated sound data ASL by performing a mixing process on the overlapping part of the 8-bit amounts of the modulated sound data ASH and the modulated sound data ASL. The number of bits of the first sound data AS1 generated by this synthesis process (i.e., the first number of bits) is 40 bits. Thus, by synthesizing the modulated sound data ASH and the modulated sound data ASL with different synthesis gain amounts, the first sound data AS1 with an expanded dynamic range of volume is obtained.

[0075] The data format conversion unit 62 converts the first sound data AS1 in 40-bit fixed-point format into the first sound data AS1F in 32-bit floating-point format (so-called 32-bit floating-point number). The 32-bit floating-point number consists of 1-bit sign, 8-bit exponent part, and 23-bit mantissa part. The conversion from fixed-point format to floating-point format can use a known method. With the floating-point format, a wide range of numerical expressions can be performed.

[0076] Figure 6Conceptually shown is the editing process based on the editorial department 65. The volume range setting unit 65A sets the volume range VR according to the device information 13A included in the attached information SI. Specifically, the volume range setting unit 65A sets the volume range VR according to the gain amount of the gain process based on the amplifier 42 (i.e., the sensitivity of the external microphone 13), which is an example of the device information 13A. For example, the larger the gain amount (i.e., the higher the sensitivity), the more the volume range setting unit 65A sets the volume range VR on the low volume side, and the smaller the gain amount (i.e., the lower the sensitivity), the more the volume range setting unit 65A sets the volume range VR on the high volume side. Thus, sound distortion and the like can be suppressed.

[0077] The data extraction unit 65B extracts the data of the volume range VR from the first sound data AS1F, thereby creating the second sound data AS2. In this way, by creating the second sound data AS2 according to the device information 13A, it is possible to generate the 24-bit fixed-point form second sound data AS2 in which sound distortion and the like are suppressed.

[0078] Figure 7 It is a flowchart showing an example of the operation of the imaging device 10. Figure 7 It shows the operation when the moving image shooting mode is selected as the operation mode and the external microphone 13 is connected to the connection part 11B.

[0079] First, the main control unit 60 determines whether there is an instruction to start moving image shooting by the user (step S10). When it is determined that there is a start instruction (step S10: Yes), the shooting process (step S11) and the acquisition process (step S12) are executed simultaneously. In the shooting process, the imaging sensor 20 shoots the subject to generate the image data PD. In the acquisition process, the first sound data AS1F in floating point form is acquired based on the sound collected by the external microphone 13.

[0080] After the shooting process and the acquisition process, the main control unit 60 determines whether there is an instruction to end moving image shooting by the user (step S13). When it is determined that there is no end instruction (step S13: No), the process returns to steps S11 and S12. Steps S11 and S12 are repeatedly executed until it is determined that there is an end instruction in step S13.

[0081] When it is determined that there is an end instruction (step S13: Yes), the attached information creation process (step S14) is executed. In the attached information creation process, the attached information creation unit 63 creates the attached information SI including the device information 13A.

[0082] After the supplementary information creation process, the sound data file creation process (step S15) is executed. In the sound data file creation process, the sound data file creation unit 64 creates a sound data file 67 that includes the first sound data AS1F and the supplementary information SI. In addition, the sound data file creation process corresponds to the "first file creation process" related to the technology of the present invention.

[0083] After the sound data file creation process, the editing process (step S16) is executed. In the editing process, the second sound data AS2 is created by editing the first sound data AS1F according to the supplementary information SI. Then, a moving image file 28 that includes the video data PD and the second sound data AS2 is created and recorded in the storage device 26. Thus, the operation of the imaging device 10 ends.

[0084] As described above, the creation method of the present invention includes: an acquisition process of acquiring the first sound data in floating point form based on the sound emitted from a sound source such as a subject and collected by the sound collection device; and a creation process of creating device information as supplementary information attached to the first sound data and as information related to the sound collection device. Thereby, the quality of the sound data can be improved.

[0085] In addition, in the above-described embodiment, the device information 13A is set as information related to the sound collection device, but the device information 13A may also be information related to the performance of the sound collection device. The information related to the performance of the sound collection device includes a noise level, a maximum sound pressure level, an output impedance, and the like. In this case, the volume range setting unit 65A sets the volume range VR according to the information related to the performance of the external microphone 13.

[0086] The noise level indicates the ease of noise mixing into the external microphone 13. When the performance of the sound collection device is the noise level, the volume range setting unit 65A restricts the upper limit of the volume range VR according to the noise level, for example. This is because, when the noise level is low, if the volume range VR is set on the high volume side, it will be difficult to hear the sound with the subject as the sound source due to the noise.

[0087] The maximum sound pressure level indicates the maximum sound pressure level that the external microphone 13 can collect. When the performance of the sound collection device is the maximum sound pressure level, the volume range setting unit 65A restricts the upper limit of the volume range VR according to the maximum sound pressure level, for example. This is because sounds exceeding the maximum sound pressure level may cause distortion.

[0088] The output impedance represents the magnitude of the internal resistance of the external microphone 13. When the performance of the sound collection device is the output impedance, for example, the smaller the output impedance, the more the volume range setting unit 65A sets the volume range VR on the low volume side. This is because the lower the output impedance, the smaller the voltage drop and the less the deterioration of the sound signal AS output from the external microphone 13. Also, the volume range setting unit 65A can limit the upper limit of the volume range VR according to the output impedance. This is because the higher the output impedance, the greater the voltage drop and the greater the possibility that the noise contained in the sound signal AS will be. For example, the higher the output impedance, the more the upper limit of the set volume range VR is reduced.

[0089] Also, the information related to the performance of the sound collection device can be a polar diagram. The polar diagram represents the sensitivity (i.e., directivity) of the external microphone 13 with respect to the sound collection direction. The polar diagram includes unidirectional, bidirectional, and omnidirectional. The volume range setting unit 65A sets the volume range VR according to the type of the polar diagram, for example. Also, by considering the image data PD in addition to the polar diagram, the volume range VR can be set within the optimal range in a manner that includes the sound emitted from the sound source such as the subject.

[0090] Also, the information related to the performance of the sound collection device can be a frequency range. The frequency range is the range of frequencies of the sound that the external microphone 13 can collect and reproduce.

[0091] [Second Embodiment]

[0092] Next, the second embodiment will be described. In the first embodiment, the accessory information creation unit 63 created the accessory information SI including the device information 13A. In the second embodiment, the accessory information creation unit 63 creates the accessory information SI including the sound source information instead of the device information 13A. The sound source information refers to the information related to the sound source such as the subject.

[0093] The structure of the imaging device 10 according to the second embodiment other than the processor 25 is the same as that of the first embodiment. Hereinafter, the same components as those in the first embodiment will be denoted by the same reference numerals, and the description will be appropriately omitted.

[0094] Figure 8This shows an example of the functional structure of the processor 25 according to the second embodiment. In this embodiment, only the function of the additional information creation unit 63 is different from that of the first embodiment. In this embodiment, the additional information creation unit 63 acquires sound source information based on the video data PD output from the image processing circuit 21, and creates additional information SI including the acquired sound source information. In this embodiment, the sound source is the main subject selected from among the multiple subjects included in the video data PD. For example, during the imaging operation of the imaging device 10, the additional information creation unit 63 acquires the category of the main subject as the sound source information for each frame constituting the video data PD.

[0095] The main subject refers to the subject among the multiple subjects included in the video data PD that is determined by the user or the main control unit 60 to have a high degree of importance. In this embodiment, the volume range setting unit 65A sets the volume range VR according to the category of the main subject. For example, when the category of the main subject is a category where the sound emitted is considered to be a high volume, such as "airplane", the volume range setting unit 65A sets the volume range VR on the high volume side. On the other hand, when the category of the main subject is a category where the sound emitted is considered to be a low volume, such as "person", the volume range setting unit 65A sets the volume range VR on the low volume side.

[0096] As the process of the additional information creation unit 63 selecting the main subject from among the multiple subjects included in the video data PD, various processes can be applied. As an example, the additional information creation unit 63 selects the main subject according to the sizes of the multiple subjects included in the video data PD. In this case, the additional information creation unit 63 selects the subject with the largest size among the multiple subjects as the main subject.

[0097] Moreover, the additional information creation unit 63 can also select the main subject according to the categories of the multiple subjects included in the video data PD. In this case, for example, the additional information creation unit 63 determines the category of each subject, and selects the subject that matches the category set by the user using the operation unit 16 as the main subject. For example, when the portrait shooting mode is set, the additional information creation unit 63 selects a person as the main subject from among the multiple subjects included in the video data PD.

[0098] Moreover, the additional information creation unit 63 can also select the main subject according to the positions of the multiple subjects included in the video data PD within the viewing angle. In this case, for example, the additional information creation unit 63 calculates the positions of each subject within the viewing angle, and selects the subject located at the center of the viewing angle as the main subject.

[0099] Further, the accessory information creation unit 63 may also select the main subject according to the focusing position of the imaging lens 12. In this case, for example, the accessory information creation unit 63 acquires information related to the focusing position from the main control unit 60, and selects, as the main subject, the subject closest to the focusing position from among the plurality of subjects included in the image data PD.

[0100] Further, the accessory information creation unit 63 may also select the main subject according to the user's input information. In this case, for example, the accessory information creation unit 63 selects, as the main subject, the subject located in the subject area selected by the user using the operation unit 16. Additionally, the accessory information creation unit 63 may also select, as the main subject, the subject located in the AF area automatically detected by the imaging device 10.

[0101] Further, the accessory information creation unit 63 may also select the main subject according to the user's line-of-sight information. In this case, for example, the imaging device 10 has a function of detecting the line of sight of the user looking into the viewfinder 14. The accessory information creation unit 63 acquires the user's line-of-sight information, and selects, as the main subject, the subject located at the position of the line of sight within the viewing angle.

[0102] Figure 9 An example of the relationship between the sound source information and the first sound data AS1F is shown. As Figure 9 shown, the first sound data AS1F is data representing the change in volume with respect to time (i.e., the change in amplitude). A correspondence relationship is established between the sound source information and the time information included in the first sound data AS1F. In Figure 9 the example shown, a correspondence relationship is established between the main subject A as the sound source information between t1 and t2, and a correspondence relationship is established between the main subject B as the sound source information between t2 and t3. The categories of the main subject A and the main subject B are different. For example, the main subject A is a "person" whose emitted sound is of low volume, and the main subject B is an "airplane" whose emitted sound is of high volume.

[0103] Figure 10 It is a flowchart showing an example of the operation of the imaging device 10 according to the second embodiment. Figure 10 It shows the operation when the moving image shooting mode is selected as the operation mode and the external microphone 13 is connected to the connection unit 11B.

[0104] In the operation of the imaging device 10 according to the present embodiment, the timing when only the additional information creation process (step S14) is executed is different from that of the above-described embodiment. In the present embodiment, since the sound source information changes over time, the additional information creation process is executed after the acquisition process. That is, in the present embodiment, steps S11, S12, and S14 are repeatedly executed until it is determined in step S13 that there is an end instruction. When it is determined that there is an end instruction (step S13: YES), the sound data file creation process (step S15) and the editing process (step S16) are executed.

[0105] [Third Embodiment]

[0106] Next, the third embodiment will be described. In the second embodiment, the sound source of the sound source information acquired by the additional information creation unit 63 is the subject. In the third embodiment, the sound source information acquired by the additional information creation unit 63 is the type of driving sound accompanying the driving of the imaging device 10. For example, the types of driving sounds are the driving sound of the focusing lens 31, the driving sound of the aperture 32, and the like. When the imaging device 10 includes a cooling fan, the types of driving sounds include the driving sound of the cooling fan.

[0107] The configuration of the imaging device 10 according to the third embodiment other than the processor 25 is the same as that of the first embodiment. Hereinafter, the same reference numerals are given to the same components as those in the first embodiment, and the description will be appropriately omitted.

[0108] Figure 11 An example of the functional configuration of the processor 25 according to the third embodiment is shown. In the present embodiment, only the function of the additional information creation unit 63 is different from that of the first embodiment. In the present embodiment, the additional information creation unit 63 acquires the sound source information based on the first sound data AS1F output from the data format conversion unit 62, and creates the additional information SI including the acquired sound source information. The additional information creation unit 63 acquires the type of driving sound accompanying the driving of the imaging device 10 as the sound source information during the imaging operation of the imaging device 10. For example, the additional information creation unit 63 determines the type of driving sound by analyzing the first sound data AS1F.

[0109] In addition, the additional information creation unit 63 is not limited to the first sound data AS1F, and may determine the type of driving sound based on sound data such as the first sound data AS1, the modulation sound data ASH, and ASL. Further, the additional information creation unit 63 may also determine the type of driving sound by acquiring information indicating the type of device being driven from the main control unit 60.

[0110] Figure 12An example of the relationship between the sound source information related to the third embodiment and the first sound data AS1F is shown. A correspondence relationship is established between the sound source information and the time information included in the first sound data AS1F. In Figure 9 In the example shown, a driving sound A is established as the sound source information corresponding to the time between t1 and t2, and a driving sound B is established as the sound source information corresponding to the time between t2 and t3. The types of the driving sound A and the driving sound are different. For example, the driving sound A is the "driving sound of the aperture 32" with a low volume sound, and the driving sound B is the "driving sound of the hot fan" with a high volume sound.

[0111] The operation of the imaging device 10 according to the third embodiment obtains the type of the driving sound as the sound source information in the accessory information creation process, and is the same as the second embodiment except for this.

[0112] [Modification Example]

[0113] Next, various modification examples will be described. In each of the above embodiments, the editorial department 65 created the second sound data AS2 by editing the first sound data AS1F according to the accessory information SI. The editorial department 65 may also create degradation information DI as information related to the sound of the second sound data AS2 degraded due to editing, and as Figure 13 shown, create a sound data file 68 including the second sound data AS2 and the degradation information DI. The editorial department 65 records the created sound data file 68 in the storage device 26. In addition, the sound data file 68 corresponds to the "second file" related to the technology of the present invention.

[0114] Degradation means that at least a part of the sound of the sound source included in the first sound data AS1F is not included in the volume range of the second sound data AS2. The degradation information is the missing sound information when creating the second sound data AS2 with less information amount than the first sound data AS1F from the first sound data AS1F. In addition, the sound data file 68 may also include the accessory information SI.

[0115] The technology of the present invention is not limited to digital cameras, and can also be applied to electronic devices such as smartphones and tablet terminals having an imaging function.

[0116] In each of the above embodiments, as the hardware structure of the control unit taking the processor 25 as an example, various processors shown below can be used. The above various processors include a CPU which is a common processor that functions as an execution software (program), and in addition, also include processors such as an FPGA that can change the circuit structure after manufacturing. The FPGA includes a dedicated circuit, etc., and the dedicated circuit is a processor such as a PLD or an ASIC having a circuit structure specifically designed for executing a specific process.

[0117] The control unit may be constituted by one of these various processors, or may be constituted by a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Further, multiple control units may also be constituted by one processor.

[0118] There are multiple examples of constituting multiple control units by one processor. The first example has the following method: Represented by a computer such as a client and a server, one processor is constituted by a combination of one or more CPUs and software, and this processor functions as multiple control units. The second example has the following method: Represented by a System On Chip (SOC) or the like, a processor that uses one IC chip to implement the functions of an entire system including multiple control units is used. Thus, as a hardware structure, the control unit may be constituted by using one or more of the above various processors.

[0119] Furthermore, as the hardware structure of these various processors, more specifically, a circuit formed by combining circuit elements such as semiconductor elements may be used.

[0120] The description content and illustrated content shown above are detailed descriptions of parts related to the technology of the present invention, and are only an example of the technology of the present invention. For example, the descriptions related to the above structure, function, action, and effect are descriptions related to an example of the structure, function, action, and effect of parts related to the technology of the present invention. Therefore, of course, it is also possible to delete unnecessary parts, add new elements, or make replacements to the description content and illustrated content shown above without departing from the gist of the technology of the present invention. Also, in order to avoid trouble and facilitate understanding of the parts related to the technology of the present invention, in the description content and illustrated content shown above, descriptions related to common technical knowledge that does not require special explanation when implementing the technology of the present invention are omitted.

[0121] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each document, patent application, and technical standard were specifically and separately incorporated by reference.

[0122] Through the above description, the following technology can be grasped.

[0123] [Appended Note Item 1]

[0124] A creation method, comprising:

[0125] An acquisition process of acquiring first audio data in floating-point form based on the sound emitted from a sound source and collected by a sound collection device; and

[0126] An additional information creation process creates additional information that is attached to the first sound data and includes device information as information related to the sound collection device or sound source information as information related to the sound source.

[0127] [Supplementary Note Item 2]

[0128] According to the creation method described in Supplementary Note Item 1, wherein,

[0129] The first sound data is used to create second sound data having a smaller number of bits than the first sound data.

[0130] [Supplementary Note Item 3]

[0131] According to the creation method described in Supplementary Note Item 1 or 2, wherein,

[0132] The additional information includes the device information.

[0133] [Supplementary Note Item 4]

[0134] According to the creation method described in Supplementary Note Item 3, wherein,

[0135] The device information is information related to gain processing used for the sound collected by the sound collection device or information related to the performance of the sound collection device.

[0136] [Supplementary Note Item 5]

[0137] According to the creation method described in Supplementary Note Item 2, wherein,

[0138] The additional information includes the sound source information.

[0139] [Supplementary Note Item 6]

[0140] According to the creation method described in Supplementary Note Item 5, wherein,

[0141] The sound source information has a corresponding relationship with the time information of the first sound data.

[0142] [Supplementary Note Item 7]

[0143] According to the creation method described in Supplementary Note Item 5 or 6, it further includes a photographing process,

[0144] In the photographing process, image data corresponding to the first sound data is created by a photographing device.

[0145] [Supplementary Note Item 8]

[0146] According to the creation method described in Supplementary Note Item 7, wherein,

[0147] The sound source is the subject included in the image data.

[0148] [Supplementary Note Item 9]

[0149] The creation method according to Supplementary Note Item 7, wherein,

[0150] The sound source is the main subject selected from among a plurality of subjects included in the image data.

[0151] [Supplementary Note Item 10]

[0152] The creation method according to Supplementary Note Item 7, wherein,

[0153] The sound source information is the type of driving sound accompanying the driving of the imaging device.

[0154] [Supplementary Note Item 11]

[0155] The creation method according to any one of Supplementary Note Items 1 to 10, further comprising a first file creation step,

[0156] In the first file creation step, a first file including the first sound data and the attached information is created.

[0157] [Supplementary Note Item 12]

[0158] The creation method according to Supplementary Note Item 11, further comprising an editing step,

[0159] In the editing step, a second sound data having a number of bits smaller than that of the first sound data is created by editing the first sound data according to the attached information.

[0160] [Supplementary Note Item 13]

[0161] The creation method according to Supplementary Note Item 12, wherein,

[0162] In the editing step, degradation information as information related to the sound of the second sound data degraded due to editing is created, and a second file including the second sound data and the degradation information is created.

[0163] [Supplementary Note Item 14]

[0164] The creation method according to Supplementary Note Item 13, wherein,

[0165] The second file includes the attached information.

Claims

1. A creation method, comprising: An acquisition process of acquiring first audio data in floating-point form based on the sound emitted from a sound source and collected by a sound collection device; And An additional information creation process of creating additional information attached to the first audio data and including device information as information related to the sound collection device or sound source information as information related to the sound source.

2. The creation method according to claim 1, wherein The first audio data is used to create second audio data having a smaller number of bits than the number of bits of the first audio data.

3. The creation method according to claim 2, wherein The additional information includes the device information.

4. The creation method according to claim 3, wherein The device information is information related to the gain process applied to the sound collected by the sound collection device or information related to the performance of the sound collection device.

5. The creation method according to claim 2, wherein The additional information includes the sound source information.

6. The creation method according to claim 5, wherein The sound source information has a corresponding relationship with the time information of the first audio data.

7. The creation method according to claim 5, further comprising a photographing process, In the photographing process, an imaging device creates imaging data corresponding to the first audio data.

8. The creation method according to claim 7, wherein The sound source is the subject included in the imaging data.

9. The creation method according to claim 7, wherein The sound source is the main subject selected from among a plurality of subjects included in the imaging data.

10. The creation method according to claim 7, wherein The sound source information is the type of drive sound accompanying the drive of the imaging device.

11. The creation method according to claim 1, further comprising a first file creation process, In the first file creation process, a first file including the first audio data and the additional information is created.

12. The creation method according to claim 11, further comprising an editing process, In the editing process, second audio data having a smaller number of bits than the number of bits of the first audio data is created by editing the first audio data according to the additional information.

13. The creation method according to claim 12, wherein In the editing process, degradation information as information related to the sound of the second audio data degraded due to editing is created, and a second file including the second audio data and the degradation information is created.

14. The creation method according to claim 13, wherein The second file includes the additional information.

15. A creation device, comprising a processor, The processor executes: An acquisition process of acquiring first audio data in floating-point form based on the sound emitted from a sound source and collected by a sound collection device; and An additional information creation process of creating additional information attached to the first audio data and including device information as information related to the sound collection device or sound source information as information related to the sound source.

Citation Information

Patent Citations

  • Data processing device, data processing method and digital audio mixer

    JP2002246913A

  • Voice signal converter

    JP2012073435A