Sound data creation method and sound data creation device

By recording and creating the sound data in the sound data creation method, and generating and processing sound data, the problem of insufficient sound data quality in the prior art is solved, and higher quality sound data creation is achieved.

CN120226341APending Publication Date: 2025-06-27FUJIFILM CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380079759.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-22
Filing Date
2023-10-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to improve the quality of sound data during the creation of sound data.

Method used

The first sound data of the first digit number is generated and recorded by the recording step, and the second sound data having the second digit number smaller than the first digit number and having directivity information is created based on the first sound data in the creation step.

Benefits of technology

Improve the quality of sound data, and enhance the stereo effect of sound data by enhancing the volume dynamic range and adding directional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226341A_ABST
    Figure CN120226341A_ABST
Patent Text Reader

Abstract

This sound data creation method comprises: a recording step for generating and recording first sound data of a first number of bits on the basis of a first sound signal output from a first sound collection element; and a creation step for creating, from the first voice data, second voice data having directivity information and a second number of digits smaller than the first number of digits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present invention relates to a method for creating sound data and a device for creating sound data. Background Art

[0002] Japanese Patent Application Laid-Open No. 2012-073435 discloses a sound signal conversion device that samples input analog sound signals of the L channel and the R channel at a sampling frequency of 192 kHz and a quantization bit number of 24 Bit using an A / D conversion device to generate digital signals. A signal processing device is connected to the output side of the A / D conversion device. The signal processing device performs processing for downsampling the frequency to 1 / 4 (48 kHz) and processing for converting the downsampled signal into a floating-point format with a quantization bit number of 32 Bit.

[0003] Japanese Patent Application Laid-Open No. 2002-246913 discloses a data processing device that converts input data from a fixed-point form to a floating-point form by a conversion unit. Summary of the Invention

[0004] Technical Problem to be Solved by the Invention

[0005] An object of an embodiment of the technology of the present invention is to provide a method for creating sound data and a device for creating sound data that can improve the quality of sound data.

[0006] Means for Solving the Technical Problem

[0007] To achieve the above object, the method for creating sound data of the present invention includes: a recording step of generating and recording first sound data of a first number of bits based on a first sound signal output from a first sound collecting element; and a creating step of creating second sound data having a second number of bits smaller than the first number of bits and having directivity information based on the first sound data.

[0008] Preferably, in the recording step, the first sound data is created by synthesizing a plurality of modulated sound data created by performing gain processing on the first sound signal a plurality of times.

[0009] Preferably, the first sound data is in a floating-point form.

[0010] Preferably, the second sound data is in a pulse code modulation form.

[0011] Preferably, the first sound data is in a mono form and the second sound data is in a stereo form.

[0012] Preferably, in the creating step, directivity information is obtained based on a plurality of second sound signals output from a plurality of second sound collecting elements.

[0013] Preferably, in the creation process, a sound data file including the first sound data is created.

[0014] Preferably, the second sound data is included in a moving image file created based on the image data output from the imaging element.

[0015] Preferably, the sound data file includes link information related to the moving image file.

[0016] In the creation process, the second sound data can be created from the first sound data using a pre-trained machine learning model.

[0017] Preferably, the pre-trained machine learning model is a model generated by performing machine learning using a plurality of learning sound data and correct answer data of directivity information, and the plurality of learning sound data is generated by collecting sound while changing the sound collection direction of the first sound collection element.

[0018] The sound data creation device of the present invention includes a processor that executes: a recording process of generating and recording first sound data of a first number of bits based on a first sound signal output from a first sound collection element; and a creation process of creating second sound data having a second number of bits smaller than the first number of bits and having directivity information based on the first sound data.

[0019] The sound data creation method of the present invention includes: a recording process of generating and recording first sound data of a first number of bits based on a first sound signal output from a first sound collection element; an acquisition process of acquiring device information of an output device that outputs sound based on second sound data having a second number of bits smaller than the first number of bits created from the first sound data; and a creation process of creating second sound data based on the first sound data and the device information.

[0020] Preferably, the device information is information related to the volume of the output device, the directivity angle information of the output device, or information related to the number of channels of the output device.

[0021] Preferably, the device information is information related to the volume, and the information related to the volume is information related to the efficiency of the output device.

[0022] The sound data creation device of the present invention includes a processor that executes: a recording process of generating and recording first sound data of a first number of bits based on a first sound signal output from a first sound collection element; an acquisition process of acquiring device information of an output device that outputs sound based on second sound data having a second number of bits smaller than the first number of bits created from the first sound data; and a creation process of creating second sound data based on the first sound data and the device information. Description of the Drawings

[0023] Figure 1This is a diagram showing an example of the structure of the imaging device according to the first embodiment.

[0024] Figure 2 This is a diagram showing an example of the structure of the sound signal processing circuit.

[0025] Figure 3 This is a diagram conceptually showing the sound signal processing.

[0026] Figure 4 This is a diagram showing an example of the functional structure of the processor.

[0027] Figure 5 This is a diagram conceptually showing the synthesis process and the data format conversion process.

[0028] Figure 6 This is a diagram conceptually showing the directivity information acquisition process.

[0029] Figure 7 This is a diagram conceptually showing the volume range setting process.

[0030] Figure 8 This is a diagram conceptually showing the data extraction process.

[0031] Figure 9 This is a diagram conceptually showing the conversion from the mono form to the stereo form.

[0032] Figure 10 This is a flowchart showing an example of the operation of the imaging device.

[0033] Figure 11 This is a diagram showing a modified example of the directivity information acquisition process.

[0034] Figure 12 This is a diagram showing an example of the functional structure of the processor according to the second embodiment.

[0035] Figure 13 This is a diagram conceptually showing an example of the learning process of the machine learning model.

[0036] Figure 14 This is a diagram showing an example of the functional structure of the processor according to the third embodiment.

[0037] Figure 15 This is a diagram conceptually showing the data extraction process based on the data extraction unit according to the third embodiment.

[0038] Figure 16 This is a flowchart showing an example of the operation of the imaging device according to the third embodiment. Detailed Embodiments

[0039] An example of an embodiment of the technology of the present invention will be described with reference to the accompanying drawings.

[0040] First, the terms used in the following description will be explained.

[0041] In the following description, "AF" is an abbreviation for "Auto Focus". "MF" is an abbreviation for "Manual Focus". "IC" is an abbreviation for "Integrated Circuit". "CPU" is an abbreviation for "Central Processing Unit". "RAM" is an abbreviation for "Random Access Memory". "CMOS" is an abbreviation for "Complementary Metal Oxide Semiconductor".

[0042] "FPGA" is an abbreviation for "Field Programmable Gate Array". "PLD" is an abbreviation for "Programmable Logic Device". "ASIC" is an abbreviation for "Application Specific Integrated Circuit". "OVF" is an abbreviation for "Optical View Finder". "EVF" is an abbreviation for "Electronic View Finder". "ADC" is an abbreviation for "Analog to Digital Converter". "LPCM" is an abbreviation for "Linear Pulse Code Modulation".

[0043] As an embodiment of the imaging device, an interchangeable-lens digital camera will be exemplified to describe the technology of the present invention. In addition, the technology of the present invention is not limited to interchangeable-lens type, and can also be applied to a lens-integrated type digital camera.

[0044] [First Embodiment]

[0045] Figure 1FIG. 0 shows an example of the structure of the imaging device 10 according to the first embodiment. The imaging device 10 is an interchangeable-lens digital camera. The imaging device 10 includes a housing 11 and an imaging lens 12 that is detachably attached to the housing 11 and includes a focusing lens 31. The imaging lens 12 is attached to the front surface side of the housing 11 via a bayonet 11A. In addition, the imaging device 10 is an example of a "sound data creation device" according to the technology of the present invention.

[0046] In addition, an external microphone 13 can be detachably attached to the housing 11. The external microphone 13 is attached to the housing 11 via a connection portion 11B provided on the upper surface of the housing 11. The external microphone 13 is, for example, a shotgun microphone, a zoom microphone, or the like. The connection portion 11B is, for example, a hot shoe.

[0047] An operation unit 16 including a dial, a release button, etc. is provided on the housing 11. As operation modes of the imaging device 10, for example, a still image shooting mode, a moving image shooting mode, and an image display mode are included. The operation unit 16 is operated by the user when setting the operation mode. In addition, the operation unit 16 is operated by the user when starting to perform still image shooting or moving image shooting.

[0048] In addition, the operation unit 16 is operated by the user when selecting a focusing mode. The focusing modes include an AF mode and an MF mode. The AF mode is a mode in which focusing control is performed by setting a subject area selected by the user or a subject area automatically detected by the imaging device 10 as a focus detection area (hereinafter, referred to as an AF area). The MF mode is a mode in which the user manually performs focusing control by operating a focusing ring (not shown).

[0049] In addition, a viewfinder 14 is provided on the housing 11. For example, the viewfinder 14 is a hybrid viewfinder (registered trademark). The hybrid viewfinder is, for example, a viewfinder that selectively uses an optical viewfinder (hereinafter, referred to as an "OVF") and an electronic viewfinder (hereinafter, referred to as an "EVF"). The user can observe an optical image or an instant preview image of the subject projected through the viewfinder 14 via a viewfinder eyepiece portion (not shown).

[0050] In addition, a display 15 is provided on the back side of the housing 11. An image based on the image data PD obtained by shooting and various menu screens, etc. are displayed on the display 15. The user can also observe the instant preview image projected through the display 15 instead of observing the viewfinder 14.

[0051] In addition, a speaker 17 is provided on the housing 11. The speaker 17 outputs sound according to the sound data included in the moving image file 28 described later. In addition, the speaker 17 is an example of an "output device" according to the technology of the present invention.

[0052] The housing 11 and the imaging lens 12 are electrically connected via the electrical contacts 11C provided on the bayonet 11A.

[0053] The imaging lens 12 includes a focusing lens 31, a diaphragm 32, and a lens drive control unit 33. The lens drive control unit 33 is electrically connected to the processor 25 accommodated in the housing 11 via the electrical contacts 11C.

[0054] The lens drive control unit 33 drives the focusing lens 31 and the diaphragm 32 according to the control signal sent from the processor 25. In order to adjust the position of the focusing lens 31, the lens drive control unit 33 performs drive control of the focusing lens 31 according to the focusing control signal sent from the processor 25.

[0055] The diaphragm 32 has an aperture with a variable opening diameter. In order to adjust the amount of incident light on the imaging sensor 20, the lens drive control unit 33 performs drive control of the diaphragm 32 according to the aperture adjustment control signal sent from the processor 25.

[0056] Moreover, an imaging sensor 20, an image processing circuit 21, a built-in microphone 22, a sound signal processing circuit 23, a processor 25, and a storage device 26 are provided inside the housing 11. The imaging sensor 20, the image processing circuit 21, the built-in microphone 22, the sound signal processing circuit 23, the storage device 26, the display 15, and the speaker 17 are controlled in operation by the processor 25.

[0057] The processor 25 is constituted by a CPU, for example. A RAM 25A as a primary storage memory is connected to the processor 25. The storage device 26 is constituted by a non-volatile memory such as a flash memory, for example. The processor 25 executes various processes according to the program 27 stored in the storage device 26. In addition, the processor 25 may be constituted by an aggregate of multiple IC chips. And, for example, a moving image file 28 generated as a result of the moving image photographing operation performed by the photographing device 10 is stored in the storage device 26.

[0058] The imaging sensor 20 is a CMOS image sensor, for example. The light (subject image) that has passed through the imaging lens 12 is incident on the light-receiving surface 20A of the imaging sensor 20. A plurality of pixels for generating a photographing signal by performing photoelectric conversion are formed on the light-receiving surface 20A. The imaging sensor 20 generates and outputs image data PD by performing photoelectric conversion on the light incident on each pixel. In addition, the imaging sensor 20 is an example of the "imaging element" related to the technology of the present invention.

[0059] The image processing circuit 21 performs image processing including white balance correction, gamma correction processing, etc. on the image data PD output from the imaging sensor 20.

[0060] The built-in microphone 22 is a stereo microphone having a pair of sound collecting elements 22A and 22B. The sound collecting elements 22A and 22B are sound sensors for the left channel (hereinafter referred to as the L channel) and the right channel (hereinafter referred to as the R channel). The sound collecting elements 22A and 22B are sound sensors such as electrostatic type, piezoelectric type, and electrodynamic type, and output the collected sound as sound signals AL and AR. The sound signal processing circuit 23 performs sound signal processing including gain processing, A / D conversion processing, etc. on the sound signals AL and AR output from the sound collecting elements 22A and 22B. In addition, the sound collecting elements 22A and 22B correspond to the "plurality of second sound collecting elements" related to the technology of the present invention. And the sound signals AL and AR correspond to the "plurality of second sound signals" related to the technology of the present invention.

[0061] The external microphone 13 includes a sound collecting element 41, an amplifier 42, and a microphone control unit 43. In the present embodiment, the external microphone 13 is a monaural microphone having one sound collecting element 41. The sound collecting element 41 is a sound sensor such as electrostatic type, piezoelectric type, and electrodynamic type, and outputs the collected sound as a sound signal. The amplifier 42 performs gain processing on the sound signal output from the sound collecting element 41. The microphone control unit 43 controls the gain amount based on the gain processing of the amplifier 42. In addition, the sound collecting element 41 corresponds to the "first sound collecting element" related to the technology of the present invention. And the sound signal output from the sound collecting element 41 corresponds to the "first sound signal" related to the technology of the present invention.

[0062] And the microphone control unit 43 supplies the sound signal subjected to gain processing by the amplifier 42 to the sound signal processing circuit 23 in the housing 11 via the connection portion 11B. A monaural and analog sound signal AS is supplied from the external microphone 13 to the sound signal processing circuit 23. In addition, the operation of the microphone control unit 43 is controlled by the processor 25.

[0063] Figure 2 An example of the structure of the sound signal processing circuit 23 is shown. The sound signal processing circuit 23 includes a first preamplifier 51A, a first ADC 52A, a second preamplifier 51B, and a second ADC 52B.

[0064] The first preamplifier 51A and the first ADC 52A are processing units for the L channel that perform gain processing and A / D conversion processing on the sound signal AL output from the sound collecting element 22A included in the built-in microphone 22. The second preamplifier 51B and the second ADC 52B are processing units for the R channel that perform gain processing and A / D conversion processing on the sound signal AR output from the sound collecting element 22B included in the built-in microphone 22.

[0065] The first preamplifier 51A has its gain amount G1 controlled by the processor 25. The second preamplifier 51B has its gain amount G2 controlled by the processor 25. When performing gain processing on the sound signals AL and AR output from the built-in microphone 22, the gain amount G1 and the gain amount G2 are set to the same value by the processor 25. The first ADC 52A and the second ADC 52B convert the analog sound signals into digital signals in the form of 24-bit LPCM, for example, by sampling with a quantization bit number of 24 bits. Additionally, the LPCM form is an example of the "pulse code modulation form" related to the technology of the present invention.

[0066] The sound signal AS output from the external microphone 13 is input to the first preamplifier 51A and the second preamplifier 51B. The first preamplifier 51A performs gain processing on the sound signal AS with the gain amount G1. The second preamplifier 51B performs gain processing on the sound signal AS with the gain amount G2. When performing gain processing on the sound signal AS output from the external microphone 13, the gain amount G1 and the gain amount G2 are set to different values by the processor 25. Hereinafter, the gain processing performed by the first preamplifier 51A is referred to as the first gain processing, and the gain processing performed by the second preamplifier 51B is referred to as the second gain processing.

[0067] The first ADC 52A converts the sound signal AS that has undergone the first gain processing by the first preamplifier 51A into a digital signal. The second ADC 52B converts the sound signal AS that has undergone the second gain processing by the second preamplifier 51B into a digital signal. Hereinafter, the sound signal AS digitized by the first ADC 52A is referred to as the modulated sound data ASH, and the sound signal AS digitized by the second ADC 52B is referred to as the modulated sound data ASL. The modulated sound data ASH and ASL are output from the sound signal processing circuit 23 to the processor 25.

[0068] Figure 3 Conceptually shows the sound signal processing of the sound signal AS based on the sound signal processing circuit 23. The sound signal AS output from the external microphone 13 is input to the processing unit for the L channel and the processing unit for the R channel. The sound signal AS input to the processing unit for the L channel is converted into a digital signal after undergoing the first gain processing with the gain amount G1, and thus is output from the sound signal processing circuit 23 as the modulated sound data ASH. The sound signal AS input to the processing unit for the R channel is converted into a digital signal after undergoing the second gain processing with the gain amount G2, and thus is output from the sound signal processing circuit 23 as the modulated sound data ASL. In the present embodiment, the number of bits of the modulated sound data ASH and ASL is 24 bits.

[0069] For example, the gain amount G1 is set to +48 dB, and the gain amount G2 is set to -48 dB. 48 dB corresponds to an 8-bit volume width. Therefore, as Figure 3 shown, the modulated sound data ASH with high gain and the modulated sound data ASL with low gain produce a deviation of 16 bits. In other words, the modulated sound data ASH and the modulated sound data ASL produce an 8-bit overlap.

[0070] Figure 4 An example of the functional structure of the processor 25 is shown. The processor 25 realizes various functional units by executing processing according to the program 27 stored in the storage device 26. Figure 4 The various functional units shown are realized in the moving image shooting mode. As Figure 4 shown, for example, the processor 25 realizes the main control unit 60, the synthesis processing unit 61, the data format conversion unit 62, the directivity information acquisition unit 63, the sound data file creation unit 64, the editing unit 65, and the file creation unit 66. The editing unit 65 includes a volume range setting unit 65A and a data extraction unit 65B.

[0071] The main control unit 60 uniformly controls each unit of the imaging device 10. The main control unit 60 controls the operation of the imaging device 10 according to the instruction signal input from the operation unit 16. The main control unit 60 controls the imaging sensor 20 to cause the imaging sensor 20 to perform a shooting operation. The imaging sensor 20 outputs the image data PD generated by shooting through the imaging lens 12. In the moving image shooting mode, the imaging sensor 20 outputs the image data PD at a frame period. The image data PD output from the imaging sensor 20 is input to the processor 25 after being subjected to image processing by the image processing circuit 21. In the case of the moving image shooting mode, the image data PD is data composed of a plurality of frames.

[0072] Moreover, in the moving image shooting mode, when the external microphone 13 is connected to the connection unit 11B, the main control unit 60 controls the external microphone 13 to perform a sound collection operation. The external microphone 13 outputs the sound signal AS to the sound signal processing circuit 23 via the connection unit 11B during the shooting operation of the imaging sensor 20. The sound signal processing circuit 23 outputs the modulated sound data ASH and ASL by performing the above-described sound signal processing. The modulated sound data ASH and ASL are sound data corresponding to the image data PD obtained by shooting the subject with the imaging sensor 20.

[0073] The synthesis processing unit 61 acquires the modulated sound data ASH and ASL output from the sound signal processing circuit 23, and synthesizes the modulated sound data ASH and ASL to create the first sound data AS1 of the first number of bits. The first sound data AS1 is digital data in the LPCM format.

[0074] The data format conversion unit 62 converts the data format of the first audio data AS1 into a floating-point format. Hereinafter, the first audio data AS1 converted into the floating-point format is referred to as the first audio data AS1F.

[0075] The directivity information acquisition unit 63 acquires directivity information DI based on a pair of audio signals AL and AR output from the built-in microphone 22 and subjected to audio signal processing by the audio signal processing circuit 23. For example, the directivity information DI is information indicating the volume difference between the L channel and the R channel.

[0076] The audio data file creation unit 64 creates an audio data file 67 including the first audio data AS1F created by the data format conversion unit 62 and the directivity information DI acquired by the directivity information acquisition unit 63. The audio data file creation unit 64 records the created audio data file 67 in the storage device 26.

[0077] The editing unit 65 refers to the audio data file 67 recorded in the storage device 26 and creates a second audio data AS2 having a second number of bits smaller than the first number of bits and having the directivity information DI based on the first audio data AS1F. For example, the second number of bits is 24 bits.

[0078] Specifically, the volume range setting unit 65A sets a volume range VR having a width of the second number of bits for the dynamic range of the first audio data AS1F. In the present embodiment, the volume range setting unit 65A sets the volume range VR based on the directivity information DI. The data extraction unit 65B extracts the data of the volume range VR set by the volume range setting unit 65A from the first audio data AS1F, thereby creating the second audio data AS2. The second audio data AS2 is digital data in a stereo format and an LPCM format.

[0079] The file creation unit 66 creates a moving image file 28 including the video data PD output from the image processing circuit 21 and the second audio data AS2 output from the data extraction unit 65B, and stores it in the storage device 26. In this way, the moving image file 28 includes the second audio data AS2 virtualized based on the directivity information DI acquired from the pair of audio signals AL and AR.

[0080] In addition, the file creation unit 66 may also create a normal moving image file 29 including the video data PD output from the image processing circuit 21 and a pair of audio signals AL and AR output from the built-in microphone 22 and subjected to audio signal processing by the audio signal processing circuit 23. In this way, the pair of audio signals AL and AR for acquiring the directivity information DI are the audio signals included in the normal moving image file 29.

[0081] Figure 5Conceptually shown are the synthesis process based on the synthesis processing unit 61 and the data format conversion process based on the data format conversion unit 62. The synthesis processing unit 61 synthesizes the modulated sound data ASH and the modulated sound data ASL by mixing the overlapping portions of the 8-bit amounts of the modulated sound data ASH and the modulated sound data ASL. The number of bits of the first sound data AS1 generated by this synthesis process (i.e., the first number of bits) is 40 bits. Thus, the first sound data AS1 with an expanded dynamic range of volume is obtained by synthesizing the modulated sound data ASH and the modulated sound data ASL with different synthesis gain amounts.

[0082] The data format conversion unit 62 converts the 40-bit fixed-point form first sound data AS1 into 32-bit floating-point form (so-called 32-bit floating-point number) first sound data AS1F. A 32-bit floating-point number consists of 1-bit sign, 8-bit exponent part, and 23-bit mantissa part. The conversion from fixed-point form to floating-point form can use a known method. Wide-range numerical expression can be performed in floating-point form.

[0083] Figure 6 Conceptually shown is the directivity information acquisition process based on the directivity information acquisition unit 63. The sound signals AL, AR are data representing the change in volume with respect to time (i.e., the change in amplitude). The above directivity information DI includes first difference information D1 and second difference information D2.

[0084] The directivity information acquisition unit 63 acquires the first difference information D1 by performing a difference operation of subtracting the sound signal AR from the sound signal AL. And, the directivity information acquisition unit 63 acquires the second difference information D2 by performing a difference operation of subtracting the sound signal AL from the sound signal AR. In Figure 6 the example shown, the first difference information D1 includes the signal in the time region mainly enclosed by a dashed line in the sound signal AL. The second difference information D2 includes the signal in the time region mainly enclosed by a dashed line in the sound signal AR. The first difference information D1 represents the information of the sound with a larger volume in the L channel than in the R channel. The second difference information D2 represents the information of the sound with a larger volume in the R channel than in the L channel.

[0085] Figure 7 Conceptually shown is the volume range setting process based on the volume range setting unit 65A. The above volume range VR includes a first volume range VR1 and a second volume range VR2.

[0086] The volume range setting unit 65A sets the first volume range VR1 based on the first difference information D1. Specifically, the volume range setting unit 65A sets the first volume range VR1 over time according to the volume included in the first difference information D1. For example, the larger the volume included in the first difference information D1, the higher on the high-volume side the volume range setting unit 65A sets the first volume range VR1. Similarly, the volume range setting unit 65A sets the second volume range VR2 based on the second difference information D2. Specifically, the volume range setting unit 65A sets the second volume range VR2 over time according to the volume included in the second difference information D2. For example, the larger the volume included in the second difference information D2, the higher on the high-volume side the volume range setting unit 65A sets the second volume range VR2.

[0087] Therefore, within the time range where the volume is larger in the L channel than in the R channel, the first volume range VR1 is set on the high-volume side. Within the time range where the volume is larger on the R channel side than in the L channel, the second volume range VR2 is set on the high-volume side.

[0088] Figure 8 Conceptually shows the data extraction process by the data extraction unit 65B. The data extraction unit 65B extracts data of the first volume range VR1 from the first audio data AS1F, thereby creating the second audio data AS2L in 24-bit fixed-point format. Specifically, the data extraction unit 65B selects the values of the sign and exponent parts of the 32-bit floating-point number according to the first volume range VR1, thereby creating the 24-bit second audio data AS2L represented by the mantissa part. And the data extraction unit 65B extracts data of the second volume range VR2 from the first audio data AS1F, thereby creating the second audio data AS2R in 24-bit fixed-point format. Specifically, the data extraction unit 65B selects the values of the sign and exponent parts of the 32-bit floating-point number according to the second volume range VR2, thereby creating the 24-bit second audio data AS2R represented by the mantissa part. The above second audio data AS2 includes the second audio data AS2L and the second audio data AS2R.

[0089] As Figure 9 shown, the first audio data AS1F is in mono format. By respectively extracting data of the first volume range VR1 and the second volume range VR2 from the first audio data AS1F, it is possible to create the second audio data AS2 in stereo format including the second audio data AS2L corresponding to the L channel and the second audio data AS2R corresponding to the R channel. That is, the second audio data AS2 is stereo audio data with directivity information DI.

[0090] Figure 10 Is a flowchart showing an example of the operation of the imaging device 10. Figure 10This represents the operation when the dynamic image shooting mode is selected as the operation mode and the external microphone 13 is connected to the connection part 11B.

[0091] First, the main control unit 60 determines whether there is an instruction to start shooting dynamic images by the user (step S10). When it is determined that there is a start instruction (step S10: Yes), the shooting process (step S11) and the recording process (step S12) are executed simultaneously. In the shooting process, the imaging sensor 20 captures the subject to generate image data PD. In the recording process, the external microphone 13 and the built-in microphone 22 collect sound. And, in the recording process, first sound data AS1 with the first number of digits is created based on the sound signal output from the sound collection element 41 of the external microphone 13. In the present embodiment, the first sound data AS1 is converted into the first sound data AS1F in floating-point form. And, in the recording process, the directivity information DI is obtained based on the sound signals AL and AR output from the pair of sound collection elements 22A and 22B of the built-in microphone 22. Furthermore, a sound data file 67 including the first sound data AS1F and the directivity information DI is created and recorded in the storage device 26.

[0092] After the shooting process and the recording process, the main control unit 60 determines whether there is an instruction to end shooting dynamic images by the user (step S13). When it is determined that there is no end instruction (step S13: No), the process returns to steps S11 and S12. Steps S11 and S12 are repeatedly executed until it is determined that there is an end instruction in step S13.

[0093] When it is determined that there is an end instruction (step S13: Yes), the creation process (step S14) is executed. In the creation process, the sound data file 67 recorded in the storage device 26 is read out, and second sound data AS2 with the second number of digits smaller than the first number of digits and having the directivity information DI is created based on the first sound data AS1F. And, in the creation process, a dynamic image file 28 including the image data PD and the second sound data AS2 is created and recorded in the storage device 26. Thus, the operation of the imaging device 10 ends.

[0094] As described above, the sound data creation method of the present invention includes: a recording process of generating and recording first sound data with the first number of digits based on the sound signal output from the first sound collection element; and a creation process of creating second sound data with the second number of digits smaller than the first number of digits and having the directivity information based on the first sound data. Thereby, the quality of the sound data can be improved.

[0095] In addition, in the above-described embodiment, the directivity information acquisition unit 63 acquires the directivity information DI based on the sound signals AL and AR input from the image processing circuit 21 to the processor 25. However, the directivity information DI may also be acquired based on the sound signals AL and AR included in the moving image file 29. In this case, as Figure 11 shown, preferably, the sound data file 67 includes the first sound data AS1F and the link information 68 related to the moving image file 29. The link information 68 is information indicating the link destination of the moving image file 29. For example, the link information 68 is the address information of the moving image file 29, the file name information of the moving image file 29, or the like.

[0096] As Figure 11 shown, the directivity information acquisition unit 63 supplies the directivity information DI acquired based on the sound signals AL and AR included in the moving image file 29 to the volume range setting unit 65A of the editing unit 65. The processing of the editing unit 65 is the same as that in the above-described embodiment.

[0097] Moreover, in the above-described embodiment, the built-in microphone 22 includes a pair of sound collecting elements 22A and 22B. However, the number of sound collecting elements is not limited to two, and the built-in microphone 22 may also include three or more sound collecting elements. That is, the directivity information acquisition unit 63 may acquire directivity information DI for three or more channels based on three or more sound signals output from the built-in microphone 22. In this case, the second sound data AS2 is multi-channel sound data. Also, the built-in microphone 22 may be a digital microphone that outputs sound signals AL and AR in digital form.

[0098] [Second Embodiment]

[0099] Next, the second embodiment will be described. In the first embodiment, the mono-form first sound data AS1F is converted into stereo-form second sound data AS2 using the directivity information DI acquired by the directivity information acquisition unit 63. In the second embodiment, the directivity information acquisition unit 63 is not provided, and the mono-form first sound data AS1F is converted into stereo-form second sound data AS2 using a machine learning model.

[0100] The configuration of the imaging device 10 according to the second embodiment other than the processor 25 is the same as that of the first embodiment. Hereinafter, the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof will be appropriately omitted.

[0101] Figure 12An example of the functional structure of the processor 25 according to the second embodiment is shown. In the present embodiment, the processor 25 implements the main control unit 60, the synthesis processing unit 61, the data format conversion unit 62, the sound data file creation unit 64, and the machine-learned model 70. In the present embodiment, the directivity information acquisition unit 63 is not configured in the processor 25. Therefore, the sound data file creation unit 64 creates a sound data file 67 that includes only the first sound data AS1F created by the data format conversion unit 62, and records it in the storage device 26.

[0102] The main control unit 60 reads the first sound data AS1F from the sound data file 67 recorded in the storage device 26, and inputs it to the machine-learned model 70. The machine-learned model 70 is, for example, a neural network that has been machine-learned by deep learning. The machine-learned model 70 converts the input first sound data AS1F in mono form into second sound data AS2 in stereo form, and outputs it.

[0103] Figure 13 An example of the learning process of the machine-learned model 70 is conceptually shown. As Figure 13 shown, the machine-learned model 70 is generated by causing the machine learning model 71 to perform machine learning using the training data 72 in the learning phase. The training data 72 is composed of a set of a plurality of learning sound data 72A and a plurality of correct answer data 72B. For example, the learning sound data 72A is sound data generated by collecting sound while changing the sound collection direction of the sound collection element 41. For example, the correct answer data 72B is correct answer data for directivity information.

[0104] The machine learning model 71 performs machine learning using, for example, the error backpropagation method. In the learning phase, the error operation and the update setting are repeatedly performed. The error operation is an operation of calculating the error between the directivity information included in the sound data output from the machine learning model 71 when the learning sound data 72A is input to the machine learning model 71 and the correct answer data 72B. The update setting is a process of setting the weights and biases of the machine learning model 71 in such a way as to reduce the error. The machine learning of the machine learning model 71 is performed, for example, in an information processing device outside the imaging device 10. The machine learning model 71 after performing machine learning is stored in the storage device 26 of the imaging device 10 as the above-mentioned machine-learned model 70. The machine-learned model 70 stored in the storage device 26 is used by the processor 25.

[0105] [Third Embodiment]

[0106] Next, a description will be given of the third embodiment. In the first embodiment, the editorial department 65 created the second sound data AS2 based on the first sound data AS1F and the directivity information DI. In the third embodiment, the second sound data AS2 is created based on the first sound data AS1F and the device information of the speaker 17.

[0107] The structure of the imaging device 10 other than the processor 25 according to the third embodiment is the same as that of the first embodiment. Hereinafter, the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof will be appropriately omitted.

[0108] Figure 14 An example of the functional structure of the processor 25 according to the third embodiment is shown. In the present embodiment, the processor 25 implements the main control unit 60, the synthesis processing unit 61, the data format conversion unit 62, the sound data file creation unit 64, and the editorial department 65. In the present embodiment, the directivity information acquisition unit 63 is not configured in the processor 25. Therefore, the sound data file creation unit 64 creates a sound data file 67 that includes only the first sound data AS1F created by the data format conversion unit 62, and records it in the storage device 26.

[0109] The storage device 26 stores the device information 80 of the speaker 17. The device information 80 is information related to the characteristics of the speaker 17. For example, the device information 80 is information related to the volume of the speaker 17, the directivity angle information of the speaker 17, or the number of channels of the speaker 17. And, for example, the information related to the volume of the speaker 17 is information related to the efficiency of the speaker 17. When a signal power of 1 W is input to the speaker 17, the efficiency is represented by the sound pressure (dB) at a position 1 m away from the speaker 17. The directivity angle is represented by the angle up to the position where the sound pressure is reduced by 6 dB with respect to the reference sound pressure at the position directly below the speaker 17.

[0110] In the present embodiment, the volume range setting unit 65A acquires the device information 80 from the storage device 26, and sets the volume range VR according to the acquired device information 80. For example, the higher the efficiency of the speaker 17, the more the volume range setting unit 65A sets the volume range VR on the high-volume side. And, the larger the directivity angle of the speaker 17, the more the volume range setting unit 65A sets the volume range VR on the high-volume side. Further, the more the number of channels of the speaker 17, the more the volume range setting unit 65A sets the volume range VR on the high-volume side.

[0111] Figure 15Conceptually shows the data extraction process of the data extraction unit 65B according to the third embodiment. In the present embodiment, the data extraction unit 65B extracts data in the volume range VR from the first audio data AS1F, thereby creating the second audio data AS2 in 24-bit fixed-point format. In the present embodiment, the second audio data AS2 is in mono format.

[0112] Figure 16 Is a flowchart showing an example of the operation of the imaging device 10 according to the third embodiment. Figure 16 Represents the operation when the moving image shooting mode is selected as the operation mode and the external microphone 13 is connected to the connection unit 11B.

[0113] First, the main control unit 60 determines whether there is an instruction to start moving image shooting by the user (step S20). When it is determined that there is a start instruction (step S20: Yes), the shooting process (step S21) and the recording process (step S22) are executed simultaneously. In the shooting process, the imaging sensor 20 shoots the subject to generate image data PD. In the recording process, the external microphone 13 collects sound. And, in the recording process, the first audio data AS1 of the first number of bits is created based on the sound signal output from the sound collection element 41 of the external microphone 13. In the present embodiment, the first audio data AS1 is converted into the first audio data AS1F in floating-point format. Further, an audio data file 67 including the first audio data AS1F is created and recorded in the storage device 26.

[0114] After the shooting process and the recording process, the main control unit 60 determines whether there is an instruction to end moving image shooting by the user (step S23). When it is determined that there is no end instruction (step S23: No), the process returns to steps S21 and S22. Steps S21 and S22 are repeatedly executed until it is determined that there is an end instruction in step S23.

[0115] When it is determined that there is an end instruction (step S23: Yes), the acquisition process (step S24) is executed. In the acquisition process, the volume range setting unit 65A acquires the device information 80 from the storage device 26. The volume range setting unit 65A sets the volume range VR according to the acquired device information 80.

[0116] After the acquisition process, a creation process (step S25) is performed. In the creation process, the data extraction unit 65B extracts data in the volume range VR from the first audio data AS1F, thereby creating the second audio data AS2. And, in the creation process, a moving image file 28 including the image data PD and the second audio data AS2 is created and recorded in the storage device 26. Thus, the operation of the imaging device 10 ends.

[0117] [Modification Example]

[0118] The technology of the present invention is not limited to digital cameras, and can also be applied to electronic devices such as smartphones and tablet terminals with a camera function.

[0119] In the above embodiments, as the hardware structure of the control unit exemplified by the processor 25, various processors shown below can be used. The above various processors include a CPU, which is a common processor that functions as an execution software (program), and in addition, processors such as FPGAs that can change the circuit structure after manufacturing are also included. The FPGA includes a dedicated circuit, etc., and the dedicated circuit is a processor such as a PLD or an ASIC having a circuit structure specifically designed for executing a specific process.

[0120] The control unit can be constituted by one of these various processors, or can be constituted by a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, multiple control units can be constituted by one processor.

[0121] Multiple examples can be conceived of a single processor constituting multiple control units. The first example has the following method: represented by computers such as clients and servers, one or more CPUs and software are combined to form one processor, and this processor functions as multiple control units. The second example has the following method: represented by a System On Chip (SOC), etc., a processor that realizes the functions of the entire system including multiple control units using one IC chip is used. Thus, as the hardware structure, the control unit can be constituted by using one or more of the above various processors.

[0122] Furthermore, as the hardware structure of these various processors, more specifically, a circuit formed by combining circuit elements such as semiconductor elements can be used.

[0123] The above-described recorded content and illustrated content are detailed descriptions of parts related to the technology of the present invention, and are only examples of the technology of the present invention. For example, the descriptions related to the above structure, function, action, and effect are descriptions related to an example of the structure, function, action, and effect of the parts related to the technology of the present invention. Therefore, of course, it is also possible to delete unnecessary parts, add new elements, or make replacements to the above-described recorded content and illustrated content without departing from the gist of the technology of the present invention. Also, in order to avoid trouble and facilitate understanding of the parts related to the technology of the present invention, descriptions related to common technical knowledge that do not require special explanation when implementing the technology of the present invention are omitted in the above-described recorded content and illustrated content.

[0124] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each document, patent application, and technical standard were specifically and individually set forth.

[0125] From the above description, the following technology can be grasped.

[0126] [Supplementary Note Item 1]

[0127] A method for creating voice data, comprising:

[0128] A recording process for generating and recording first voice data of a first number of digits based on a first voice signal output from a first sound collection element; and

[0129] A creation process for creating second voice data having a second number of digits smaller than the first number of digits and having directivity information based on the first voice data.

[0130] [Supplementary Note Item 2]

[0131] The method for creating voice data according to Supplementary Note Item 1, wherein

[0132] In the recording process, the first voice data is created by synthesizing a plurality of modulated voice data created by performing gain processing on the first voice signal multiple times.

[0133] [Supplementary Note Item 3]

[0134] The method for creating voice data according to Supplementary Note Item 1 or 2, wherein

[0135] The first voice data is in floating-point form.

[0136] [Supplementary Note Item 4]

[0137] The method for creating voice data according to any one of Supplementary Note Items 1 to 3, wherein

[0138] The second voice data is in pulse code modulation form.

[0139] [Supplementary Note Item 5]

[0140] The method for creating voice data according to any one of Supplementary Note Items 1 to 4, wherein

[0141] The first voice data is in mono form and the second voice data is in stereo form.

[0142] [Supplementary Note Item 6]

[0143] The method for creating voice data according to any one of Supplementary Note Items 1 to 5, wherein

[0144] In the creation process, the directivity information is obtained based on a plurality of second sound signals output from a plurality of second sound collection elements.

[0145] [Supplementary Note Item 7]

[0146] According to the sound data creation method described in Supplementary Note Item 6, wherein

[0147] In the creation process, a sound data file containing the first sound data is created.

[0148] [Supplementary Note Item 8]

[0149] According to the sound data creation method described in Supplementary Note Item 7, wherein

[0150] The second sound data is included in a moving image file created based on image data output from an imaging element.

[0151] [Supplementary Note Item 9]

[0152] According to the sound data creation method described in Supplementary Note Item 8, wherein

[0153] The sound data file contains link information related to the moving image file.

[0154] [Supplementary Note Item 10]

[0155] According to the sound data creation method described in any one of Supplementary Note Items 1 to 5, wherein

[0156] In the creation process, the second sound data is created from the first sound data using a pre-trained machine learning model.

[0157] [Supplementary Note Item 11]

[0158] According to the sound data creation method described in Supplementary Note Item 10, wherein

[0159] The pre-trained machine learning model is a model generated by performing machine learning using a plurality of learning sound data and correct answer data of the directivity information, and the plurality of learning sound data is generated by collecting sound while changing the sound collection direction of the first sound collection element.

Claims

1. A method for creating sound data, comprising: A recording process of generating and recording first sound data of a first number of digits based on a first sound signal output from a first sound collection element; And A creation process of creating second sound data having a second number of digits smaller than the first number of digits and having directivity information based on the first sound data.

2. The method for creating sound data according to claim 1, wherein In the recording process, the first sound data is created by synthesizing a plurality of modulated sound data created by performing gain processing on the first sound signal multiple times.

3. The method for creating sound data according to claim 2, wherein The first sound data is in floating-point form.

4. The method for creating sound data according to claim 3, wherein The second sound data is in pulse code modulation form.

5. The method for creating sound data according to claim 1, wherein The first sound data is in mono form and the second sound data is in stereo form.

6. The method for creating sound data according to claim 1, wherein In the creation process, the directivity information is obtained based on a plurality of second sound signals output from a plurality of second sound collection elements.

7. The method for creating sound data according to claim 6, wherein In the creation process, a sound data file containing the first sound data is created.

8. The method for creating sound data according to claim 7, wherein The second sound data is included in a moving image file created based on image data output from an imaging element.

9. The method for creating sound data according to claim 8, wherein The sound data file contains link information related to the moving image file.

10. The method for creating sound data according to claim 1, wherein In the creation process, the second sound data is created from the first sound data using a trained machine learning model.

11. The method for creating sound data according to claim 10, wherein The trained machine learning model is a model generated by performing machine learning using a plurality of learning sound data and correct answer data of the directivity information, and the plurality of learning sound data are generated by collecting sound while changing the sound collection direction of the first sound collection element.

12. A sound data creation device comprising a processor, The processor executes: A recording process of generating and recording first sound data of a first number of digits based on a first sound signal output from a first sound collection element; and A creation process of creating second sound data having a second number of digits smaller than the first number of digits and having directivity information based on the first sound data.

13. A method for creating sound data, comprising: A recording process of generating and recording first sound data of a first number of digits based on a first sound signal output from a first sound collection element; An acquisition process of acquiring device information of an output device that outputs sound based on second sound data having a second number of digits smaller than the first number of digits created from the first sound data; And A creation process of creating the second sound data based on the first sound data and the device information.

14. The method for creating sound data according to claim 13, wherein the device information is information related to the volume of the output device, the pointing angle information of the output device, or information related to the number of channels of the output device.

15. The method for creating sound data according to claim 14, wherein the device information is information related to the volume, the information related to the volume is information related to the efficiency of the output device.

16. A sound data creation device, comprising a processor, wherein the processor executes: a recording process of generating and recording first sound data of a first number of bits based on a first sound signal output from a first sound collecting element; an acquisition process of acquiring device information of an output device that outputs sound based on second sound data of a second number of bits created from the first sound data and smaller than the first number of bits; and a creation process of creating the second sound data based on the first sound data and the device information.

Citation Information

Patent Citations

  • Data processing device, data processing method and digital audio mixer

    JP2002246913A

  • Voice signal converter

    JP2012073435A