Information processing method, information processing program, and information processing device
By converting inaudible sound data into audible sound data and processing it through a learning model, the method addresses the challenge of high learning load, facilitating accurate feature derivation in sound data processing.
Patent Information
- Application Number
- PCT/JP2025/003591
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-02-04
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional technologies face challenges in deriving highly accurate features from sound data in the inaudible band due to the large learning load of the learning model, making it difficult to easily obtain precise features.
An information processing method that converts first-band sound data, including an inaudible sound band, into second-band sound data, which includes an audible sound band, and inputs the converted data into a learning model to derive feature quantities.
Enables the easy derivation of highly accurate feature quantities of sound data by reducing the learning load and improving the accuracy of feature extraction.
Smart Images

Figure JP2025003591_30102025_PF_FP_ABST
Abstract
Description
Information processing method, information processing program, and information processing device
[0001] The present disclosure relates to an information processing method, an information processing program, and an information processing device.
[0002] Conventionally, a technology for obtaining features of sound data using a learning model has been disclosed. For example, a system has been disclosed in which sound data in the ultrasonic band, which is an inaudible sound band, is input into a first learning model, and based on the estimation result, which is a feature obtained by inputting sound data into a second learning model, the output result obtained is estimated as human behavior. Also, for example, a system has been disclosed in which human behavior information is estimated by inputting sound data in the inaudible sound band into a learning model that shows the relationship between the sound data in the inaudible sound band and behavioral information, which is a feature.
[0003] International Publication No. WO 2023 / 008260 International Publication No. WO 2022 / 054407
[0004] Here, sound data in the inaudible band may be used to derive highly accurate features because, compared to sound data in the audible band, the inaudible band does not contain noise in the audible band and is more susceptible to attenuation. However, the learning load of a learning model that inputs sound data in the inaudible band is large. For this reason, in conventional technologies that require a learning model that inputs sound data in the inaudible band, the learning load of the learning model is large, making it difficult to easily derive highly accurate features of sound data.
[0005] The problem to be solved by the present disclosure is to provide an information processing method, an information processing program, and an information processing device that can easily derive highly accurate features of sound data.
[0006] The information processing method according to the present disclosure is an information processing method executed by a computer, which acquires sound data, converts first-band sound data of a first band including an inaudible sound band contained in the sound data into converted second-band sound data of a second band including an audible sound band, and inputs the converted second-band sound data to a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band data.
[0007] According to the information processing method, information processing program, and information processing device disclosed herein, highly accurate feature quantities of sound data can be easily derived.
[0008] FIG. 1 is a functional block diagram illustrating an example of an information processing apparatus according to an embodiment. FIG. 2 is a hardware configuration diagram of the example of the information processing apparatus. FIG. 3A is a schematic diagram illustrating an example of sound data including first band sound data and second band sound data. FIG. 3B is a schematic diagram illustrating an example of first band sound data extracted from sound data. FIG. 4 is a schematic diagram of an example of first band sound data. FIG. 5A is an explanatory diagram of an example of a conversion process. FIG. 5B is an explanatory diagram of an example of a conversion process. FIG. 5C is an explanatory diagram of an example of a conversion process. FIG. 5D is an explanatory diagram of an example of a conversion process. FIG. 5E is an explanatory diagram of an example of a conversion process. FIG. 6 is a schematic diagram of an example of first band sound data. FIG. 7 is a schematic diagram of an example of the data configuration of a feature amount management DB. FIG. 8 is a flowchart illustrating an example of the flow of information processing executed by the information processing apparatus.
[0009] Hereinafter, embodiments of an information processing method, an information processing program, and an information processing device according to the present disclosure will be described with reference to the accompanying drawings.
[0010] FIG. 1 is a functional block diagram showing an example of an information processing apparatus 10 according to the present embodiment.
[0011] The information processing device 10 is an information processing device that derives feature amounts of sound data from the sound data and determines the sound source type of the sound data using the feature amounts (described in detail below).
[0012] FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing device 10. As shown in FIG.
[0013] The information processing device 10 has a hardware configuration that utilizes a typical computer, with a CPU (Central Processing Unit) 10A, a ROM (Read Only Memory) 10B, a RAM (Random Access Memory) 10C, and an I / F (Interface) 10D, etc., all interconnected by a bus 10E.
[0014] The CPU 10A is a computing device that controls the information processing device 10 of this embodiment. The ROM 10B stores programs and the like that realize various processes by the CPU 10A. The RAM 10C stores data necessary for various processes by the CPU 10A. The I / F 10D is an interface for sending and receiving data.
[0015] The program for executing information processing executed by the information processing device 10 of this embodiment is provided by being pre-installed in the ROM 10B, etc. The program executed by the information processing device 10 of this embodiment may be provided by being recorded on a computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, or a digital versatile disc (DVD) in a format that can be installed or executed by the information processing device 10.
[0016] Returning to FIG. 1, the explanation will be continued.
[0017] The information processing device 10 includes a sound collection unit 11, a communication unit 12, an input unit 13, an output unit 14, a storage unit 15, and a processing unit 20. The sound collection unit 11, the communication unit 12, the input unit 13, the output unit 14, the storage unit 15, and the processing unit 20 are communicatively connected via a bus 16 or the like.
[0018] The sound collection unit 11 is sensitive to sounds in both the audible and inaudible frequency bands and collects sounds in these frequency bands. The audible band is a frequency band of audible sounds that can be heard by the human ear. The audible band is, for example, a frequency band of 0 kHz to 20 kHz, but is not limited to this numerical range. The inaudible band is a frequency band of inaudible sounds that cannot be heard by the human ear. The inaudible band is, for example, a frequency band above 20 kHz, but is not limited to this numerical range. The sound collection unit 11 is, for example, a microphone such as a MEMS (Micro Electronics Mechanical System).
[0019] The sound collection unit 11 may be provided outside the information processing device 10. In this case, the sound collection unit 11 and the information processing device 10 may be connected to each other via a network or the like so as to be able to communicate with each other by wire or wirelessly.
[0020] The sound collection unit 11 collects sounds including audible sounds and inaudible sounds, converts the collected sounds into electrical signals to generate sound data, and outputs the generated sound data to the processing unit 20.
[0021] The sound data includes sounds in both audible and inaudible frequency bands. The sound data is digital sound pressure data that has been AD (Analog-to-Digital conversion) converted at a predetermined sampling period and arranged in time series, and is data of a predetermined time width that corresponds to the duration of sound collection. The sound data may also include periods of silence.
[0022] The communication unit 12 communicates with an external information processing device via a network or the like. The input unit 13 accepts various operations by a user. The input unit 13 is, for example, a pointing device such as a mouse, a keyboard, etc. The output unit 14 outputs various information. The output unit 14 is, for example, a display that displays various information, a speaker that outputs various sounds, etc. At least one of the input unit 13 and the output unit 14 may be provided outside the information processing device 10. In this case, the input unit 13 and the output unit 14 provided outside the information processing device 10 may be connected to the information processing device 10 via a network or the like so as to be able to communicate with each other via wired or wireless communication.
[0023] The storage unit 15 stores various types of data. The storage unit 15 is, for example, a semiconductor memory element such as a RAM or a flash memory, a hard disk, an optical disk, or the like. The storage unit 15 may be a storage device provided outside the information processing device 10. The storage unit 15 may also be a storage medium that stores or temporarily stores programs and various types of information after downloading them via a LAN (Local Area Network), the Internet, or the like. At least a portion of the information stored in the storage unit 15 may also be stored in a storage device such as an external server device communicably connected to the information processing device 10.
[0024] In this embodiment, the storage unit 15 pre-stores a learning model 15A and a feature management DB (database) 15B. Details of the learning model 15A and the feature management DB 15B will be described later.
[0025] The processing unit 20 is a computer that executes information processing in the information processing device 10 .
[0026] The processing unit 20 includes an acquisition unit 20A, an extraction unit 20B, a determination unit 20C, a conversion unit 20D, a feature derivation unit 20E, a sound source type determination unit 20F, and an output control unit 20G. The acquisition unit 20A, the extraction unit 20B, the determination unit 20C, the conversion unit 20D, the feature derivation unit 20E, the sound source type determination unit 20F, and the output control unit 20G are implemented by one or more processors. For example, each of the above units may be implemented by a processor such as a CPU executing a program, i.e., software. Each of the above units may be implemented by a processor such as a dedicated IC, i.e., hardware. Each of the above units may be implemented by a combination of software and hardware. When multiple processors are used, each processor may implement one of the units or two or more of the units. Furthermore, at least one of the above units included in the information processing device 10 may be installed in an external information processing device communicatively connected to the information processing device 10 via a network or the like.
[0027] The acquisition unit 20A acquires sound data collected by the sound collection unit 11. As described above, the sound data includes sounds in both the audible and inaudible frequency bands. The acquisition unit 20A may acquire sound data by reading sound data collected by the sound collection unit 11 or the like and stored in the storage unit 15 from the storage unit 15. The acquisition unit 20A may also acquire sound data by receiving sound data collected by the sound collection unit 11 or the like from an external information processing device or the like via the communication unit 12.
[0028] The extraction unit 20B extracts first-band sound data of a first band included in the sound data acquired by the acquisition unit 20A. The first band is a frequency band that includes at least an inaudible sound band. In this embodiment, a case where the first band is an inaudible sound band will be described as an example. The extraction unit 20B also extracts second-band sound data of a second band included in the sound data. The second band is a frequency band that includes at least an audible sound band. In this embodiment, a case where the second band is an audible sound band will be described as an example.
[0029] The extraction unit 20B may extract each of the first band sound data and the second band sound data contained in the sound data from the sound data using a known FIR (Finite Impulse Response) filter, etc. Examples of FIR filters include a low-pass filter, a high-pass filter, a band-pass filter, and a band-eliminate filter.
[0030] Fig. 3A is a schematic diagram showing an example of sound data 30 including first-band sound data 31 and second-band sound data 32. Fig. 3B is a schematic diagram showing an example of first-band sound data 31 extracted from the sound data 30. The horizontal axis of Fig. 3A and Fig. 3B represents time, and the vertical axis represents amplitude.
[0031] Through the extraction process by the extraction unit 20B, first band sound data 31 and second band sound data 32 are extracted from the sound data 30. Note that the extraction unit 20B extracts the first band sound data 31 from the sound data 30, and extracts the second band sound data 32 as needed through a process described below.
[0032] Returning to FIG. 1, the explanation will be continued.
[0033] The determination unit 20C, conversion unit 20D, feature derivation unit 20E, and sound source type determination unit 20F perform the following processes for each processing unit time on the sound data 30 acquired by the acquisition unit 20A. The processing unit time may be a time that matches one sampling period during the AD conversion or a time that is equal to or longer than this period. The processing unit time is, for example, one second or two seconds, but is not limited to these times.
[0034] The judgment unit 20C judges for each processing unit time whether to use converted second-band sound data (described in detail later) obtained by converting first-band sound data 31 included in the sound data 30 into the second band, second-band sound data 32 included in the sound data 30, or the sound data 30 itself to derive the feature quantities of the sound data 30.
[0035] Fig. 4 is a schematic diagram of an example of the first band sound data 31. Fig. 4 is a graph that schematically shows the first band sound data 31 represented by a waveform such as that shown in Fig. 3B etc. In Fig. 4, the horizontal axis represents time and the vertical axis represents amplitude.
[0036] For example, it is assumed that the first band sound data 31 included in the sound data 30 acquired by the acquisition unit 20A is represented by the graph shown in FIG.
[0037] The determination unit 20C determines, for each processing unit time, whether the amplitude of the first-band sound data 31 is equal to or greater than the second amplitude T2. Then, the determination unit 20C identifies a section of the processing unit time in which it is determined that the amplitude of the first-band sound data 31 is equal to or greater than the second amplitude T2 as a sound detection section in which some sound has been detected. Furthermore, in this embodiment, the determination unit 20C determines that, within the sound detection section, converted second-band sound data (described later) obtained by converting the first-band sound data 31 into a second band should be used to derive the feature amount.
[0038] The second amplitude T2 may be set to a value less than the first amplitude T1 (described later) and may have a preset threshold value for determining whether or not sound is being generated. The second amplitude T2 may be changeable as needed within a range that satisfies the above conditions, for example, by a user operating the input unit 13.
[0039] The first-band sound data 31 tends to be less susceptible to the influence of noise than the second-band sound data 32. For this reason, the determination unit 20C determines whether or not a sound detection section exists for each processing unit time using the first-band sound data 31, thereby making it possible to determine with high accuracy whether or not the timing of each processing unit time in the sound data 30 exists as a sound detection section in the sound data 30.
[0040] In the example shown in FIG. 4, the determination unit 20C sequentially performs the above determination for each processing unit time for the entire period of the first band sound data 31, and thereby identifies, among the sections s1 to s12 that make up the entire period, sections s2, s4, s6, s8, s10, and s12 that have amplitudes equal to or greater than the second amplitude T2, as sound detection sections.
[0041] When the determining unit 20C determines that the amplitude of the first band sound data 31 is equal to or greater than the second amplitude T2, the determining unit 20C further determines whether the amplitude of the first band sound data 31 is equal to or greater than the first amplitude T1.
[0042] The first amplitude T1 is an example of a predetermined amplitude. The first amplitude T1 is larger than the second amplitude T2, and a threshold value for determining whether to use the converted second-band sound data obtained by converting the first-band sound data 31, the second-band sound data 32, or the sound data 30 to derive the feature may be set in advance. For example, the first amplitude T1 may be set in advance as an amplitude value at which the accuracy of the feature derived using the converted second-band sound data obtained by converting the first-band sound data 31 into the second band switches from higher to lower than the accuracy of the feature derived using the second-band sound data 32 or the sound data 30. The first amplitude T1 may be changeable as needed within a range that satisfies the above-described condition, for example, by a user operating the input unit 13.
[0043] Returning to FIG. 1, the explanation will be continued.
[0044] If the conversion unit 20D determines that the amplitude of the first band sound data 31 in the processing unit time is equal to or greater than the first amplitude T1, it converts the first band sound data 31 extracted from the sound data 30 into converted second band sound data of the second band for that processing unit time.
[0045] 5A to 5E are explanatory diagrams of an example of conversion processing by the conversion unit 20D. The conversion unit 20D converts first-band sound data 31 of the first band into converted second-band sound data 33 of the second band using the conversion processing method shown in Figs. 5A to 5E, etc. Figs. 5A to 5E are diagrams that schematically show sound data 30, first-band sound data 31, second-band sound data 32, and converted second-band sound data 33, etc., represented by waveforms such as those shown in Fig. 3B, etc.
[0046] 5A to 5E, the explanation is based on the assumption that the audible sound band is 0 kHz to 20 kHz and the inaudible sound band is a frequency band exceeding 20 kHz. However, the audible sound band is not limited to the above numerical range, as long as it corresponds to the audible sound band from which features can be derived using the learning model 15A. For example, if the audible sound band from which features can be derived using the learning model 15A is 0 kHz to 16 kHz, the conversion unit 20D may execute the following conversion process by setting the audible sound band to 0 kHz to 16 kHz and the inaudible sound band to a frequency band exceeding 16 kHz. Furthermore, for example, if the audible sound band from which features can be derived using the learning model 15A is 0 kHz to 24 kHz, the conversion unit 20D may execute the following conversion process by setting the audible sound band to 0 kHz to 24 kHz and the inaudible sound band to a frequency band exceeding 24 kHz.
[0047] 5A is an explanatory diagram of an example of conversion processing by the conversion unit 20D. Fig. 5A shows a schematic diagram illustrating first-band sound data 31 before conversion and converted second-band sound data 33 after conversion to the second band in a processing unit time. In the graph of Fig. 5A, the horizontal axis represents time and the vertical axis represents frequency.
[0048] For example, the conversion unit 20D extracts sound pressure data of a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first band sound data 31, and moves the sound pressure data to the second band while maintaining the sound pressure distribution in the frequency band, thereby generating the converted second band sound data 33. The predetermined amplitude may be set in advance to a value greater than the second amplitude T2.
[0049] The conversion unit 20D generates converted second-band sound data 33 by moving sound pressure data of a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first-band sound data 31, to the second band, so that the converted second-band sound data 33 can be generated from sound pressure data of a characteristic frequency band included in the first-band sound data 31.
[0050] If frequency bands having amplitudes equal to or greater than a predetermined amplitude included in the first-band sound data 31 are dispersed in the frequency direction, the converter 20D arranges the sound pressure data of the dispersed frequency bands so that they are consecutive in the frequency direction. The converter 20D then compresses and moves the group of sound pressure data of the arranged frequency bands to the second-band frequency band while maintaining the sound pressure distribution in each of the consecutively arranged frequency bands, thereby generating the converted second-band sound data 33.
[0051] Furthermore, there may be cases where a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first band of the first-band sound data 31, is smaller than the width of the frequency band of the second band. In this case, the conversion unit 20D may generate converted second-band sound data 33 by expanding the sound pressure data of the frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first-band sound data 31, in the frequency direction to the frequency bandwidth of the two bands while maintaining the sound pressure distribution in the frequency band, and then moving the expanded data to the second band.
[0052] In this case, the conversion unit 20D may generate the converted second-band sound data 33 by moving sound pressure data of a frequency band having a predetermined amplitude or more, which is included in the first-band sound data 31, into the second band while maintaining the sound pressure distribution in the frequency band.
[0053] Furthermore, the bandwidth of a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first-band sound data 31, may exceed the frequency bandwidth of the second band. In this case, the conversion unit 20D may compress the sound pressure data of the frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first-band sound data 31, in the frequency direction to the frequency bandwidth of the two bands while maintaining the sound pressure distribution in the frequency band, and then move the compressed data to the second band, thereby generating the converted second-band sound data 33.
[0054] Fig. 5B is an explanatory diagram of an example of conversion processing by the conversion unit 20D. Fig. 5B shows a schematic diagram illustrating each of the sound data 30 and the converted second-band sound data 33 in a processing unit time. In the graph of Fig. 5B, the horizontal axis represents time and the vertical axis represents frequency.
[0055] For example, the conversion unit 20D divides the first-band sound data 31 included in the sound data 30 into a plurality of divided first-band sound data 34 in the frequency direction. The frequency bandwidth of each of the divided first-band sound data 34 is preferably equal to or less than the width of the second band of the second-band sound data 32. FIG. 5B shows an example in which the second band, which is the audible sound band, has a frequency bandwidth of 20 kHz from 0 kHz to 20 kHz, and each of the plurality of divided first-band sound data 34 has a frequency bandwidth of 20 kHz. FIG. 5B also shows an example in which the first-band sound data 31 is divided into divided first-band sound data 34A to 34D. The divided first-band sound data 34A to 34D are examples of the divided first-band sound data 34.
[0056] The conversion unit 20D converts the first-band sound data 31 into converted second-band sound data 33 by adding and arranging the multiple divided first-band sound data 34 in the second band while maintaining the sound pressure distribution in the frequency direction of each of the divided first-band sound data 34. In this example, the conversion unit 20D does not use the second-band sound data 32 included in the sound data 30, but uses only the first-band sound data 31 included in the sound data 30, moves the multiple divided first-band sound data 34 obtained by dividing the first-band sound data 31 to the second band, and adds up the sound pressure data of each, thereby generating the converted second-band sound data 33.
[0057] 5C is an explanatory diagram of an example of conversion processing by the conversion unit 20D. Fig. 5C shows schematic diagrams of the first-band sound data 31, the converted second-band sound data 33', and the converted second-band sound data 33 in a processing unit time. In the graph of Fig. 5C, the horizontal axis represents frequency and the vertical axis represents amplitude.
[0058] For example, the conversion unit 20D converts the first-band sound data 31 represented by the spectral waveform shown in FIG. 5C into converted second-band sound data 33′ by undersampling the first-band sound data 31 into the converted second-band sound data 33′. In undersampling, the conversion unit 20D uses the frequency bandwidth of the second band (e.g., 20 kHz) as a basic unit and divides the first-band sound data 31 into frequency unit widths, which are the frequency bandwidths of the basic units. The conversion unit 20D then places the spectrum of sound pressure data of the lowest frequency unit width among the multiple frequency unit widths included in the first-band sound data 31 in the second band, and superimposes the spectra of sound pressure data of frequency unit widths higher than the frequency unit width on the spectrum of the sound pressure data of the frequency unit width by alternately inverting the spectra in the frequency direction and folding them over. Through these processes, the converted second-band sound data 33′ after undersampling becomes data in which the first-band sound data 31 of the first band is aggregated into the second band.
[0059] The conversion unit 20D may use converted second-band sound data 33' generated by undersampling as the converted second-band sound data 33. Note that, from the viewpoint of achieving higher accuracy, the conversion unit 20D may use converted second-band sound data 33 obtained by spectrally inverting the spectrum represented by the converted second-band sound data 33' within the second band as the converted second-band sound data 33, as shown in Fig. 5C. Achieving higher accuracy means improving compatibility with the learning model 15A.
[0060] The conversion unit 20D generates the converted second-band sound data 33 by the above-described undersampling, so that the converted second-band sound data 33 that expresses frequency characteristics closer to the second-band sound data 32 extracted from the sound data 30 can be generated from the first-band sound data 31.
[0061] Fig. 5D is an explanatory diagram of an example of conversion processing by the conversion unit 20D. Fig. 5D shows a schematic diagram illustrating each of the sound data 30 and the converted second-band sound data 33 in a processing unit time. In the graph shown in Fig. 5D, the horizontal axis represents time and the vertical axis represents frequency.
[0062] For example, the conversion unit 20D compresses the first-band sound data 31 included in the sound data 30 in the frequency direction to a frequency bandwidth of the second band or less and moves the first-band sound data 31 to the second-band sound data 32. In detail, the conversion unit 20D reduces the density of the sound pressure distribution of the sound pressure data of the first-band sound data 31, compresses the first-band sound data 31 in the frequency direction to the frequency bandwidth of the second band, and moves the first-band sound data 31 to the second band. In this example, the conversion unit 20D does not use the second-band sound data 32 included in the sound data 30, but uses only the first-band sound data 31 included in the sound data 30, and compresses the first-band sound data 31 to the second band, thereby generating converted second-band sound data 33.
[0063] In this case, the conversion unit 20D can generate converted second-band sound data 33 in which the spectral features representing the first-band sound data 31 of the first band are compressed in the frequency direction, and therefore can generate converted second-band sound data 33 that is suitable for deriving the feature amount.
[0064] Fig. 5E is an explanatory diagram of an example of conversion processing by the conversion unit 20D. Fig. 5E shows a schematic diagram illustrating each of the sound data 30 and the converted second-band sound data 33 in a processing unit time. In the graph shown in Fig. 5E, the horizontal axis represents time and the vertical axis represents frequency.
[0065] For example, the conversion unit 20D compresses the sound data 30 in the frequency direction to a frequency bandwidth of the second band or less and moves the sound data 30 to the second band, thereby converting the first-band sound data 31 into second-band sound data 32. In particular, the conversion unit 20D reduces the density of the sound pressure distribution of the sound pressure data constituting the sound data 30, and then compresses the sound data in the frequency direction to the frequency bandwidth of the second band. That is, in this example, the conversion unit 20D generates converted second-band sound data 33 by compressing both the first-band sound data 31 and the second-band sound data 32 included in the sound data 30 into the second band.
[0066] In this case, the conversion unit 20D can generate converted second-band sound data 33 in which the spectral features representing the sound data 30 in the first and second frequency bands are compressed, and thus can generate converted second-band sound data 33 that is suitable for deriving the features.
[0067] Returning to FIG. 1, the explanation will be continued.
[0068] The conversion unit 20D converts the first-band sound data 31 into converted second-band sound data 33 of the second band by any of the conversion processes described above. That is, when the conversion unit 20D determines that the amplitude of the first-band sound data 31 in a processing unit time is equal to or greater than the first amplitude T1, it converts the first-band sound data 31 into converted second-band sound data 33 for that processing unit time.
[0069] Then, the conversion unit 20D converts the first band sound data 31 into the second band sound data 33, and outputs the converted second band sound data 33 to the feature derivation unit 20E as data used to derive a feature.
[0070] When the feature derivation unit 20E receives the converted second-band sound data 33 from the conversion unit 20D, the feature derivation unit 20E inputs the converted second-band sound data 33 to the learning model 15A.
[0071] The learning model 15A is a learning model that receives second-band data, which is sound data in a second band, as input and outputs feature quantities of the second-band data. The feature quantities are information that represent the characteristics of the sound and are expressed, for example, as a multidimensional vector. The learning model 15A may be any learning model that receives at least second-band sound data as input and outputs feature quantities of the second-band sound data. The learning model 15A is constructed, for example, by machine learning using one or more data sets, each consisting of a pair of second-band data and a feature quantity, as training data. The learning model 15A may be any known machine learning model and is not limited to any particular model. For example, the learning model 15A may be a convolutional neural network (CNN), a recurrent neural network (RNN) such as a long short term memory (LSTM), or an attention mechanism.
[0072] The learning model 15A is trained in advance and stored in advance in the storage unit 15. Note that the learning model 15A may be a commercially available or publicly available learning model such as OpenL3, which receives second band data as input and outputs feature quantities of the second band data.
[0073] When the feature derivation unit 20E receives the converted second-band sound data 33 from the conversion unit 20D, the feature derivation unit 20E inputs the converted second-band sound data 33 to the learning model 15A, thereby deriving the first feature, which is the feature output from the learning model 15A, as a feature of the sound data 30.
[0074] That is, when the amplitude of the first-band sound data 31 included in the sound data 30 is equal to or greater than the first amplitude T1, the feature derivation unit 20E inputs converted second-band sound data 33 obtained by converting the first-band sound data 31 into the second band to the learning model 15A. Then, the feature derivation unit 20E derives a first feature, which is a feature of the converted second-band sound data 33, as an output from the learning model 15A.
[0075] On the other hand, it is assumed that the determining unit 20C determines that the amplitude of the first band sound data 31 is equal to or greater than the second amplitude T2 but less than the first amplitude T1.
[0076] When it is determined that the amplitude of the first band sound data 31 included in the sound data 30 is less than the first amplitude T1, the feature derivation unit 20E inputs the sound data 30 to the learning model 15A. Then, the feature derivation unit 20E derives a second feature as a feature output from the learning model 15A.
[0077] Furthermore, when it is determined that the amplitude of the first-band sound data 31 included in the sound data 30 is less than the first amplitude T1, the feature derivation unit 20E may input the second-band sound data 32 included in the sound data 30 to the learning model 15A. Then, the feature derivation unit 20E may derive a third feature as a feature output from the learning model 15A.
[0078] 4 , for example, among sections s2, s4, s6, s8, s10, and s12 that are sound detection sections with the second amplitude T2 or more among sections s1 to s12 that constitute the entire period of the first band sound data 31, the amplitudes of sections s2, s4, s6, s8, and s12 other than section s10 are equal to or greater than the first amplitude T1. Therefore, by performing the above-described processing for each processing unit time by the determination unit 20C, the conversion unit 20D, and the feature derivation unit 20E, the first feature output by inputting converted second band sound data 33 obtained by converting the first band sound data 31 into the second band into the learning model 15A is derived as a feature for these sections s2, s4, s6, s8, and s12 in the sound data 30.
[0079] 4, among the sound detection sections of sections s1 to s12 constituting the entire period in which the sound is detected at an amplitude equal to or greater than the second amplitude T2, the section with an amplitude less than the first amplitude T1 is section s10. Therefore, by performing the above-described processes for each processing unit time by the determination unit 20C, the conversion unit 20D, and the feature derivation unit 20E, the second feature or the third feature output by inputting the sound data 30 or the second band sound data 32 into the learning model 15A is derived as a feature for section s10 in the sound data 30.
[0080] The judgment unit 20C may not judge whether the amplitude of the first band sound data 31 is equal to or greater than the second amplitude T2, but may judge whether the amplitude of the first band sound data 31 is equal to the first amplitude T1 for each processing unit time.
[0081] In this case, by executing the above-mentioned processes for each processing unit time by the determination unit 20C, the conversion unit 20D, and the feature derivation unit 20E, for sections s2, s4, s6, s8, and s12 in the first band sound data 31 that are equal to or greater than the first amplitude T1, the converted second band sound data 33 obtained by converting the first band sound data 31 into the second band is input to the learning model 15A, and the first feature output is derived as a feature, in the same manner as described above.
[0082] In this case, by executing the above-described processing for each processing unit time by the determination unit 20C, the conversion unit 20D, and the feature derivation unit 20E, for the sections s1, s3, s5, s7, s9, s10, and s11 in the first band sound data 31 that are less than the first amplitude T1, the second feature or the third feature output by inputting the sound data 30 or the second band sound data 32 into the learning model 15A is derived as a feature.
[0083] It may be specified in advance that periodic sound generation intervals are included in the sound data 30. For example, if the sound collection environment of the sound data 30 is an environment in which a sound source that generates periodic sounds is located, it is assumed in advance that the sound data 30 will include periodic sound generation intervals.
[0084] This will be explained using Fig. 6. Fig. 6 is a schematic diagram of an example of the first band sound data 31. Fig. 6 is a graph that schematically shows the first band sound data 31 represented by a waveform such as that shown in Fig. 3B etc. In Fig. 6, the horizontal axis represents time and the vertical axis represents amplitude.
[0085] For example, assume that the first band sound data 31 included in the sound data 30 acquired by the acquisition unit 20A is represented by the graph shown in Fig. 6. Also assume that the sound data 30 is specified in advance to include periodic sound generation intervals, that is, intervals s2, s4, s6, s8, s10, s12, s14, and s16.
[0086] Similarly to the above, the determination unit 20C determines for each processing unit time whether the amplitude of the first-band sound data 31 is equal to or greater than the second amplitude T2. Then, the determination unit 20C identifies a section of the processing unit time in which it is determined that the amplitude of the first-band sound data 31 is equal to or greater than the second amplitude T2 as a sound detection section in which some sound has been detected. In the example shown in FIG. 6 , the determination unit 20C identifies sections s2, s4, s6, s8, s10, s14, and s16, which have amplitudes equal to or greater than the second amplitude T2, among sections s1 to s16 that constitute the entire period of the first-band sound data 31, as sound detection sections.
[0087] For the sound detection section, the determination unit 20C, the conversion unit 20D, and the feature amount derivation unit 20E may perform the same processing as above.
[0088] Then, for a section s12, which is a sound generation section that does not overlap with the detected sound detection section, among the pre-specified periodic sound generation sections represented by the sound data 30, the feature derivation unit 20E inputs the sound data 30 of the sound generation section to the learning model. Then, for the section s12, the feature derivation unit 20E derives a second feature, which is a feature of the sound data 30, as an output from the learning model 15A.
[0089] Note that the feature derivation unit 20E may input, to the learning model, second-band sound data 32 included in the sound data 30 of a sound generation section s12, which is a sound generation section that does not overlap with a detected sound detection section among pre-specified periodic sound generation sections represented by the sound data 30. Then, for the section s12, the feature derivation unit 20E may derive a third feature, which is a feature of the second-band sound data 32, as an output from the learning model 15A.
[0090] Returning to FIG. 1, the explanation will be continued.
[0091] The sound source type determination unit 20F determines the sound source type of the sound data 30 from a first feature amount, which is a feature amount output by inputting the converted second band sound data 33 to the learning model 15A. That is, for a processing unit time in which the amplitude of the first band sound data 31 is equal to or greater than the first amplitude T1, the sound source type determination unit 20F determines the sound source type of the sound data 30 using the first feature amount obtained by inputting the converted second band sound data 33, which is obtained by converting the first band sound data 31 of the processing unit time into the second band, to the learning model 15A.
[0092] The type of sound source represents the type of medium or state of the medium that is the source of the generated sound. The medium may be anything that generates sound. For example, the medium may be various types of driven equipment or devices. The type of medium state represents the type of driving state of the equipment or device. For example, if the medium is a device that cuts and stamps, the type of medium state may be sounds related to operation, such as cutting sounds, stamping sounds, etc.
[0093] Furthermore, for a processing unit time in which the amplitude of the first band sound data 31 is less than the first amplitude T1, the sound source type determination unit 20F uses a second feature amount obtained by inputting the sound data 30 for that processing unit time into the learning model 15A to determine the sound source type of the sound data 30. Furthermore, for a processing unit time in which the amplitude of the first band sound data 31 is less than the first amplitude T1, the sound source type determination unit 20F may use a third feature amount obtained by inputting the second band sound data 32 for that processing unit time into the learning model 15A to determine the sound source type of the sound data 30.
[0094] The sound source type determination unit 20F determines the sound source type of the sound data 30 based on the similarity between the feature output from the learning model 15A and the reference feature, which is the feature of the generated sound data of the sound source type.
[0095] More specifically, the sound source type determining unit 20F determines the similarity using the feature amount management DB 15B.
[0096] 7 is a schematic diagram showing an example of the data structure of the feature management DB 15B. The feature management DB 15B is a database in which sound source types are associated with reference features. The data format of the feature management DB 15B is not limited to a database.
[0097] The reference features are features output from the learning model 15A by inputting generated sound data of a sound source type to the learning model 15A. The reference features include a first reference feature, a second reference feature, and a third reference feature. The generated sound data is sound data collected by the sound collection unit 11 from a medium that is the sound source of the sound source type in a state specified by the sound source type.
[0098] The first reference feature is a reference feature output from the learning model 15A by extracting first-band generated sound data of the first band from the generated sound data of the sound source type by the extraction unit 20B, converting it to the second band by the conversion unit 20D, and then inputting it into the learning model 15A.
[0099] The second reference feature is a reference feature that is output from the learning model 15A by inputting generated sound data of a sound source type to the learning model 15A.
[0100] The third reference feature is a reference feature output from the learning model 15A by extracting second-band generated sound data of the second band from the generated sound data of the sound source type by the extraction unit 20B and inputting the data into the learning model 15A.
[0101] The processing unit 20 may generate reference features in advance for each of a plurality of different sound source types using the generated sound data of each sound source type and the learning model 15A, and store the reference features in advance in association with the feature management DB 15B.
[0102] Then, for a processing unit time in which the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1, the sound source type determination unit 20F calculates a similarity between a first feature obtained by inputting converted second-band sound data 33 obtained by converting the first-band sound data 31 of the processing unit time into a second band into the learning model 15A, and the first reference feature registered in the feature management DB 15B. For example, cosine similarity may be used as the similarity. That is, the sound source type determination unit 20F calculates a cosine similarity between the first feature represented by a multidimensional vector and each of the multiple first reference feature represented by a multidimensional vector registered in the feature management DB 15B.
[0103] Then, the sound source type determination unit 20F determines the sound source type of the first reference feature whose similarity to the first feature is equal to or greater than a predetermined similarity as the sound source type of the sound data 30 in the processing unit time. A threshold value capable of determining the sound source type may be set in advance as the predetermined similarity. The predetermined similarity may be changeable as appropriate by a user's operation instruction via the input unit 13, etc. Furthermore, when there are multiple first reference features whose similarity to the first feature is equal to or greater than a predetermined similarity, the sound source type determination unit 20F may determine the sound source type of the first reference feature with the highest similarity as the sound source type of the sound data 30 in the processing unit time.
[0104] For a processing unit time in which the amplitude of the first band sound data 31 is less than the first amplitude T1, the sound source type determination unit 20F calculates the similarity between the second feature obtained by inputting the sound data 30 for that processing unit time into the learning model 15A and the second reference feature registered in the feature management DB 15B. That is, the sound source type determination unit 20F calculates the cosine similarity between the second feature represented by a multidimensional vector and each of the multiple second reference feature represented by a multidimensional vector registered in the feature management DB 15B.
[0105] Then, the sound source type determination unit 20F determines the sound source type of the second reference feature whose similarity to the second feature is equal to or greater than a predetermined similarity as the sound source type of the sound data 30 in that processing unit time. Alternatively, the sound source type determination unit 20F may determine the sound source type of the second reference feature whose similarity to the second feature is highest as the sound source type of the sound data 30 in that processing unit time. Alternatively, if no second reference feature whose similarity to the second feature is equal to or greater than a predetermined similarity is registered in the feature management DB 15B, the sound source type determination unit 20F may determine the sound source type of the sound data 30 in that processing time to be "none."
[0106] Furthermore, for a processing unit time in which the amplitude of the first band sound data 31 is less than the first amplitude T1, the sound source type determination unit 20F may calculate the similarity between a third feature obtained by inputting the second band sound data 32 included in the sound data 30 for that processing unit time into the learning model 15A and a third reference feature registered in the feature management DB 15B. That is, the sound source type determination unit 20F may calculate the cosine similarity between the third feature represented by a multidimensional vector and each of a plurality of third reference features represented by multidimensional vectors registered in the feature management DB 15B.
[0107] In this case, the sound source type determination unit 20F determines the sound source type of the third reference feature whose similarity to the third feature is equal to or greater than a predetermined similarity as the sound source type of the sound data 30 in that processing unit time. Alternatively, the sound source type determination unit 20F may determine the sound source type of the third reference feature whose similarity to the third feature is highest as the sound source type of the sound data 30 in that processing unit time. Alternatively, if no third reference feature whose similarity to the third feature is equal to or greater than a predetermined similarity is registered in the feature management DB 15B, the sound source type determination unit 20F may determine the sound source type of the sound data 30 in that processing time to be "none."
[0108] In addition, for a processing unit time in which the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1, the similarity between the first feature obtained by inputting the converted second-band sound data 33 obtained by converting the first-band sound data 31 of the processing unit time into the second band into the learning model 15A and the first reference feature may be less than the predetermined similarity. That is, all of the similarities between the first feature and the first reference feature of each of the multiple sound source types registered in the feature management DB 15B may be less than the predetermined similarity.
[0109] In this case, the sound source type determination unit 20F determines the sound source type of the sound data 30 in the same manner as described above by calculating the similarity between the second feature output by inputting the sound data 30 to the learning model 15A or the third feature output by inputting the second band sound data 32 included in the sound data 30 to the learning model 15A, and the first reference feature or the second reference feature, for the processing unit time.
[0110] That is, if the similarity between the first feature derived from the converted second band sound data 33 obtained by converting the first band sound data 31 into the second band and the reference feature is less than a predetermined similarity, the sound source type determination unit 20F determines the sound source type using the feature (second feature or third feature) derived using the sound data 30 or the second band sound data 32.
[0111] Through these processes, the sound source type determination unit 20F can determine the sound source type based on the converted second-band sound data 33 obtained by converting the first-band sound data 31, which is less susceptible to noise and more susceptible to attenuation, into the second band for a section where the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1. Furthermore, for a section where the accuracy of determining the sound source type may decrease when using the first-band sound data 31, such as when the amplitude of the first-band sound data 31 is less than the first amplitude T1, the sound source type determination unit 20F can determine the sound source type based on the sound data 30 or the second-band sound data 32. Furthermore, when the accuracy of determining the similarity based on the converted second-band sound data 33 is low, the sound source type determination unit 20F can determine the sound source type based on the sound data 30 or the second-band sound data 32.
[0112] Returning to FIG. 1, the explanation will be continued.
[0113] The output control unit 20G outputs the determination result by the sound source type determination unit 20F. The determination result is, for example, information that indicates the sound source type determined for each processing unit time in the entire section of the sound data 30 for each processing unit time.
[0114] The output control unit 20G outputs the determination result to the output unit 14. For example, the output control unit 20G displays an image representing the determination result on the display of the output unit 14. The output control unit 20G also causes a speaker in the storage unit 15 to output a sound representing the determination result. The output control unit 20G may also transmit the determination result to an external information processing device via the communication unit 12. The output control unit 20G may also store the determination result in the storage unit 15.
[0115] Next, an example of the flow of information processing executed by the information processing device 10 of this embodiment will be described.
[0116] FIG. 8 is a flowchart showing an example of the flow of information processing executed by the information processing device 10 of this embodiment.
[0117] The acquisition unit 20A acquires the sound data 30 collected by the sound collection unit 11 (step S100).
[0118] The extraction unit 20B extracts the first band sound data 31 from the sound data 30 acquired in step S100 (step S102).
[0119] Then, the processing unit 20 executes the processes of steps S104 to S128 for each processing unit time for the sound data 30 acquired in step S100.
[0120] Specifically, the determination unit 20C determines whether the amplitude of the first band sound data 31 of the processing unit time to be processed in the first band sound data 31 extracted in step S102 is equal to or greater than the second amplitude T2 (step S104). If the determination in step S104 is affirmative (step S104: Yes), the determination unit 20C determines that the processing unit time is a sound detection section, and proceeds to step S106.
[0121] In step S106, the determination unit 20C determines whether the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1 (step S106). If it is determined that the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1 (step S106: Yes), the process proceeds to step S108.
[0122] In step S108, the conversion unit 20D converts the first-band sound data 31 extracted from the sound data 30 for the processing unit time into converted second-band sound data 33 (step S108).
[0123] The feature derivation unit 20E inputs the converted second band sound data 33 generated in step S108 into the learning model 15A, thereby deriving the first feature output from the learning model 15A (step S110).
[0124] The sound source type determination unit 20F calculates the similarity between the first feature derived in step S110 and the first reference feature registered in the feature management DB 15B (step S112).
[0125] The sound source type determination unit 20F determines whether the similarity calculated in step S112 is equal to or greater than a predetermined similarity (step S114). If the determination in step S114 is negative (step S114: No), the process proceeds to step S118. If the determination in step S114 is positive (step S114: Yes), the process proceeds to step S116.
[0126] In step S116, the sound source type determination unit 20F determines the sound source type of the first reference feature whose similarity to the first feature is equal to or greater than a predetermined similarity as the sound source type of the sound data 30 in the processing unit time (step S116). Furthermore, when there are multiple first reference features whose similarity to the first feature is equal to or greater than a predetermined similarity, the sound source type determination unit 20F may determine the sound source type of the first reference feature with the highest similarity as the sound source type of the sound data 30 in the processing unit time.
[0127] On the other hand, if the determination in step S106 is negative (step S106: No), the process proceeds to step S118. In step S118, the extraction unit 20B extracts second-band sound data 32 from the sound data 30 of the processing unit time to be processed (step S118). Then, the feature derivation unit 20E inputs the second-band sound data 32 extracted in step S118 into the learning model 15A, thereby deriving a third feature output from the learning model 15A (step S120). Then, the sound source type determination unit 20F calculates the similarity between the third feature derived in step S120 and a third reference feature registered in the feature management DB 15B (step S122). The sound source type determining unit 20F determines the sound source type of the second reference feature having the highest similarity to the second feature calculated in step S122 as the sound source type of the sound data 30 in the processing unit time (step S124).
[0128] As described above, the processing of steps S118 to S124 may be performed using the second feature amount output from the learning model 15A by inputting the sound data 30 into the learning model 15A.
[0129] On the other hand, if the determination in step S104 is negative (step S104: No), the process proceeds to step S126. In step S126, the determination unit 20C determines whether the processing unit time to be processed is within a predetermined periodic sound generation interval of the sound data 30 (step S126). If the determination in step S126 is positive (step S126: Yes), the process proceeds to step S118. If the determination in step S126 is negative (step S126: No), the process proceeds to step S128.
[0130] In step S128, the sound source type determining unit 20F determines that there is no sound source type for the processing unit time to be processed (step S128).
[0131] The processing unit 20 executes the processes of steps S104 to S128 for each processing unit time for the sound data 30 acquired in step S100, thereby determining the type of sound source for each processing unit time.
[0132] The output control unit 20G outputs a determination result indicating the type of sound source determined for each processing unit time (step S130), and ends this routine.
[0133] In the above routine, an example has been described in which the conversion unit 20D converts the first-band sound data 31 into the second-band sound data 32 at the timing of step S108. However, the conversion unit 20D may convert the first-band sound data 31 for the entire period included in the sound data 30 into converted second-band sound data 33 at the timing between step S102 and step S104, i.e., before the processing for each processing unit time is executed. In this case, in step S108, the conversion unit 20D may output the converted second-band sound data 33 for the processing unit time to be processed, which has been converted in advance, to the feature derivation unit 20E.
[0134] As described above, the information processing device 10 of this embodiment acquires sound data 30 and converts first-band sound data 31 of a first band including an inaudible sound band contained in the sound data 30 into converted second-band sound data 33 of a second band including an audible sound band. The information processing device 10 inputs the converted second-band sound data 33 to a learning model 15A that receives second-band data of the second band as input and outputs feature quantities of the second-band data.
[0135] Here, sound data in the inaudible sound band is used to derive highly accurate features because it does not contain audible noise and is more susceptible to attenuation than sound data in the audible sound band. However, because the sampling frequency of the inaudible sound band is higher than that of the audible sound band, the learning load of a learning model that inputs sound data in the inaudible sound band is large. For this reason, in conventional technologies that require a learning model that inputs sound data in the inaudible sound band, the learning load of the learning model is large, making it difficult to easily derive highly accurate features of sound data.
[0136] On the other hand, in the information processing device 10 of this embodiment, the first band sound data 31 included in the sound data 30 is converted into converted second band sound data 33 of a second band including the audible sound band, and the converted second band sound data 33 is input to a learning model 15A that outputs the feature quantities of the second band data of the second band.
[0137] In this way, in the information processing device 10 of this embodiment, by using the converted second band sound data 33, it is possible to derive highly accurate feature quantities based on the first band sound data 31 converted by the converted second band sound data 33. Furthermore, in the information processing device 10 of this embodiment, it is possible to easily derive feature quantities of the sound data 30 by using the learning model 15A that inputs the converted second band sound data 33 and second band data of the second band including the audible band, without separately preparing a learning model that inputs sound data of the inaudible band.
[0138] Therefore, the information processing device 10 of this embodiment can easily derive highly accurate feature quantities of the sound data 30.
[0139] Furthermore, the information processing apparatus 10 of this embodiment determines the sound source type of the sound data 30 from the first feature amount, which is the feature amount output by inputting the converted second band sound data 33 to the learning model 15A.
[0140] Therefore, in addition to the above-mentioned effects, the information processing device 10 of this embodiment can easily and accurately determine the sound source type of the sound data 30 .
[0141] Furthermore, the information processing device 10 of this embodiment determines the sound source type of the sound data 30 based on the similarity between the first feature and a reference feature, which is a feature of generated sound data of the sound source type. The information processing device 10 of this embodiment determines the sound source type of the sound data 30 using the similarity between features represented by, for example, multidimensional vectors, and therefore can determine the sound source type of the sound data 30 with high accuracy.
[0142] Furthermore, it is preferable that the first band is an inaudible band and the second band is an audible band. Note that the frequency bandwidth of the second band to be processed by the learning model 15A may be a frequency bandwidth of a frequency band that includes at least the audible band, and may be either completely identical to or mismatched with the frequency bandwidth of the converted second-band sound data 33. Specifically, it is preferable that the frequency bandwidth of the converted second-band sound data 33 of the second band is equal to or smaller than the frequency bandwidth of the second band to be processed by the learning model 15A.
[0143] Furthermore, in the information processing device 10 of this embodiment, if the amplitude of first-band sound data 31 included in sound data 30 is equal to or greater than the first amplitude T1, converted second-band sound data 33 obtained by converting the first-band sound data 31 is input to the learning model 15A, and if the amplitude of the first-band sound data 31 included in sound data 30 is less than the first amplitude T1, the information processing device 10 inputs the sound data 30 or second-band sound data 32 included in the sound data 30 to the learning model 15A.
[0144] Therefore, in the information processing device 10 of this embodiment, when it is possible to derive highly accurate features from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is equal to or greater than the first amplitude T1, the converted second-band sound data 33 obtained by converting the first-band sound data 31 into the second band can be used to derive the features. Furthermore, in the information processing device 10 of this embodiment, when it is difficult to derive highly accurate features from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is less than the first amplitude T1, the sound data 30 or the second-band sound data 32 can be used to derive the features. Therefore, in addition to the above-described effects, the information processing device 10 of this embodiment can derive features of the sound data 30 with even higher accuracy.
[0145] Furthermore, when the similarity between the first feature and the reference feature is less than a predetermined similarity, the information processing device 10 of this embodiment determines the sound source type of the sound data 30 based on the similarity between the reference feature and the second feature or the third feature, which is a feature output by inputting the sound data 30 or the second band sound data 32 included in the sound data 30 into the learning model 15A.
[0146] For this reason, in the information processing device 10 of this embodiment, when the similarity between the first feature and the reference feature is equal to or greater than a predetermined similarity, the similarity can be used to determine the sound source type of the sound data 30. Furthermore, in the information processing device 10 of this embodiment, when it is difficult to calculate a high-accuracy similarity from the converted second-band sound data 33, such as when the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data 30 can be determined using the feature derived from the sound data 30 or the second-band sound data 32 using the learning model 15A. Therefore, in addition to the above effects, the information processing device 10 of this embodiment can determine the sound source type of the sound data 30 with even higher accuracy.
[0147] Furthermore, the information processing device 10 of this embodiment identifies, based on the first band sound data 31, a section whose amplitude is equal to or greater than the second amplitude T2 as a sound detection section included in the sound data 30, and inputs converted second band sound data 33 obtained by converting the first band sound data 31 of the identified sound detection section into a second band to the learning model 15A. Furthermore, the information processing device 10 of this embodiment inputs, for a sound generation section that does not overlap with a sound detection section among pre-specified periodic sound generation sections represented by the sound data 30, the sound data 30 of the sound generation section or second band sound data 32 of the second band included in the sound data 30 to the learning model 15A.
[0148] Therefore, for a sound generation section in which a sound having an amplitude equal to or greater than the second amplitude T2 is not detected from the first-band sound data 31 among the periodic sound generation sections, the feature derivation unit 20E can derive a feature from the sound data 30 or the second-band sound data 32 included in the sound data 30. Therefore, in addition to the above-described effects, the information processing device 10 of this embodiment can derive a feature from the sound data 30 with even higher accuracy.
[0149] (Variation 1) In the above embodiment, an example was described in which the feature quantities of the sound data 30 are derived using the learning model 15A, which receives second band data of the second band as input and outputs the feature quantities of the second band data.
[0150] However, the information processing device 10 may derive the sound source type determination result of the sound data 30 using a second learning model that receives second band data of the second band and outputs the sound source type determination result.
[0151] In this case, the information processing device 10 may perform machine learning using known transfer learning with the learning model 15A to create in advance a second learning model that inputs second-band data and outputs a sound source type determination result. The training data used in the transfer learning may include at least one of the following: one or more data sets each including a pair of sound data 30 and a sound source type determination result; one or more data sets each including a pair of converted second-band sound data 33 obtained by converting first-band sound data 31 included in the sound data 30 into the second band and a sound source type determination result; and one or more data sets each including a pair of second-band sound data 32 included in the sound data 30 and a sound source type determination result. The conversion unit 20D and the feature derivation unit 20E may then use the second learning model to derive the sound source type determination result for the sound data 30 from the sound data 30.
[0152] (Variation 2) In the above embodiment, the processing unit 20 of the information processing device 10 has been described as determining the type of sound source based on the second-band sound data 32 of the second band including the audible sound band included in the sound data 30, when determining that the amplitude of the first-band sound data 31 of the first band including the inaudible sound band included in the sound data 30 is less than the first amplitude T1, which is the amplitude value at which the accuracy of the feature decreases. That is, in the above embodiment, the type of sound source is determined based on the second-band sound data 32 of the second band including the audible sound band included in the sound data 30, when there is a possibility of degradation in the determination of the type of sound source using the converted second-band sound data 33 obtained by converting the first-band sound data 31 of the first band including the inaudible sound band into the second band.
[0153] However, conversely, when the processing unit 20 of the information processing device 10 of this embodiment determines that the amplitude of the second-band sound data 32 of the second band including the audible sound band included in the sound data 30 is less than the first amplitude T1, which is the amplitude value at which the accuracy of the feature decreases, the processing unit 20 may determine the type of sound source based on the first-band sound data 31 of the first band including the inaudible sound band included in the sound data 30. In other words, when there is a possibility of degradation in the determination of the type of sound source using the second-band sound data 32 of the second band including the audible sound band included in the sound data 30, the processing unit 20 may determine the type of sound source based on the converted second-band sound data 33 obtained by converting the first-band sound data 31 of the first band including the inaudible sound band into the second band.
[0154] In addition, in the above embodiment, the processing unit 20 of the information processing device 10, when deriving the first feature by inputting the converted second-band sound data 33 into the learning model 15A, determines the sound source type of the sound data 30 based on the similarity between the first feature and the first reference feature registered in the feature management DB 15B.
[0155] Furthermore, the processing unit 20 of the information processing device 10 has been described as deriving a third feature by inputting the sound data 30 or the second band sound data 32 included in the sound data 30 into the learning model 15A, and determining the sound source type of the sound data 30 based on the similarity between the third feature and a third reference feature registered in the feature management DB 15B.
[0156] Furthermore, the processing unit 20 of the information processing device 10 has been described as determining the sound source type of the sound data 30 based on the similarity between the second feature and the second reference feature registered in the feature management DB 15B when the second feature is derived by inputting the sound data 30 into the learning model 15A.
[0157] However, the processing unit 20 of the information processing device 10 of this embodiment may determine the sound source type of the sound data 30 by using any two of the similarity between the first feature and the first reference feature (first similarity), the similarity between the second feature and the second reference feature (second similarity), and the similarity between the third feature and the third reference feature (third similarity), or by using three of the first to third similarities. In this case, for example, the processing unit 20 may determine the sound source type of the sound data 30 by identifying, from the feature management DB 15B, the sound source type corresponding to the reference feature (first reference feature, second reference feature, or third reference feature) used to calculate the highest similarity among the multiple similarities used to determine the sound source type.
[0158] The present technology may also be configured as follows. (1) An information processing method executed by a computer, comprising: acquiring sound data; converting first-band sound data of a first band including an inaudible sound band included in the sound data into converted second-band sound data of a second band including an audible sound band; and inputting the converted second-band sound data to a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band sound data. In this way, the information processing method of this embodiment uses the converted second-band sound data 33, thereby enabling highly accurate derivation of feature quantities based on the first-band sound data 31 converted by the converted second-band sound data 33. Furthermore, the information processing method of this embodiment does not require a separate learning model that receives sound data of the inaudible sound band as input, but instead uses a learning model 15A that receives the converted second-band sound data 33 and second-band data of a second band including the audible sound band as input, thereby enabling easy derivation of feature quantities of the sound data 30. (2) The information processing method according to (1), in which the sound source type of the sound data is determined from a first feature, which is the feature output by inputting the converted second-band sound data to the learning model. Therefore, in addition to the above-described effect, the information processing method of this embodiment can easily and accurately determine the sound source type of the sound data 30. (3) The information processing method according to (1) or (2), in which the sound source type of the sound data is determined based on the similarity between the first feature and a reference feature, which is the feature of the generated sound data of the sound source type. Therefore, in addition to the above-described effect, the information processing method of this embodiment can easily and accurately determine the sound source type of the sound data 30. (4) The information processing method according to any one of (1) to (3), in which the first band is an inaudible sound band, and the second band is an audible sound band. Therefore, in addition to the above-described effect, the information processing method of this embodiment can easily derive the feature of the sound data 30 using a learning model 15A that inputs second-band data in the audible sound band.(5) The information processing method according to any one of (1) to (4), wherein, when the amplitude of the first-band sound data included in the sound data is equal to or greater than a first amplitude, the converted second-band sound data obtained by converting the first-band sound data is input to the learning model; and, when the amplitude of the first-band sound data included in the sound data is less than the first amplitude, the sound data is input to the learning model. Therefore, in the information processing method of this embodiment, when highly accurate features can be derived from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is equal to or greater than a first amplitude T1, the converted second-band sound data 33 obtained by converting the first-band sound data 31 to a second band can be used to derive the features. Furthermore, in the information processing method of this embodiment, when it is difficult to derive highly accurate features from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is less than the first amplitude T1, the sound data 30 can be used to derive the features. Therefore, in addition to the above-described effects, the information processing method of this embodiment can derive features of the sound data 30 with even greater accuracy. (6) The information processing method according to any one of (1) to (5), wherein, when the amplitude of the first-band sound data included in the sound data is equal to or greater than a first amplitude, the converted second-band sound data obtained by converting the first-band sound data is input to the learning model; and, when the amplitude of the first-band sound data included in the sound data is less than the first amplitude, the converted second-band sound data of the second band included in the sound data is input to the learning model. Therefore, in the information processing method of this embodiment, when highly accurate features can be derived from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is equal to or greater than a first amplitude T1, the converted second-band sound data 33 obtained by converting the first-band sound data 31 to a second band can be used to derive the features. Furthermore, in the information processing method of this embodiment, when it is difficult to derive highly accurate features from the first-band sound data 31, for example, when the amplitude of the first-band sound data 31 is less than the first amplitude T1, the second-band sound data 32 can be used to derive the features. Therefore, in addition to the above-mentioned effects, the information processing method of this embodiment can derive the feature amount of the sound data 30 with even higher accuracy.(7) The information processing method according to any one of (3) to (6), wherein, when the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data is determined based on the similarity between the second feature, which is the feature output by inputting the sound data to the learning model, and the reference feature. Therefore, in the information processing method of this embodiment, when the similarity between the first feature and the reference feature is equal to or greater than a predetermined similarity, the sound source type of the sound data 30 can be determined using the similarity. Furthermore, in the information processing method of this embodiment, when it is difficult to calculate a high-accuracy similarity from the converted second-band sound data 33, such as when the similarity between the first feature and the reference feature is less than the predetermined similarity, the sound source type of the sound data 30 can be determined using the feature derived from the sound data 30 using the learning model 15A. Therefore, in addition to the above-described effects, the information processing method of this embodiment can determine the sound source type of the sound data 30 with even greater accuracy. (8) The information processing method according to any one of (3) to (6), wherein, when the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data is determined based on the similarity between the reference feature and a third feature, which is the feature output by inputting second-band sound data of the second band included in the sound data to the learning model. In the information processing method of this embodiment, when it is difficult to calculate a high-accuracy similarity from the converted second-band sound data 33, such as when the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data 30 can be determined using a feature derived from the second-band sound data 32 using the learning model 15A. Therefore, in addition to the above-mentioned effects, the information processing method of this embodiment can determine the sound source type of the sound data 30 with even higher accuracy. (9) The information processing method according to any one of (1) to (8), further comprising: identifying a sound detection interval included in the sound data based on the first band sound data; inputting the converted second band sound data obtained by converting the first band sound data of the identified sound detection interval into the learning model; and inputting the sound data of a sound generation interval that does not overlap with the sound detection interval among pre-specified periodic sound generation intervals represented by the sound data into the learning model.Therefore, for a sound generation section in which a sound having an amplitude equal to or greater than the second amplitude T2 is not detected in the first band sound data 31 among the periodic sound generation sections, the feature derivation unit 20E can derive a feature from the sound data 30. Therefore, in addition to the above-described effect, the information processing method of the present embodiment can derive a feature of the sound data 30 with even higher accuracy. (10) The information processing method according to any one of (1) to (9), comprising: identifying a sound detection section included in the sound data based on the first band sound data; inputting the converted second band sound data obtained by converting the first band sound data of the detected sound detection section into the learning model; and inputting the second band sound data of the second band included in the sound data of the sound generation section into the learning model for a sound generation section that does not overlap with the detected sound detection section among pre-specified periodic sound generation sections represented by the sound data. Therefore, for sound generation sections in which a sound having an amplitude equal to or greater than the second amplitude T2 is not detected in the first-band sound data 31 among the periodic sound generation sections, the feature derivation unit 20E can derive features from the second-band sound data 32 included in the sound data 30. Therefore, in addition to the above-described effects, the information processing method of this embodiment can derive features of the sound data 30 with even higher accuracy. (11) The information processing method according to any one of (1) to (10), further comprising: extracting sound pressure data of a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first-band sound data, and converting the first-band sound data into the converted second-band sound data by shifting the extracted sound pressure data to the second band. Therefore, the information processing method of this embodiment can generate converted second-band sound data 33 that efficiently utilizes a frequency band in which the characteristics of inaudible sounds are prominent. Therefore, in addition to the above-described effects, the information processing method of this embodiment can derive features of the sound data 30 with even higher accuracy. (12) The information processing method according to any one of (1) to (11), wherein the first-band sound data is converted into the converted second-band sound data by adding a plurality of divided first-band sound data obtained by dividing the first-band sound data in a frequency direction to the second band. Therefore, the information processing method of this embodiment can use a wide range of components obtained by adding together a wide range of components in the inaudible sound band as the converted second-band sound data 33.Therefore, in addition to the above-mentioned effects, the information processing method of this embodiment can derive the feature quantities of the sound data 30 with even higher accuracy. (13) The information processing method according to any one of (1) to (12), in which the first-band sound data is converted into the converted second-band sound data by undersampling the first-band sound data to the second band. Therefore, the information processing method of this embodiment can use the converted second-band sound data 33 obtained by adding together components from a wide range of the inaudible sound band. Therefore, in addition to the above-mentioned effects, the information processing method of this embodiment can derive the feature quantities of the sound data 30 with even higher accuracy. (14) The information processing method according to any one of (1) to (13), in which the first-band sound data is converted into the converted second-band sound data by compressing the first-band sound data in the frequency direction to the frequency bandwidth of the second band and moving the data to the second band. Therefore, the information processing method of this embodiment can convert a wide range of components in the inaudible sound band into the converted second-band sound data 33 in a form compressed directly in the frequency direction, and therefore, a wide range of features of the spectrogram can be used as the converted second-band sound data 33. Therefore, the information processing method of this embodiment can derive the feature quantities of the sound data 30 with even higher accuracy. (15) The information processing method according to any one of (1) to (13), wherein the first-band sound data is converted into the converted second-band sound data by compressing the sound data in the frequency direction to the frequency bandwidth of the second band and moving the sound data to the second band. Therefore, the information processing method of this embodiment can convert a wide range of components in the inaudible sound band and the audible sound band in a form compressed directly in the frequency direction, and therefore, a wide range of features of the spectrogram can be used as the converted second-band sound data 33. Therefore, the feature quantities of the sound data 30 can be derived with even higher accuracy. Note that when the sound data 30 is compressed in the frequency direction to the frequency bandwidth of the second band and moved to the second band, noisy audible sound components may also be included, but because the sound data is compressed in the frequency direction, the influence of these noise components can be suppressed.(16) An information processing program causing a computer to execute the steps of: acquiring sound data; converting first-band sound data of a first band including an inaudible sound band included in the sound data into converted second-band sound data of a second band including an audible sound band; and inputting the converted second-band sound data into a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band sound data. As described above, in the information processing program of this embodiment, by using the converted second-band sound data 33, it is possible to derive highly accurate feature quantities based on the first-band sound data 31 converted by the converted second-band sound data 33. Furthermore, in the information processing program of this embodiment, it is possible to easily derive feature quantities of the sound data 30 by using the learning model 15A that receives the converted second-band sound data 33 and second-band data of the second band including the audible sound band as input, without separately preparing a learning model that receives sound data of the inaudible sound band as input. Therefore, the information processing program of this embodiment can easily derive highly accurate feature quantities of the sound data 30. (17) An information processing device comprising: an acquisition unit that acquires sound data; a conversion unit that converts first-band sound data of a first band including an inaudible band included in the sound data into converted second-band sound data of a second band including an audible band; and a feature derivation unit that inputs the converted second-band sound data to a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band sound data. In this way, the information processing device 10 of this embodiment uses the converted second-band sound data 33, thereby enabling highly accurate derivation of feature quantities based on the first-band sound data 31 converted by the converted second-band sound data 33. Furthermore, the information processing device 10 of this embodiment does not require a separate learning model that receives sound data of the inaudible band as input, but instead uses the learning model 15A that receives the converted second-band sound data 33 and second-band data of the second band including the audible band as input, thereby enabling easy derivation of feature quantities of the sound data 30. Therefore, the information processing device 10 of this embodiment can easily derive highly accurate feature quantities of the sound data 30.
[0159] Although the above describes an embodiment, the above embodiment is presented as an example and is not intended to limit the scope of the invention. The above novel embodiment can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the 20N invention. The above embodiment is included within the scope and spirit of the invention, and is also included in the invention and its equivalents as defined in the claims.
[0160] 10 Information processing device 15A Learning model 20A Acquisition unit 20D Conversion unit 20E Feature derivation unit 20F Sound source type determination unit
Claims
1. An information processing method executed by a computer, comprising: acquiring sound data; converting first-band sound data of a first band including an inaudible sound band contained in the sound data into converted second-band sound data of a second band including an audible sound band; and inputting the converted second-band sound data into a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band data.
2. The information processing method according to claim 1, wherein the sound source type of the sound data is determined from a first feature amount, which is the feature amount output by inputting the converted second-band sound data into the learning model.
3. The information processing method according to claim 2, wherein the sound source type of the sound data is determined based on the similarity between the first feature and a reference feature which is the feature of the generated sound data of the sound source type.
4. The information processing method according to claim 1, wherein the first band is an inaudible sound band, and the second band is an audible sound band.
5. The information processing method according to claim 1, wherein, when the amplitude of the first band sound data included in the sound data is equal to or greater than a first amplitude, the converted second band sound data obtained by converting the first band sound data is input to the learning model, and when the amplitude of the first band sound data included in the sound data is less than the first amplitude, the sound data is input to the learning model.
6. The information processing method according to claim 1, wherein, when the amplitude of the first band sound data included in the sound data is equal to or greater than a first amplitude, the converted second band sound data obtained by converting the first band sound data is input to the learning model, and when the amplitude of the first band sound data included in the sound data is less than the first amplitude, the second band sound data of the second band included in the sound data is input to the learning model.
7. An information processing method as described in claim 3, wherein, if the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data is determined based on the similarity between the second feature, which is the feature output by inputting the sound data into the learning model, and the reference feature.
8. The information processing method of claim 3, wherein, if the similarity between the first feature and the reference feature is less than a predetermined similarity, the sound source type of the sound data is determined based on the similarity between the reference feature and a third feature, which is the feature output by inputting second-band sound data of the second band included in the sound data to the learning model.
9. The information processing method according to claim 1, further comprising: identifying a sound detection interval included in the sound data based on the first band sound data; inputting the converted second band sound data obtained by converting the first band sound data of the identified sound detection interval into the learning model; and inputting the sound data of a sound occurrence interval that does not overlap with the sound detection interval among pre-specified periodic sound occurrence intervals represented by the sound data into the learning model.
10. The information processing method according to claim 1, further comprising: identifying a sound detection interval included in the sound data based on the first band sound data; inputting the converted second band sound data obtained by converting the first band sound data of the detected sound detection interval into the learning model; and inputting the second band sound data of the second band included in the sound data of the sound occurrence interval into the learning model for a sound occurrence interval that does not overlap with the detected sound detection interval among pre-specified periodic sound occurrence intervals represented by the sound data.
11. The information processing method according to claim 1, wherein sound pressure data of a frequency band having an amplitude equal to or greater than a predetermined amplitude, which is included in the first band sound data, is extracted, and the extracted sound pressure data is moved to the second band, thereby converting the first band sound data into the converted second band sound data.
12. The information processing method according to claim 1, wherein the first-band sound data is converted into the converted second-band sound data by adding a plurality of divided first-band sound data obtained by dividing the first-band sound data in a frequency direction to the second band.
13. The information processing method according to claim 1, further comprising converting the first-band sound data into the converted second-band sound data by undersampling the first-band sound data into the second band.
14. The information processing method according to claim 1, wherein the first-band sound data is converted into the converted second-band sound data by compressing the first-band sound data in a frequency direction to the frequency bandwidth of the second band and moving the data to the second band.
15. The information processing method according to claim 1, wherein the first-band sound data is converted into the converted second-band sound data by compressing the sound data in a frequency direction to the frequency bandwidth of the second band and moving the sound data to the second band.
16. An information processing program for causing a computer to execute the steps of: acquiring sound data; converting first-band sound data of a first band including an inaudible sound band contained in the sound data into converted second-band sound data of a second band including an audible sound band; and inputting the converted second-band sound data into a learning model that receives second-band data of the second band as input and outputs feature quantities of the second-band data.
17. An information processing device comprising: an acquisition unit that acquires sound data; a conversion unit that converts first-band sound data of a first band including an inaudible sound band contained in the sound data into converted second-band sound data of a second band including an audible sound band; and a feature derivation unit that inputs the converted second-band sound data to a learning model that receives second-band data of the second band as input and outputs features of the second-band data.
Citation Information
Patent Citations
Signal generation device and signal processor
JP2011215212A
Speech processing method, speech processing program, and speech processing device
WO2008015800A1
Learning model generation device, inference processing device, learning model generation method, and inference processing method
WO2024014318A1