Noise determination device, and noise determination program
By analyzing tone color data through Fourier and cepstrum transforms, and using a recurrent neural network, the device improves the accuracy of abnormal sound detection by considering timbre and other sound elements, aligning with human perception.
Patent Information
- Application Number
- JP2024000621
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-17
AI Technical Summary
Existing sound determination technologies, such as those described in Patent Document 1, may not accurately reflect human perception of abnormal sounds due to their reliance on pitch and loudness alone, failing to consider other sound elements like timbre, which can lead to inappropriate determination results.
The device employs a model that analyzes tone color data derived from Fourier-transformed sound data, using cepstrum analysis to extract sound pressure levels across a specific frequency range, and utilizes a recurrent neural network (LSTM) to determine the presence or absence of abnormal sounds based on these characteristics.
This approach provides sound determination results that align closer to human perception by considering timbre and other sound elements, enhancing the accuracy of abnormal sound detection.
Smart Images

Figure 2025106973000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a abnormal sound determination device and an abnormal sound determination program.
Background Art
[0002] The abnormal sound determination device disclosed in Patent Document 1 uses a pre-machine-learned model to determine whether an abnormal sound is included in the sound under inspection. The model takes, as input, data representing the temporal change of the sound pressure level for each frequency with respect to a sound over a certain period, and outputs information regarding the presence or absence of an abnormal sound.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The input data of the technology of Patent Document 1 corresponds to the frequency and the sound pressure level and reflects the pitch and loudness of the sound. Here, a person can perceive not only the pitch and loudness of the sound but also other elements related to the sound. Therefore, even if the pitch and loudness of the sound are the same, a person may recognize a certain sound as an abnormal sound and not recognize another sound as an abnormal sound. Therefore, from the viewpoint of obtaining a determination result close to a person's sense regarding the presence or absence of an abnormal sound, there is a possibility that the technology of Patent Document 1 may not necessarily obtain appropriate information.
Means for Solving the Problems
[0005] The abnormal sound determination device for solving the above problems includes an execution unit and a storage unit. The storage unit stores a model that outputs abnormal sound information, which is information regarding the presence or absence of abnormal sounds in the sound data, when tone color data indicating a waveform in a specific frequency range, which is a waveform of a frequency spectrum obtained by performing a Fourier transform on the sound data of the sound under inspection, is input. The execution unit executes a first process of acquiring the sound data in a unit period for the sound under inspection, a second process of acquiring waveform information of the frequency spectrum by further performing an inverse Fourier transform on a cepstrum obtained by applying cepstrum analysis to the sound data in the unit period to Fourier-transform the frequency spectrum of the sound data, a third process of extracting, as the tone color data, the sound pressure level for each frequency arranged in the order of frequency in the specific frequency range from the waveform information, and a fourth process of outputting the abnormal sound information by inputting the tone color data into the model.
[0006] The abnormal sound determination program for solving the above problems targets a computer including an execution unit and a storage unit. The storage unit stores a model that outputs abnormal sound information, which is information regarding the presence or absence of abnormal sounds in the sound data, when tone color data indicating a waveform in a specific frequency range, which is a waveform of a frequency spectrum obtained by performing a Fourier transform on the sound data of the sound under inspection, is input. The execution unit is caused to execute a first process of acquiring the sound data in a unit period for the sound under inspection, a second process of acquiring waveform information of the frequency spectrum by further performing an inverse Fourier transform on a cepstrum obtained by applying cepstrum analysis to the sound data in the unit period to Fourier-transform the frequency spectrum of the sound data, a third process of extracting, as the tone color data, the sound pressure level for each frequency arranged in the order of frequency in the specific frequency range from the waveform information, and a fourth process of outputting the abnormal sound information by inputting the tone color data into the model.
Advantages of the Invention
[0007] In each of the above technical ideas, regarding the presence or absence of abnormal sounds, information can be obtained based on criteria close to human perception.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0009] Hereinafter, an embodiment of an abnormal sound determination device and an abnormal sound determination program will be described with reference to the drawings. <Overall Configuration> As shown in FIG. 1, the abnormal sound determination system 100 includes an abnormal sound determination device 10, a user interface 20, and a sound collection device 30.
[0010] The abnormal sound determination device 10 is a computer including a CPU 12 and a memory 14. The CPU 12 is an execution unit. The memory 14 is a storage unit. The memory 14 includes three types: a RAM, a ROM, and an electrically rewritable non-volatile memory. In this embodiment, these three types are collectively referred to as the memory 14. The memory 14 stores in advance various programs in which the processing to be executed by the CPU 12 is described, and various data necessary for the CPU 12 to execute the programs. One of the various programs is an abnormal sound determination program W.
[0011] The user interface 20 is a touch panel display. That is, the user interface 20 serves both as a display device and as an input device for the user to input information. The user interface 20 is connected to the abnormal sound determination device 10 by wire or wirelessly. The user interface 20 displays a video according to the information output by the abnormal sound determination device 10. Also, the user interface 20 outputs information according to the user's input operation to the abnormal sound determination device 10.
[0012] The sound collection device 30 is a microphone or a sensor that detects the sound pressure V, which is the pressure of the sound emitted by the object. The sound collection device 30 is connected to the abnormal sound determination device 10 by wire or wirelessly. The sound collection device 30 repeatedly outputs information regarding the detected sound pressure V to the abnormal sound determination device 10.
[0013] <Abnormal sound determination process> Hereinafter, taking as an example the case of diagnosing the sound emitted by the transaxle mounted on a hybrid vehicle, the process performed by the abnormal sound determination device 10 will be described. Note that the diagnosis of the sound emitted by the transaxle is performed, for example, as part of the inspection of the transaxle before the vehicle is shipped from the factory. The transaxle includes an electric motor that serves as a drive source of the vehicle and a gear mechanism that transmits the rotation of the electric motor. That is, the transaxle includes rotating parts that transmit power by rotation. If the assembly position of such rotating parts is displaced from the original position or the machining state of the gear teeth is not good, vibration or poor meshing of the gear teeth may occur when each rotating part rotates. And accordingly, abnormal sound may occur. Abnormal sound is a sound that does not occur if the transaxle is normal. Abnormal sound can also be said to be a sound that is usually not noticeable from a human sense.
[0014] When inspecting a transaxle, first, the operator installs a sound collecting device 30 near the transaxle to be inspected. Then the operator operates the transaxle in this state. That is, the electric motor of the transaxle is driven. Then, hereafter, the sound collecting device 30 starts to detect the sound pressure V of the sound emitted by the transaxle. When the operator operates the transaxle, the operator inputs an execution instruction for the abnormal sound determination process to the abnormal sound determination device 10. The abnormal sound determination process is a process for determining whether or not the sound emitted by the transaxle, which is the sound to be inspected, contains an abnormal sound. Upon receiving the execution instruction from the operator, the CPU 12 of the abnormal sound determination device 10 starts the abnormal sound determination process. The CPU 12 realizes various processes of the abnormal sound determination process by executing the abnormal sound determination program W. Note that the CPU 12 functions as a data collection unit 12A, a timbre extraction unit 12B, a preprocessing unit 12C, and an abnormal sound determination unit 12D when executing the abnormal sound determination process.
[0015] As shown in FIG. 2, in the abnormal noise determination process, first, the data collection unit 12A executes the process of step S10. In step S10, the data collection unit 12A acquires sound data VD regarding the sound to be inspected. Specifically, in step S10, the data collection unit 12A repeatedly acquires information on the sound pressure V output by the sound collection device 30 and stores the acquired information in the memory 14. The data collection unit 12A repeats the acquisition and storage of the information on the sound pressure V until a unit period elapses from the time when the process advances to step S10. Then, when the unit period elapses, the data collection unit 12A ends the acquisition and storage of the information on the sound pressure V. Then, the data collection unit 12A advances the process to step S20. Thereafter, the CPU 12 treats a series of time series of the sound pressure V acquired over the unit period as the sound data VD. That is, in step S10, the data collection unit 12A acquires the sound data VD for a unit period targeted at the sound emitted by the transaxle, which is the sound to be inspected. The unit period is determined to ensure a time series of the sound pressure V of a sufficient length from the viewpoint of obtaining a highly accurate output result from the determination model M described later. The unit period is, for example, 1 minute. Note that when ending the process of step S10, the data collection unit 12A displays a message indicating that the acquisition of the sound data VD has been completed through the user interface 20. Upon receiving this display, the operator can determine that it is possible to stop the operation of the transaxle. The above process of step S10 is the first process.
[0016] In step S20, the timbre extraction unit 12B generates information regarding the timbre of the sound under inspection. Specifically, the timbre extraction unit 12B applies cepstrum analysis to the sound data VD in the unit period acquired in step S10. Hereinafter, the processing performed by the timbre extraction unit 12B will be described. First, the timbre extraction unit 12B performs a Fourier transform on the sound data VD. As a result, as shown by the dotted line in FIG. 3, the timbre extraction unit 12B generates a frequency spectrum S representing the sound pressure level for each frequency. Note that the waveform of the frequency spectrum S shown in FIG. 3 schematically represents an example of the waveform and does not necessarily match the actual one. The frequency spectrum S is an arrangement of the sound pressure levels for each frequency in ascending order of frequency in a graph with the frequency on the horizontal axis and the sound pressure level on the vertical axis. The sound pressure level is an index representing the magnitude of the sound pressure V, and the unit is [dB]. When the timbre extraction unit 12B generates the frequency spectrum S, it logarithmically transforms the sound pressure level for each frequency in this frequency spectrum S. After that, the timbre extraction unit 12B performs a Fourier transform on the logarithmically transformed frequency spectrum S. The information obtained as a result is called a cepstrum. The cepstrum is an arrangement of the index values for each quefrency in a graph with the quefrency on the horizontal axis and the index value on the vertical axis. The quefrency is the reciprocal of the frequency, that is, it has the dimension of time. The index value is a value that reflects the amplitude of the fluctuation component corresponding to each quefrency with respect to the waveform of the original frequency spectrum S. When the timbre extraction unit 12B generates the cepstrum, it performs lifting to extract data in the low quefrency region where the quefrency is equal to or less than a predetermined value from this cepstrum. After that, the timbre extraction unit 12B performs an inverse Fourier transform on the cepstrum extracted by lifting. As a result, the timbre extraction unit 12B returns the cepstrum to information representing the relationship between the frequency and the sound pressure level. That is, as shown by the solid line in FIG. 3, the timbre extraction unit 12B generates the waveform information J of the frequency spectrum S. This corresponds to the timbre extraction unit 12B acquiring the waveform information J of the frequency spectrum S. Similar to the frequency spectrum S, the waveform information J is an arrangement of the sound pressure levels for each frequency in ascending order of frequency in a graph with the frequency on the horizontal axis and the sound pressure level on the vertical axis.However, in consideration of the above liftering, the waveform information J is obtained by extracting a gentle fluctuation component from the waveform indicated by the frequency spectrum S. Such waveform information J is sometimes referred to as a spectral envelope. Here, the frequency range of 20 kHz or less is defined as the audible frequency range of humans, and the frequency range greater than 20 kHz is defined as the inaudible frequency range of humans. The frequency range in the waveform information J spans both the audible and inaudible ranges of humans. The series of analyses for obtaining the waveform information J of the frequency spectrum S through Fourier transform and inverse Fourier transform as described above is the cepstrum analysis of the present embodiment. As shown in FIG. 2, when the timbre extraction unit 12B generates the waveform information J, the process proceeds to step S30. Note that the process of step S20 is the second process.
[0017] In step S30, the preprocessing unit 12C generates tone color data TD. First, the preprocessing unit 12C refers to the waveform information J generated in step S20. Then, the preprocessing unit 12C rearranges the sound pressure levels for each frequency in the waveform information J in descending order of frequency. As a result, as shown in FIG. 4, the preprocessing unit 12C generates converted waveform information JA. The converted waveform information JA has the sound pressure levels for each frequency arranged in the order of decreasing frequency, with the highest frequency sound pressure level being the first. When the preprocessing unit 12C generates the converted waveform information JA, it selects the sound pressure levels within a predetermined specific frequency range R for this converted waveform information JA. The specific frequency range R in this embodiment is the entire frequency range targeted by the waveform information J and thus the converted waveform information JA. That is, the specific frequency range R includes both the frequency range of the human audible range and the frequency range of the human inaudible range. When selecting the sound pressure levels within the specific frequency range R, the preprocessing unit 12C essentially selects all the sound pressure levels included in the converted waveform information JA. Note that the preprocessing unit 12C may select the sound pressure levels by thinning them out at a predetermined frequency interval. When the preprocessing unit 12C selects the sound pressure levels, it extracts the set of the selected sound pressure levels as the tone color data TD. That is, in step S30, the preprocessing unit 12C rearranges the sound pressure levels and then extracts the tone color data TD from the waveform information J. As shown in FIG. 4, the tone color data TD has the sound pressure levels X(1) to X(n) for each frequency within the specific frequency range R arranged in descending order of frequency. Reflecting the above waveform information J, the tone color data TD shows the waveform of the gentle fluctuation component in the frequency spectrum S within the specific frequency range R. Hereinafter, the symbol "X" is used to denote the sound pressure levels targeted by the tone color data TD. Also, if necessary, the order of the frequencies is indicated in parentheses for the sound pressure level X. The above "n" is the number assigned to the last frequency and corresponds to the number of frequencies in the tone color data TD. Note that the order of the frequencies in the tone color data TD is defined for each frequency in descending order of the frequencies in the tone color data TD, with the highest frequency in the tone color data TD being the first, as described above. As shown in FIG. 2, when the preprocessing unit 12C generates the tone color data TD, it proceeds with the process to step S40.Note that the process of step S30 is the third process.
[0018] In step S40, the abnormal sound determination unit 12D outputs abnormal sound information regarding the sound data VD. The abnormal sound information is information regarding the presence or absence of abnormal sounds in the sound data VD. The abnormal sound determination unit 12D outputs the abnormal sound information using the determination model M prestored in the memory 14. The determination model M is a pre-trained learned model that outputs abnormal sound information when tone color data TD is input. The determination model M of the present embodiment is a so-called LSTM (Long Short-Term Memory) neural network. The LSTM neural network is a type of recurrent neural network.
[0019] As shown in FIG. 5, the determination model M includes an input layer, a hidden layer, and an output layer. The input layer is provided with the number of frequencies in the tone color data TD. That is, the input layer is provided for each frequency of the tone color data TD. The sound pressure level X of the corresponding frequency in the tone color data TD is input to each input layer.
[0020] The hidden layers are provided for each input layer. The sound pressure level X is input to the corresponding hidden layer from the input layer. The hidden layer performs operations using an activation function or the like on the input sound pressure level X and outputs the obtained operation result. Here, regarding the order of frequencies in the timbre data TD, the hidden layer to which the sound pressure level X of the "U" - th frequency is input is called the front - stage hidden layer. Also, the hidden layer to which the sound pressure level X of the "(U + 1)" - th frequency is input is called the rear - stage hidden layer. "U" is a natural number. In the determination model M, the output of the front - stage hidden layer is input to the rear - stage hidden layer. And in the rear - stage hidden layer, the input from the front - stage hidden layer is also used for the operations in the rear - stage hidden layer. And the output of the rear - stage hidden layer reflects the output of the front - stage hidden layer. In the determination model M, it is configured such that such information propagation between hidden layers occurs in all hidden layers. Each hidden layer is provided with a mechanism called an LSTM block that can adjust such information propagation between hidden layers. The LSTM block consists of a cell for retaining errors internally to prevent the vanishing gradient, an input gate for controlling the input to the cell, an output gate for controlling the output from the cell, and a forget gate for preventing the error from staying in the cell excessively.
[0021] Only one output layer is provided. The output layer outputs abnormal - sound information according to the output of the hidden layer to which the sound pressure level X(n) of the last frequency in the timbre data TD is input. The abnormal - sound information in this embodiment represents the probability y of the presence of abnormal sound in the sound data VD as a value in the range from "0" to "1". That is, the abnormal - sound information quantifies the likelihood that the sound data VD contains abnormal sound as a value in the range from "0" to "1".
[0022] In step S40, the abnormal sound determination unit 12D inputs the sound pressure levels X(1) to X(n) in the tone color data TD to the determination model M configured as described above. Thereby, the abnormal sound determination unit 12D outputs the presence probability y of an abnormal sound. After that, as shown in FIG. 2, the abnormal sound determination unit 12D advances the process to step S50. The process of step S40 is the fourth process. Note that the abnormal sound determination unit 12D also functions as an output unit that outputs the output result of the determination model M.
[0023] In step S50, the abnormal sound determination unit 12D determines the presence or absence of an abnormal sound based on the abnormal sound information. When the presence probability y of the abnormal sound is equal to or greater than the threshold value, the abnormal sound determination unit 12D determines that the sound data VD contains an abnormal sound. On the other hand, when the presence probability y of the abnormal sound is less than the threshold value, the abnormal sound determination unit 12D determines that the sound data VD does not contain an abnormal sound. The threshold value is a predetermined value, for example, "0.5". When the abnormal sound determination unit 12D determines the presence or absence of an abnormal sound, it displays the determination result through the user interface 20. After that, the abnormal sound determination unit 12D ends a series of processes of the abnormal sound determination process.
[0024] <Learning process> The learning process for obtaining the determination model M will be described. In the learning process, supervised learning using training data is performed. The learning process is performed before the execution of the abnormal sound determination process described above. That is, the determination model M used in the abnormal sound determination process has been learned in advance.
[0025] On the premise that the CPU 12 performs learning processing, the memory 14 stores a plurality of training data in advance. One piece of training data is one piece of tone color data TD with the probability of presence y of abnormal noise as correct answer information attached thereto. The tone color data TD constituting the training data is the same as that described in step S30. The correct answer information is "1" if the tone color data TD includes abnormal noise, and "0" if the tone color data TD does not include abnormal noise. When creating such training data, for a transfer axle of the same type as the transfer axle to be inspected, a plurality of test design contents with various changes in the assembly position and processing state of each rotating part are determined in advance. Then, the tone color data TD regarding the sound emitted by the transfer axle when the transfer axle is actually operated with each of such design contents is acquired by the same method as in steps S10 to S30 above. At the same time, it is determined by a person whether the sound when the transfer axle is operated with each design content includes abnormal noise. Then, by attaching such determination results to the respective tone color data TD, individual training data is generated. In this way, a large number of tone color data TD in the case of including abnormal noise and tone color data TD in the case of not including abnormal noise are generated respectively.
[0026] By executing a program, the CPU 12 realizes a series of processes of learning processing described below. As shown in FIG. 1, when performing the learning processing, the CPU 12 functions as a model generation unit 12E. In the learning processing, the model generation unit 12E sequentially compares, for a plurality of pieces of training data, the value output by the determination model M with the correct answer information attached to the tone color data TD in one piece of training data, using the tone color data TD as an input. At this time, the model generation unit 12E adjusts the value of the parameter in the determination model M so that the difference between the value output by the determination model M and the correct answer information becomes small. Along with this, the value of the parameter is gradually updated. Then, the model generation unit 12E determines that the learning is completed when the above-described difference becomes sufficiently small. When the learning is completed, the model generation unit 12E stores the adjusted parameter value in the memory 14 together with a program representing the calculation content of the determination model M. That is, the memory 14 stores a set of such information related to the determination model M as model data defining the determination model M.
[0027] <Operations of the Embodiment> As elements of sound that can be perceived by humans, in addition to the pitch and loudness of sound, timbre can be cited. In diagnosing a test sound based on criteria closer to human senses, it is conceivable to analyze the test sound based on timbre. The characteristics of timbre in the test sound are reflected in the waveform of the frequency spectrum S. Therefore, in this embodiment, when diagnosing a test sound, the waveform of the frequency spectrum S is analyzed by the determination model M. Here, in the waveform of the frequency spectrum S indicated by the dotted line in FIG. 3, a noise component, which is a fine fluctuation component, is superimposed on the main component, which is a gentle fluctuation component characterizing the overall waveform distribution of the frequency spectrum S. The characteristics of timbre in the test sound are mainly reflected in the main component among these fluctuation components. If the frequency spectrum S including the noise component is input to the determination model M, the information of the main component will be buried when the frequency spectrum S is analyzed by the determination model M. As a result, it becomes difficult to obtain an output that more accurately reflects the characteristics of the timbre of the test sound. Therefore, in step S20 of the abnormal sound determination process, prior to generating the timbre data TD in step S30, cepstrum analysis is applied to the sound data VD. Thereby, the noise component is removed from the frequency spectrum S and only the main component is extracted. The above-mentioned predetermined value used in liftering is a value that is predetermined to be optimal for removing the noise component and extracting the main component that characterizes the timbre of the test sound.
[0028] <Effects of the Embodiment> (1) The determination model M of this embodiment outputs abnormal sound information with the timbre data TD reflecting the waveform of the frequency spectrum S as input. That is, in this embodiment, it is determined whether or not the test sound contains abnormal sounds based on the characteristics of timbre. In this way, by determining whether or not the test sound contains abnormal sounds based on the characteristics of timbre, information can be obtained on a criterion closer to human senses regarding the presence or absence of abnormal sounds.
[0029] (2) In this embodiment, a recurrent neural network is adopted as the determination model M. If a recurrent neural network is used as the determination model M, the tone color data TD can be analyzed while maintaining the order of frequencies in the tone color data TD, that is, without disturbing the order of the sound pressure levels X in the tone color data TD. Utilizing such a determination model M is suitable for obtaining an output that reflects the characteristics of the waveform of the frequency spectrum S.
[0030] (3) In order to obtain a determination result based on criteria close to human perception, it is conceivable to analyze the waveform of the frequency spectrum S by specializing in the frequency range of the human audible range among the frequency spectra S. On the other hand, although the sound in the frequency range of the human inaudible range is difficult to be heard by the human ear, the waveform of the frequency spectrum S in the inaudible frequency range also reflects the characteristics of the tone color emitted by the sound under test. Therefore, in comprehensively determining whether the sound under test contains abnormal sounds, it is preferable to analyze the waveform of the frequency spectrum S for both the human audible range and the inaudible range. From this perspective, in this embodiment, the sound pressure level X in the frequency range covering both the human audible range and the inaudible range is included in the tone color data TD. And in this embodiment, the waveform of the frequency spectrum S covering both the human audible range and the inaudible range is analyzed by the determination model M.
[0031] Now, due to the characteristics of the recurrent neural network, the output of each hidden layer propagates to the next hidden layer in the order of the timbre data TD. And finally, the output of the hidden layer of the last frequency reaches the output layer. In relation to such characteristics, in the output of the hidden layer of the last frequency and thus the output of the output layer, the output of the hidden layer in the frequency range closer to the latter half is more likely to be reflected compared to the output of the hidden layer in the frequency range closer to the former half in the order of frequencies in the timbre data TD. As described above, the timbre data TD of the present embodiment subjects the waveforms of the frequency spectra S in both the human audible range and the inaudible range to analysis by the determination model M. In this way, when analyzing the waveforms across both the human audible range and the inaudible range, if the characteristics of the waveforms in the frequency range of the audible range can be vividly reflected by the output result of the determination model M, it becomes possible to obtain a determination result based on a criterion closer to human perception while taking into account the potential information included in the inaudible range. Therefore, in the present embodiment, when generating the timbre data TD, the sound pressure level X is arranged in descending order of frequencies. And such a group of sound pressure levels X arranged in descending order is used as the input to the determination model M. And the determination model M sequentially propagates the output of the hidden layer towards the hidden layer of the lower frequencies. By adopting such a configuration, the characteristics of the waveforms in the frequency range of the audible range are more likely to be reflected in the result finally output from the output layer compared to the inaudible range.
[0032] <Modification Example> The above embodiment can be implemented with the following modifications. The above embodiment and the following modification examples can be implemented in combination with each other within a technically non - conflicting range.
[0033] · The abnormal sound information is not limited to the example of the above embodiment. For example, the abnormal sound information may represent the degree of abnormality of the sound under test at multiple levels. The abnormal sound information only needs to indicate whether abnormal sounds are included in the sound data VD. If it is known in advance that there are multiple types of abnormal sounds that may be included in the sound under test, the determination model M may be configured to output information regarding the presence or absence of each of those individual abnormal sounds.
[0034] ·The specific frequency range R is not limited to the examples of the above embodiments. For example, the specific frequency range R may be only the frequency range of the human audible range. In this case, in step S30, the tone color data TD is generated only for the human audible range. At the same time, in accordance with the number of frequencies included in the tone color data TD, the number of input layers and thus the number of hidden layers in the determination model M are also changed. When the specific frequency range R is limited to only the human audible range, there are the following advantages. That is, when the specific frequency range R is limited to only the human audible range, since the waveform of the frequency spectrum S is analyzed specifically for the human audible range, a determination result can be obtained based on a criterion closer to human perception. Moreover, when the frequency range of the tone color data TD is limited to only the audible range, accordingly, the number of input layers and thus the number of hidden layers for each input layer in the determination model M also decreases. Therefore, the number of arithmetic processes required to obtain the final output result decreases. Therefore, the processing load of the abnormal sound determination device 10 when performing arithmetic operations using the determination model M can be suppressed.
[0035] ·Including the case where the specific frequency range R is only the frequency range of the human audible range as in the above modification example, it is not essential to arrange the sound pressure levels for each frequency in descending order of frequency when generating the tone color data TD from the waveform information J. That is, the sound pressure levels in the tone color data TD may be arranged in ascending order starting from the one with the lowest frequency. In this case, the order of frequencies in the tone color data TD is defined for each frequency in ascending order of the frequencies in the tone color data TD, with the lowest frequency in the tone color data TD being the first, contrary to the above embodiment. And in this case, the output layer of the determination model M outputs abnormal sound information according to the output from the hidden layer to which the sound pressure level of the highest frequency in the tone color data TD is input.
[0036] ·The configuration of the determination model M is not limited to the examples of the above embodiments. The determination model M may be configured to be able to output appropriate abnormal sound information. Appropriate learning may be performed on the determination model M according to its configuration. Regarding the learning of the determination model M, the parameters of the determination model M may be readjusted based on the output result obtained in step S40 of the abnormal sound determination process.
[0037] ·The learning process may be performed by a computer different from the abnormal sound determination device 10. Then, the determination model M generated by the learning process may be stored in the memory 14 of the abnormal sound determination device 10 prior to the execution of the abnormal sound determination process.
[0038] ·The determination model M is not limited to one using a recurrent neural network. For example, the determination model M may be configured by a convolutional neural network. In this case, it is conceivable to use an image of the tone color data TD as input data to the determination model M. Then, the determination model M is configured to output abnormal sound information by analyzing the waveform pattern depicted in this image.
[0039] ·The determination model M is not limited to one that has been machine-learned in advance. The determination model M only needs to be configured to output abnormal sound information when the tone color data TD is input. For example, the determination model M may be configured to output abnormal sound information using a technique such as the dynamic time warping method that compares the similarity of time-series data. The input format of the tone color data TD to the determination model M may be appropriately changed according to the content of the determination model M.
[0040] ·Regarding the process of step S10, the mode in which the data collection unit 12A acquires the sound data VD is not limited to the example of the above embodiment. For example, the time series of the sound pressure V of the sound to be inspected may be stored in the memory 14 in advance, and the data collection unit 12A may acquire the sound data VD by reading out such a time series.
[0041] ·The sound collection device 30 is not limited to the example of the above embodiment. The sound collection device 30 may be any device that can detect information related to sound. · The user interface 20 is not limited to the examples of the above embodiments. The user interface 20 may be configured to be capable of providing information to the operator or allowing the operator to input information. Devices for providing information to the operator and devices for allowing the operator to input information may be provided separately. The device for providing information to the operator may use sound, such as a speaker.
[0042] · The inspection sound is not limited to the examples of the above embodiments. · The abnormal sound determination device 10 may have any of the following configurations (a) to (c). (a) The abnormal sound determination device 10 includes one or more processors that execute various processes according to a computer program. The processor includes a CPU and memories such as a RAM and a ROM. The memory stores program codes or instructions configured to cause the CPU to execute processes. The memory, that is, the computer-readable medium, includes any available medium accessible by a general-purpose or dedicated computer.
[0043] (b) The abnormal sound determination device 10 includes one or more dedicated hardware circuits that execute various processes. Examples of the dedicated hardware circuit include an application-specific integrated circuit, that is, an ASIC or an FPGA.
[0044] (c) The abnormal sound determination device 10 includes a processor that executes some of the various processes according to a computer program and a dedicated hardware circuit that executes the remaining processes of the various processes.
Explanation of Reference Numerals
[0045] 10… Abnormal sound determination device 12… CPU 14… Memory M… Determination model
Claims
1. comprising an execution unit and a storage unit; when tone color data, which is a waveform of a frequency spectrum obtained by performing a Fourier transform on sound data of a sound to be inspected and indicates a waveform in a specific frequency range, is input to the storage unit, the storage unit stores a model that outputs abnormal sound information, which is information regarding the presence or absence of an abnormal sound in the sound data; the execution unit performs a first process of acquiring the sound data for a unit period targeted at the sound to be inspected; performs a second process of acquiring waveform information of the frequency spectrum by further performing an inverse Fourier transform on a cepstrum obtained by applying cepstrum analysis to the sound data for the unit period to Fourier-transform the frequency spectrum of the sound data; performs a third process of extracting, as the tone color data, sound pressure levels for each frequency arranged in the order of frequency in the specific frequency range from the waveform information; and performs a fourth process of outputting the abnormal sound information by inputting the tone color data to the model An abnormal sound determination device.
2. The model is a recurrent neural network that has been machine-learned in advance The abnormal sound determination device according to Claim 1.
3. The model includes an input layer provided for each frequency in the tone color data; a hidden layer provided for each input layer, which performs an operation on the sound pressure level for each frequency input from the corresponding input layer and outputs the operation result; and one output layer, wherein, taking the largest frequency in the tone color data as the first, the order of each frequency is defined in descending order of the frequencies in the tone color data, and when "U" is a natural number, the output of the hidden layer for the "U"th frequency is configured to be reflected in the output of the hidden layer for the "(U + 1)"th frequency; the output layer outputs the abnormal sound information according to the output of the hidden layer for the last frequency in the tone color data; and the specific frequency range includes both the frequency range of the human audible range and the frequency range of the inaudible range The abnormal sound determination device according to Claim 2.
4. The specific frequency range includes only the frequency range of the human audible range The abnormal sound determination device according to Claim 2.
5. targeting a computer comprising an execution unit and a storage unit The memory unit stores a model that outputs abnormal sound information, which is information regarding the presence or absence of abnormal sounds in the sound data, when tone color data, which is a waveform of a frequency spectrum obtained by performing a Fourier transform on the sound data of the sound under inspection and indicates a waveform in a specific frequency range, is input; to the execution unit, perform a first process of acquiring the sound data of a unit period for the sound under inspection; perform a second process of acquiring waveform information of the frequency spectrum, which is obtained by performing an inverse Fourier transform on a cepstrum obtained by applying cepstrum analysis to the sound data of the unit period and then performing a Fourier transform on the sound data; perform a third process of extracting, as the tone color data, the sound pressure level for each frequency arranged in the order of frequency in the specific frequency range from the waveform information; perform a fourth process of outputting the abnormal sound information by inputting the tone color data to the model. Abnormal sound determination program.
Citation Information
Patent Citations
Inspecting device of tone color
JP1990057924A
Drive force control device
JP2009074395A
Controller, control method of controller and program
JP2023183077A
Electric artificial larynx device
WO2015019835A1
Shunt murmur analysis device, shunt murmur analysis method, computer program, and recording medium
WO2016207951A1