Abnormal sound detection method, abnormal sound detection program, and abnormal sound detection system

The abnormal sound detection method and system address the challenge of training inspectors by adjusting sound data for human hearing and using machine learning to replicate expert judgments, enhancing inspection skills and accuracy.

JP7741473B2Active Publication Date: 2025-09-18THE INSTITUTE OF PHYSICAL & CHEMICAL RESEARCH +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021158823
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-09-18
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Existing anomaly detection systems require human inspection skills that are difficult to train due to a decline in the number of experienced inspectors, leading to insufficient skill transfer.

Method used

An abnormal sound detection method and system that adjusts sound data according to human hearing ability, generating adjusted sound data for machine learning input, facilitating skill training through machine learning and visualization techniques.

Benefits of technology

Enables effective inspection skill training by replicating expert inspector judgments, allowing inspectors to improve their skills independently and accurately identify abnormalities using machine-learned determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007741473000001
    Figure 0007741473000001
  • Figure 0007741473000002
    Figure 0007741473000002
  • Figure 0007741473000003
    Figure 0007741473000003
Patent Text Reader

Abstract

To provide an abnormal sound determination method, an abnormal sound determination program, and an abnormal sound determination system, capable of being used for training an inspection skill.SOLUTION: An abnormal sound determination method includes: an acquisition step of acquiring sound data; a filter processing step of adjusting the sound data according to discrimination ability of a person; and a data generation step pf generating the adjusted sound data as input data for machine learning.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an abnormal sound detection method, an abnormal sound detection program, and an abnormal sound detection system. [Background technology]

[0002] Conventionally, a method has been proposed for inspecting structures by conducting a hammering test on concrete to detect structural abnormalities from the hammering sounds, thereby detecting concrete lifting, corrosion, etc. For example, Non-Patent Document 1 discloses an anomaly detection system that digitizes hammering sounds from a hammer or the like and enables the collection, storage, and analysis of the digitalized data. [Prior art documents] [Patent documents]

[0003] [Non-Patent Document 1] Masahiro Murakawa, "Anomaly Detection System Using Artificial Intelligence Technology and Its Industrial Applications," Journal of the Atomic Energy Society of Japan, Vol. 59, No. 6 (2017), pp. 31-35 Summary of the Invention [Problem to be solved by the invention]

[0004] The system in Non-Patent Document 1 automatically determines whether or not a structure has an abnormality, but there are cases where the abnormality determination requires the use of human inspection skills depending on the abnormality mode of the object, etc. However, in a situation where the number of inspectors has been on a downward trend in recent years, there are fewer opportunities to teach the skills of experienced inspectors, and there is a problem that skill training cannot be sufficiently provided to other inspectors, etc.

[0005] In view of the above, an object of the present disclosure is to provide an abnormal sound detection method, an abnormal sound detection program, and an abnormal sound detection system that can be used in inspection skill training. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, the abnormal sound detection method according to the present disclosure includes an acquisition step of acquiring sound data, a filtering step of adjusting the sound data in accordance with the human hearing ability, and a data generation step of generating the adjusted sound data as input data for machine learning.

[0007] In order to achieve the above-mentioned object, the abnormal sound detection program according to the present disclosure causes a computer to execute an acquisition step of acquiring sound data, a filtering step of adjusting the sound data in accordance with the human hearing ability, and a data generation step of generating the adjusted sound data as input data for machine learning.

[0008] In order to achieve the above-mentioned object, the abnormal sound detection system according to the present disclosure includes an input unit that acquires sound data, and a processing unit that adjusts the sound data in accordance with the human hearing ability and generates the adjusted sound data as input data for machine learning. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to provide an abnormal sound detection method, an abnormal sound detection program, and an abnormal sound detection system that can be used in inspection skill training. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an abnormal sound detection system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an outline of processing in an abnormal sound detection system. [Figure 3] 10 is a flowchart showing the processing of an abnormal sound detection program. [Figure 4] FIG. 10 is a diagram illustrating sound data. [Figure 5] FIG. 10 is a diagram showing primary divided data transformed into the frequency domain. [Figure 6] FIG. 10 is a diagram illustrating frequency weighting characteristics. [Figure 7] FIG. 1 is a schematic diagram showing a 1 / 3 octave band. [Figure 8] FIG. 10 is a diagram illustrating time weighting characteristics. [Figure 9] FIG. 2 is a diagram showing the configuration of image data. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The abnormal sound detection system 1 shown in Fig. 1 includes an input unit 11 that acquires sound data 21, an output unit 12 that outputs a determination result 3 (analysis result) of the sound data 21, a control unit 13, and a storage unit 14. The input unit 11 may acquire the sound data 21 from an external device using wired or wireless communication, or may function as a sensor (sound collection means) such as a microphone, vibration pickup, or vibration acceleration pickup, and be configured to convert sound (vibration) collected from the outside into an electrical signal.

[0012] The output unit 12 can display the determination result 3 determined by the abnormal sound determination program 23 described below, for example, on a display, or can output data related to the determination result 3 to another device (for example, another display unit or speaker) via wired or wireless connection.

[0013] The control unit 13 is a circuit such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), or FPGA (Field Programmable Gate Array), and functions as a processing unit that runs various programs. The control unit 13 controls the operations of the input unit 11, output unit 12, and storage unit 14. The control unit 13 can also execute various programs such as an abnormal sound detection program 23.

[0014] The storage unit 14 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), an optical disk, or a semiconductor memory. The storage unit 14 stores sound data 21, image data 22, and an abnormal sound detection program 23. The sound data 21 is data acquired by the input unit 11. The sound data 21 may include abnormal data including abnormal sounds and non-abnormal sound data not including abnormal sounds. The multiple sound data 21 input to the abnormal sound detection program 23 may include either or both abnormal data and non-abnormal data depending on the application and function of the abnormal sound detection system 1. The abnormal sound detection program 23 includes an input data generation unit 231 and a determination unit 232. The input data generation unit 231 generates image data 22 based on the sound data 21. The image data 22 is used as input data to be input to the determination unit 232.

[0015] The determination unit 232 has a function of inputting image data 22 and determining whether or not an abnormal sound is included in the original sound data 21. The determination unit 232 can also perform machine learning using the input image data 22 as training data. When using image data 22 as training data, the image data 22 can include a label indicating whether or not an abnormal sound is included in the original sound data 21. The abnormal sound detection program 23 can be stored in a computer-readable storage medium (for example, a storage device such as the storage unit 14).

[0016] The abnormal sound detection system 1 may be configured with one device or multiple devices. For example, some of the input unit 11, output unit 12, control unit 13, and storage unit 14 may be located in multiple different devices. The abnormal sound detection system 1 may also be configured such that multiple devices each including the input unit 11, output unit 12, control unit 13, and storage unit 14 are connected to each other so as to be able to communicate with each other via wired or wireless communication, and so as to be able to execute the processing described in this embodiment. The abnormal sound detection system 1 also includes a computer (not shown) for controlling the abnormal sound detection system 1.

[0017] Figure 2 is a diagram showing an overview of the abnormal sound detection method in the abnormal sound detection system 1. The sound data 21 acquired by the abnormal sound detection system 1 is converted into intermediate data (primary divided data 21n, secondary divided data 21nm) by the input data generation unit 231 (see Figure 1). Image data 22 is then generated based on the intermediate data. Multiple image data 22 can be generated corresponding to each piece of sound data 21. When the determination unit 232 performs an abnormality determination on the image data 22, the abnormal sound detection system 1 outputs a determination result 3 for the image data 22 (in other words, the original sound data 21).

[0018] Next, the abnormal sound detection method of this embodiment will be described. Fig. 3 is a flowchart showing the processing in the input data generation unit 231 of the abnormal sound detection program 23. In step S01, the control unit 13 acquires sound data 21 via the input unit 11 (acquisition step). The sound data 21 includes a plurality of hammering waveforms 210 (sound waveforms) obtained when a hammering test is performed on concrete such as the inner wall of a tunnel (see Fig. 2). When the sound data 21 includes a plurality of hammering waveforms 210, the control unit 13 can extract one of the hammering waveforms 210 from the sound data 21.

[0019] Fig. 4 is a diagram showing a hitting sound waveform 210 of the sound data 21. The vertical axis of Fig. 4 represents sound pressure [Pa], and the horizontal axis represents time [s]. The dashed lines in Fig. 4 represent timings separated by 50 ms from timing T0 = 0 [s].

[0020] In the filtering process from step S02 to step S07, the control unit 13 performs a process of adjusting the sound data 21 in accordance with the human hearing ability.

[0021] In step S02, the control unit 13 performs a time division process to divide the sound data 21 into intervals of 50 ms from timing T0 as a time adjustment. The control unit 13 divides the sound data 21 into a plurality of predetermined time domain intervals including the sound generation timing (T0). In this embodiment, the interval from -100 ms to +300 ms is divided into eight intervals (1) to (8) based on timing T0, and primary divided data 211 to 218 (hereinafter, may be collectively referred to as primary divided data 21n (n = 1 to 8)) are generated as intermediate data. Note that the interval before -100 ms and the interval after +300 ms are excluded from the primary divided data 21n. When the hitting sound waveform 210 in FIG. 4 is represented in the time domain, sound pressure values ​​are observed over approximately intervals (3) and (4).

[0022] In step S03, the control unit 13 converts each of the primary divided data 211 to 218 divided into sections (1) to (8) into data in the frequency domain. Fig. 5 is a diagram showing the primary divided data 211 to 218 converted into the frequency domain. The control unit 13 can convert the time-domain primary divided data 211 to 218 into frequency-domain data (frequency spectrum) using a Fourier transform.

[0023] In step S04, the control unit 13 uses the frequency weighting characteristic F1 shown in FIG. 6 to weight the primary divided data 211 to 218 converted into the frequency domain, and adjusts the sound pressure level.

[0024] When the sound data 21 is collected by a microphone or the like, the sound data 21 may contain sounds in a frequency range that humans cannot hear. For example, assuming a hammering waveform 210 obtained when an inspector performs a hammering test, even if the sound data 21 containing sounds in a frequency range that humans cannot hear is used as training data or data to be judged (collectively referred to as "input data"), it is difficult for the judgment unit 232 to reproduce the judgment skill of an expert inspector. Therefore, the primary divided data 211 to 218 are adjusted using the frequency weighting characteristic F1.

[0025] 6 is a sound pressure adjustment tailored to the human ear, and indicates that hearing sensitivity is low in the low-frequency and high-frequency ranges. Therefore, the control unit 13 adjusts the sound pressure level by relatively lowering the sound pressure level for each frequency of the sound data 21 in the low-frequency and high-frequency ranges, thereby achieving weighting adjustment in accordance with human hearing sensitivity.

[0026] In step S05, the control unit 13 extracts audible band data from the primary divided data 211-218 (sound data 21) using a 1 / 3 octave bandpass filter, and performs frequency band adjustment to divide the primary divided data 211-218 (sound data 21) into multiple frequency bands. The control unit 13 applies a bandpass filter with a frequency interval of 1 / 3 octave to each of the primary divided data 211-218, and generates secondary divided data 21nm (n=1 to 8, m=1 to 32, where n represents the time division region and m represents the frequency division region), which is intermediate data obtained by frequency division into multiple bands (32 regions in this embodiment). Note that one octave is the frequency interval at which the frequency ratio is doubled.

[0027] Figure 7 shows the sections [1] to

[32] of a 1 / 3 octave bandpass filter. The lowest frequency section [1] is located near the boundary between the audible and inaudible ranges on the low frequency side. The highest frequency section

[32] is located near the boundary between the audible and inaudible ranges on the high frequency side.

[0028] In addition, the control unit 13 is not limited to using a 1 / 3 octave bandpass filter, and may also apply a bandpass filter with a frequency interval of 1 / N octave (for example, N is a natural number greater than or equal to 24) to each of the primary divided data 211 to 218 to generate secondary divided data 21nm (n=1 to 8, m=1 to 32, where n represents the time division region and m represents the frequency division region) that is frequency-divided into multiple sections.

[0029] In step S06, the control unit 13 converts the second divided data 21nm into time domain data. The control unit 13 can convert the second divided data 21nm in the time domain into time domain data (time spectrum) using an inverse Fourier transform.

[0030] In step S07, the control unit 13 weights the secondary divided data 21 nm converted into the time domain using the time weighting characteristic F2 shown in FIG. 8, and adjusts the sound pressure level (time weighting process). The sound pressure of a sound changes in an extremely short time. The time weighting characteristic F2 is a characteristic (a so-called fast characteristic) that approximates the time response of the human ear, and has a rise time constant τ=125 ms. The slope of the rise time of the time weighting characteristic F2 from the timing (0 s) is 34.7 dB / s. Note that the characteristic used in the time weighting process is not limited to the time weighting characteristic F2.

[0031] 9 from the secondary divided data 21nm acquired by the filter processing of steps S02 to S07 (data generation step). The image data 22 is generated from the plurality of adjusted secondary divided data 21nm (sound data 21). The image data 22 is used as input data that can be read by the determination unit 232 of the abnormal sound detection program 23.

[0032] The image data 22 is arranged in two mutually orthogonal directions of frequency and time, with sound pressures (sound pressure levels) corresponding to each frequency band interval [1] to

[32] and each time interval (1) to (8) arranged as grayscale pixel values. The pixel value (grayscale shading) of each cell 221 indicates the sound pressure level. Note that the sound pressure level used as the pixel value of the image data 22 may be the average or integral value of the sound pressure of the secondary divided data 21 nm divided by frequency and time. For ease of explanation, FIG. 9 shows three levels of pixel values ​​for simplicity. The magnitude (brightness) of the pixel values ​​increases in the order of cell 221a, cell 221b, and cell 221c. Therefore, the sound pressure levels increase in the order of cell 221a, cell 221b, and cell 221c.

[0033] The abnormal sound detection system 1 generates one image data 22 corresponding to one piece of sound data 21 (specifically, sound data including one hitting sound waveform 210) input to the input unit 11. From the plurality of pieces of sound data 21, a plurality of image data 22 corresponding to each piece of sound data 21 are generated.

[0034] The image data 22 generated after adjusting the sound data 21 can be input as training data for machine learning to the determination unit 232. When the image data 22 is used as training data, the training data can include a label indicating whether or not the original sound data 21 corresponding to the image data 22 contains an abnormal sound.

[0035] Furthermore, the image data 22 can be input as data to be determined to the learned determination unit 232. The control unit 13 can cause the determination unit 232 to determine whether or not the input image data 22 includes an abnormal sound. Thereafter, the control unit 13 can output, via the output unit 12, a determination result 3 of the presence or absence of an abnormality in the image data 22 (sound data 21) made by the determination unit 232.

[0036] In this way, the abnormality detection method using the abnormal sound detection system 1 can include a detection step of inputting the image data 22, which is input data, into a machine learning program (determination unit 232) as teacher data or data to be detected.

[0037] In this embodiment, the judgment unit 232 is trained based on sound data 21 adjusted in imitation of the human hearing ability, and the judgment unit 232 judges whether or not there is an abnormal sound in other sound data 21. Therefore, for example, in the field of hammering inspection, by comparing the judgment result of a trainee who judges sound data 21 by listening to it with the judgment result 3 obtained by judging the same sound data 21 using the abnormal sound judgment program 23, the trainee can compare the judgment result with the judgment result that is assumed to be that of a skilled inspector, thereby improving his inspection skills.

[0038] Furthermore, by using the abnormal sound detection system 1 (abnormal sound detection program 23) at an actual inspection work site, an inspector can perform hammering inspection work while referring to the detection results 3 equivalent to those of an experienced inspector, even without being accompanied by other inspectors. Therefore, the inspection skills of inspectors can be improved in practice while working on-site.

[0039] The abnormal sound detection method of this embodiment may also include a visualization step of creating a basis image that visualizes the basis for determining whether or not an abnormality exists using Grad-CAM. Grad-CAM is a technology that focuses on features extracted by the final convolutional layer of a convolutional neural network to visualize which parts of an image machine learning is looking at to make a determination. For example, by creating and displaying a basis image in which cells that contribute most to the basis for the determination are colored using a color gradation on an image with the same matrix number as the image data 22, other inspectors can objectively understand which frequencies and timings of sounds an experienced inspector primarily uses as the basis for determining whether or not an abnormality exists. Therefore, by training a trainee using the abnormal sound detection program 23, the trainee can learn which ranges of sounds to listen to in order to make an accurate determination, even if they do not have sufficient opportunities for manual training.

[0040] As described above, the abnormal sound detection method executable in the abnormal sound detection system 1 according to the embodiment of the present disclosure includes an acquisition step (S01) of acquiring sound data 21, a filtering step (part or all of S02 to S07) of adjusting the sound data 21 in accordance with the human hearing ability, and a data generation step of generating the adjusted sound data 21 as input data for machine learning. As a result, even when an abnormality is determined using a machine-learned determination program (e.g., the determination unit 232) that has already undergone machine learning, it is possible to obtain a determination result that is close to that of an experienced inspector. In this way, it is possible to configure an abnormal sound detection method, an abnormal sound detection program, and an abnormal sound detection system 1 that can be used for training inspection skills.

[0041] In addition, the image data 22 generated as input data is adjusted to include the audible range while excluding the inaudible range, so that excess data that is less relevant to hearing discrimination skills can be omitted, reducing the increase in data volume.

[0042] This concludes the description of the embodiment of the present disclosure, but the aspects of the present disclosure are not limited to this embodiment.

[0043] For example, in the filter processing process of this embodiment, a configuration has been described in which all of the sound pressure level, frequency band, and time are adjusted, but the configuration may also be such that one or more (including some or all) of the sound pressure level, frequency band, and time are adjusted for the sound data 21.

[0044] Furthermore, in this embodiment, a configuration has been described in which the sound data 21 includes a hammering waveform 210 of a hammering test on concrete, but the sound data 21 may also include some or a plurality of sound waveforms (corresponding to the hammering waveform 210 shown in FIG. 4 ) of metal processing sounds, mechanical sounds, vibration sounds, hammering sounds, or noises emitted from automobiles, trains, bullet trains, airplanes, etc. The sound data 21 may be a sound waveform that does or does not include abnormal sounds. Possible abnormal sounds include tire sounds, malfunctioning sounds from the body of a car, etc., and abnormal road surface sounds.

[0045] Furthermore, the sound data 21 is not limited to data acquired by inspection, but may be data acquired by any other means.

[0046] Furthermore, the order of the processes in steps S02 to S07 of the abnormal sound detection method shown in FIG. 3 is just one example, and the order may be changed as appropriate. [Explanation of symbols]

[0047] 1. Abnormal sound detection system 3 Judgment result 11 Input section 12 Output section 13 Control Unit 14 Storage section 21 Sound data 21n(211~218) Primary division data 21nm secondary division data 22 Image data 23 Abnormal sound detection program 210 Impact waveform 221(221a~221c) Cell 231 Input data generation unit 232 Judgment section F1 frequency weighting F2 Time weighting characteristics T0 timing τ time constant

Claims

1. an acquisition step of acquiring sound data; a filtering process for adjusting the sound data in accordance with the human hearing ability; a data generation step of generating the adjusted sound data as input data for machine learning; Including, The filtering step adjusts all of the sound pressure level, frequency band, and time, adjusting the sound pressure level by weighting the magnitude of the sound pressure level for each frequency of the sound data in accordance with human hearing sensitivity; As the frequency band adjustment, audible band data is extracted from the sound data using a 1 / 3 octave band pass filter, and the sound data is divided into a plurality of bands; As the time adjustment, the sound data is divided into a plurality of predetermined time domains including sound generation timings. Abnormal sound determination method.

2. An acquisition step of acquiring sound data; a filtering process for adjusting the sound data in accordance with the human hearing ability; a data generation step of generating the adjusted sound data as input data for machine learning; Including, The filtering step adjusts all of the sound pressure level, frequency band, and time, The input data is image data in which the sound pressure levels corresponding to the frequency bands and the time periods are arranged as grayscale pixel values ​​in two mutually orthogonal directions of the frequency bands and the time periods. Abnormal sound determination method.

3. The abnormal sound detection method described in Claim 1, wherein the input data is image data in which the sound pressure levels corresponding to the frequency band and the time are arranged as grayscale pixel values ​​in two mutually perpendicular directions of the frequency band and the time.

4. The sound data includes abnormal data including an abnormal sound and non-abnormal sound data not including the abnormal sound, The abnormal sound detection method according to claim 1 , further comprising a determination step of inputting the input data as teacher data or data to be determined into a machine learning program.

5. A method for detecting abnormal sounds as described in claim 4 which cites claim 2, or claim 4 which cites claim 3, which includes a visualization step of creating a basis image that visualizes the basis for determining whether or not there is an abnormality using Grad-CAM.

6. A method for determining abnormal sounds described in any one of claims 1 to 5, wherein the sound data includes either mechanical sounds, vibration sounds, or hammering sounds.

7. An acquisition step of acquiring sound data; a filtering process for adjusting the sound data in accordance with the human hearing ability; a data generation step of generating the adjusted sound data as input data for machine learning; Including, The filtering step adjusts all of the sound pressure level, frequency band, and time, adjusting the sound pressure level by weighting the magnitude of the sound pressure level for each frequency of the sound data in accordance with human hearing sensitivity; As the frequency band adjustment, audible band data is extracted from the sound data using a 1 / 3 octave band pass filter, and the sound data is divided into a plurality of bands; an abnormal sound detection program for causing a computer to divide the sound data into a plurality of predetermined time domains including sound generation timings as the time adjustment;

8. An acquisition step of acquiring sound data; a filtering process for adjusting the sound data in accordance with the human hearing ability; a data generation step of generating the adjusted sound data as input data for machine learning; Including, The filtering step adjusts all of the sound pressure level, frequency band, and time, an abnormal sound detection program that causes a computer to execute the following: converting the input data into image data in which the sound pressure levels corresponding to the frequency bands and the time periods are arranged as grayscale pixel values ​​in two mutually perpendicular directions of the frequency bands and the time periods.

9. An input unit for acquiring sound data; a processing unit that adjusts the sound data in accordance with human hearing discrimination ability and generates the adjusted sound data as input data for machine learning; Equipped with The processing unit adjusts all of the sound pressure level, frequency band, and time, adjusting the sound pressure level by weighting the magnitude of the sound pressure level for each frequency of the sound data in accordance with human hearing sensitivity; As the frequency band adjustment, audible band data is extracted from the sound data using a 1 / 3 octave band pass filter, and the sound data is divided into a plurality of bands; The abnormal sound detection system divides the sound data into a plurality of predetermined time domains including sound generation timings as the time adjustment.

10. an input unit for acquiring sound data; a processing unit that adjusts the sound data in accordance with human hearing discrimination ability and generates the adjusted sound data as input data for machine learning; Equipped with The processing unit adjusts all of the sound pressure level, frequency band, and time, an abnormal sound detection system, wherein the input data is image data in which the sound pressure levels corresponding to the frequency bands and the time periods are arranged as grayscale pixel values ​​in two mutually perpendicular directions of the frequency bands and the time periods.

Citation Information

Patent Citations

  • Bearing detection method based on convolutional neural network

    CN110579354A

  • Sound analyzing method and sound analyzer

    JP2002267529A

  • Sound evaluation ai system and sound evaluation ai program

    JP2020095123A

  • Examination data processing device and examination data processing method

    WO2016117358A1