Abnormality diagnosis apparatus, abnormality diagnosis method, and program

The abnormality diagnosis device diagnoses equipment abnormalities by creating and comparing normal and abnormal sound models based on pitch features, effectively identifying abnormal operation through sound analysis.

JP2025183933APending Publication Date: 2025-12-17FUJI ELECTRIC CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025084623
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2025-05-21
Publication Date
2025-12-17

AI Technical Summary

Technical Problem

Conventional techniques are unable to diagnose abnormalities in equipment from features related to the pitch of sounds.

Method used

An abnormality diagnosis device that creates a normal model based on normal sound data and a diagnosis target model based on abnormal sound data, using feature quantities related to the pitch of sound, and compares these models to diagnose equipment abnormalities.

Benefits of technology

Enables accurate diagnosis of equipment abnormalities by analyzing pitch features, particularly distinguishing high-pitched sounds indicative of abnormal operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025183933000001_ABST
    Figure 2025183933000001_ABST
Patent Text Reader

Abstract

To diagnose an abnormality of an instrument from feature quantities relating to a pitch of sound.SOLUTION: An abnormality diagnosis apparatus according to one aspect of the present disclosure includes: a first model creation unit that creates a normal model constituted by feature quantities relating to a pitch of sound included in normal sound on the basis of the normal sound representing sound acquired from an instrument in a normal state; a second model creation unit that creates a diagnosis target model constituted by feature quantities relating to a pitch of sound included in diagnosis target sound on the basis of the diagnosis target sound representing sound acquired from a target instrument to be subjected to abnormality diagnosis; and a diagnosis unit that diagnoses an abnormality of the target instrument by comparing the diagnosis target model with the normal model. The feature quantities relating to the pitch of sound are scores representing how much sounds of frequencies corresponding to respective ones of twelve classes defined on the basis of chroma are included.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an abnormality diagnosis device, an abnormality diagnosis method, and a program. [Background technology]

[0002] There is known technology for detecting abnormalities in a device from the sound generated by the device (for example, Patent Document 1, etc.). On the other hand, a characteristic called "pitch" is known as a characteristic that represents the pitch of a sound, and it is generally known that when a device is operating abnormally, it is likely to generate a high-pitched sound (i.e., a high-pitched sound). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-144024 Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional techniques have not been able to diagnose abnormalities in equipment from features related to the pitch of sounds.

[0005] The present disclosure has been made in consideration of the above points, and aims to diagnose abnormalities in equipment from feature quantities related to the pitch of sound. [Means for solving the problem]

[0006] An abnormality diagnosis device according to one aspect of the present disclosure includes a first model creation unit that creates a normal model based on normal sound, which represents sound obtained from an equipment under normal conditions, and that creates a normal model composed of feature quantities related to the pitch of the sound contained in the normal sound; a second model creation unit that creates a diagnosis target model based on diagnosis target sound, which represents sound obtained from the equipment that is the target of abnormality diagnosis, and that creates a diagnosis target model composed of feature quantities related to the pitch of the sound contained in the diagnosis target sound; and a diagnosis unit that diagnoses an abnormality in the equipment by comparing the diagnosis target model with the normal model, wherein the feature quantities related to the pitch of the sound are scores that represent the degree to which sounds of frequencies corresponding to each of 12 classes defined based on chroma are included. [Effects of the Invention]

[0007] Equipment abnormalities can be diagnosed from the characteristics related to the pitch of the sound. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of a hardware configuration of an abnormality diagnosis device according to a first embodiment. [Figure 2] 1 is a diagram illustrating an example of a functional configuration of an abnormality diagnosis device according to a first embodiment; [Figure 3] 10 is a flowchart illustrating an example of a model creation process according to the first embodiment. [Figure 4] 4 is a flowchart illustrating an example of an abnormality diagnosis process according to the first embodiment. [Figure 5] FIG. 10 is a diagram illustrating a comparison between normal and abnormal conditions. [Figure 6] FIG. 10 is a diagram illustrating an example of a functional configuration of an abnormality diagnosis device according to a second embodiment. [Figure 7] 10 is a flowchart illustrating an example of a conversion method determination process according to the second embodiment. [Figure 8] 10 is a flowchart illustrating an example of a model creation process according to the second embodiment. [Figure 9] 10 is a flowchart illustrating an example of an abnormality diagnosis process according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] First and second embodiments of the present invention will be described in detail below with reference to the drawings. In each of the following embodiments, an abnormality diagnosis device 10 will be described that can diagnose abnormalities in equipment based on features related to the pitch of sound generated by the equipment (hereinafter also referred to as "pitch features"). Note that sound may be referred to as, for example, "acoustics." Equipment may be referred to as, for example, "device," "terminal," "facility," etc. Specific examples of equipment include machine tools (e.g., cutting machines, bending machines, etc.), industrial machinery (e.g., conveyors, rollers, etc.), semiconductor manufacturing equipment, electric heating devices, industrial robots (e.g., vertical articulated robots, horizontal articulated robots, etc.), vehicles, etc. However, these are all examples of equipment, and equipment is not limited to these specific examples.

[0010] Here, the abnormality diagnosis device 10 according to the first and second embodiments executes a "model creation process" that creates a normal model from data representing the sound of an apparatus under normal conditions (hereinafter also referred to as "sound data"), and an "abnormality diagnosis process" that diagnoses whether an abnormality has occurred based on the sound data of the apparatus to be diagnosed for abnormality and the normal model. Note that, although the following describes a case in which the abnormality diagnosis device 10 executes both the model creation process and the abnormality diagnosis process, the model creation process and the abnormality diagnosis process may be executed by different devices. For example, in an abnormality diagnosis system including a model creation device and an abnormality diagnosis device, the model creation device may execute the model creation process, and the abnormality diagnosis device may execute the abnormality diagnosis process. However, the model creation process needs to be executed before the abnormality diagnosis process.

[0011] [First embodiment] The first embodiment will be described below.

[0012] <Example of Hardware Configuration of the Abnormality Diagnostic Device 10 According to the First Embodiment> An example of the hardware configuration of the abnormality diagnosis device 10 according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the hardware configuration of the abnormality diagnosis device 10 according to the first embodiment.

[0013] 1, the abnormality diagnosis device 10 according to the first embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.

[0014] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the abnormality diagnosis device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.

[0015] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0016] The communication I / F 104 is an interface for communicating with a sensor (e.g., a microphone) that records sound generated by the device, and for communicating with other devices. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device (storage device) such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. The processor 108 is, for example, one of various arithmetic devices such as a CPU (Central Processing Unit).

[0017] 1 is an example, and the fault diagnosis device 10 may have other hardware configurations. For example, the fault diagnosis device 10 may have multiple auxiliary storage devices 107 and multiple processors 108, or may have various types of hardware other than the hardware shown in the figure.

[0018] <Example of functional configuration of the abnormality diagnosis device 10 according to the first embodiment> An example of the functional configuration of the abnormality diagnosis device 10 according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the functional configuration of the abnormality diagnosis device 10 according to the first embodiment.

[0019] As shown in FIG. 2 , the abnormality diagnosis device 10 according to the first embodiment includes a first input unit 201, a first conversion unit 202, a second conversion unit 203, a third conversion unit 204, a storage unit 205, an abnormality diagnosis unit 206, and an output unit 207. These units are realized, for example, by a process in which one or more programs installed in the abnormality diagnosis device 10 are executed by a processor 108 or the like. The abnormality diagnosis device 10 according to the first embodiment also includes a normal data storage unit 208 and a normal model storage unit 209. These storage units are realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, at least one of the normal data storage unit 208 and the normal model storage unit 209 may be realized by a storage area of ​​a storage device (e.g., a storage device provided in a database server) or the like that is communicatively connected to the abnormality diagnosis device 10.

[0020] In the model creation process, the first input unit 201 inputs a normal data set stored in the normal data storage unit 208. The normal data set is a set of sound data of a device under normal conditions. In addition, in the abnormality diagnosis process, the first input unit 201 inputs sound data of a device to be diagnosed for abnormality (hereinafter also referred to as "diagnosis target data"). Note that the sound data is data obtained by sampling and quantizing sound waves at predetermined sampling periods, and is expressed as time-series data of values ​​indicating the magnitude (e.g., amplitude) of the sound signal.

[0021] In the model creation process, the first conversion unit 202 performs a short-time Fourier transform (STFT) on sound data included in the normal data set, and converts the sound data into a first matrix. Similarly, in the abnormality diagnosis process, the first conversion unit 202 performs a short-time Fourier transform on diagnosis target data, and converts the diagnosis target data into a first matrix. The first matrix is ​​a matrix whose elements are the Fourier spectrum of the sound signal, with numbers representing frequency components as rows and numbers representing windows used in the short-time Fourier transform as columns. The Fourier spectrum is composed of an amplitude component and a phase component.

[0022] In the model creation process, the second conversion unit 203 converts a first matrix into a second matrix using a conversion method that utilizes a technique called harmonic separation. Specifically, in the model creation process, the second conversion unit 203 converts the first matrix into a second matrix by, for example, applying a median filter to the amplitude components of each row of the first matrix in the row direction. Similarly, in the abnormality diagnosis process, the second conversion unit 203 converts the first matrix into a second matrix using a conversion method that utilizes harmonic separation by, for example, applying a median filter to the amplitude components of each row of the first matrix in the row direction.

[0023] In the model creation process, the third conversion unit 204 converts the second matrix into a third matrix composed of values ​​representing classes classified by a method called chroma. Similarly, in the abnormality diagnosis process, the third conversion unit 204 converts the second matrix into a third matrix. Chroma is a method of classifying pitch into 12 classes (note names) and expressing pitch by these classes. In chroma, pitch is classified into 12 classes (note names): {C, C#, D, D#, E, F, F#, G, G#, A, A#, B}.

[0024] In the model creation process, the saving unit 205 saves the third matrix in the normal model storage unit 209 as a normal model.

[0025] In the abnormality diagnosis process, the abnormality diagnosis unit 206 compares the third matrix with the normal model stored in the normal model storage unit 209 to diagnose whether an abnormality has occurred in the device that is the target of abnormality diagnosis.

[0026] In the abnormality diagnosis process, the output unit 207 outputs the result of the abnormality diagnosis by the abnormality diagnosis unit 206 (hereinafter also referred to as the "diagnosis result") to a predetermined output destination.

[0027] The normal data storage unit 208 stores the normal data set. The normal model storage unit 209 stores the normal model saved by the saving unit 205.

[0028] <Model Creation Process According to the First Embodiment> The model creation process according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the model creation process according to the first embodiment.

[0029] The first input unit 201 inputs a normal data set stored in the normal data storage unit 208 (step S101). Hereinafter, the normal data set is defined as D={a (d) |d=1, ,|D|}, where |D| is the total number of sound data included in the normal data set D, a (d) is the d-th sound data included in the normal data set D. Also, the d-th sound data a (d) is a (d) =(a1 (d) ,···,a T (d) ) where T is the dth sound data a (d) It is assumed that the normal data set D is stored in the normal data storage unit 208 in advance.

[0030] The first input unit 201 receives one sound data a from the normal data set D. (d) ∈D is obtained (step S102).

[0031] The first conversion unit 202 converts the sound data a acquired in step S102 into (d) ∈D, and perform a short-time Fourier transform (STFT) on the sound data a (d) into the first matrix X (d) (Step S103). Here, the number of the window used in the STFT is n, the total number of windows is N, the number of the frequency component in the STFT is m, and the total number of frequency components is M. In this case, the first matrix is ​​an M×N complex value matrix X (d) =(X m,n (d) ) where X m,n (d) is X m,n (d) =|X m,n(d) |exp(iθ m,n (d) ) is expressed as |X m,n (d) | is the amplitude component, θ m,n (d) is called the phase component.

[0032] In the following, for simplicity, it is assumed that the smaller the number m, the higher the frequency component, and each frequency component is predetermined. That is, the frequency component with the number m is referred to as f m Then, f1>f2>>f M and each f m The value of is determined in advance. Also, the smaller the number n, the smaller the window's time component. That is, the window with number n is W n Then, the window W n Comparing the start time of the time series data included in W1 <W2<···<W N It is assumed that:

[0033] The second transform unit 203 transforms the first matrix X (d) into the second matrix Y (d) (Step S104). That is, the second conversion unit 203 converts, for example, the first matrix X (d) In each row of , each element X m,n (d) Amplitude component of |X m,n (d) The first matrix X is obtained by applying a median filter row-wise to | (d) into the second matrix Y (d) Specifically, the second conversion unit 203 converts the first matrix X into the following matrix X by, for example, the following steps 1-1 to 1-3. (d) into the second matrix Y (d) Convert to.

[0034] Step 1-1: The second conversion unit 203 converts X m,n (d) ←|X m,n (d) |Let's say.

[0035] Step 1-2: The second conversion unit 203 performs the calculation (X m,1 (d) ,···,X m,N (d) ) As a specific example, it is possible to apply the median filter to n=2 in order from n=N-1, with w=3 and s=1.

[0036] Step 1-3: The second transform unit 203 converts the first matrix X after applying the median filter (d) into the second matrix Y (d) This gives the second matrix Y (d) is a matrix composed of amplitude components with noise removed, making it easy to distinguish between high-pitched and low-pitched sounds. (d) =(Y m,n (d) ) where Y (d) is an M×N real-valued matrix.

[0037] The third transformation unit 204 transforms the second matrix Y (d) into the third matrix Z (d) Specifically, the third conversion unit 204 converts the second matrix Y (d) into the third matrix Z (d) Convert to.

[0038] Step 2-1: The third conversion unit 204 converts Y m,n (d) ←|Y m,n (d) | 2 Let's say.

[0039] Step 2-2: The third conversion unit 204 converts a predetermined frequency f c, and calculates the frequency bands of each of the 12 classes {C, C#, D, D#, E, F, F#, G, G#, A, A#, B} from the frequency f. That is, the third conversion unit 204 calculates the frequency of each class so that the frequency ratio between competing classes (note names) is 1:(2 to the power of 1 / 12) with a predetermined frequency as the frequency of C. Then, the third conversion unit 204, for example, defines the range equal to or higher than the frequency of C and lower than the frequency of C# as the C frequency band, the range equal to or higher than the frequency of C# and lower than the frequency of D as the C# frequency band, ..., the range equal to or higher than the frequency of A# and lower than the frequency of B as the A# frequency band, and the range equal to or higher than the frequency of B and lower than twice the frequency of C as the B frequency band. Note that, c can be determined arbitrarily, for example, f c = 261.6 Hz (i.e., the frequency at which the frequency of A becomes 440 Hz).

[0040] Below, the frequency band of B is b1, the frequency band of A# is b2, ..., the frequency band of C# is b 11 , the frequency band of C is b 12 In addition, assuming that M is a multiple of 12, let L:=M / 12. If M is not a multiple of 12, for example, the rows with all values ​​of 0 are added to the second matrix Y (d) and make M a multiple of 12.

[0041] Step 2-3: The third conversion unit 204 converts f (i-1)L+1 ~f iL A b i For each n=1, ,N and each i=1, ,12, we use a linear mapping F that projects onto (Y (i-1)L+1,n (d) ,···,Y iL,n (d) ) to b i That is, the third transformation unit 204 projects F(Y (i-1)L+1,n (d) ,···,Y iL,n (d) ) is calculated.

[0042] Step 2-4: The third conversion unit 204 converts Z i,n (d) ←F(Y (i-1)L+1,n (d) ,···,Y iL,n (d) ) as a 12×N real-valued matrix Z (d) =(Z i,n (d) ) is created. i,n (d) may be called, for example, a "chroma value."

[0043] Step 2-5: The third transformation unit 204 transforms all elements Z i,n (d) is normalized to a value between 0 and 1, and the normalized Z i,n (d) A 12×N real-valued matrix Z consisting of (d) =(Z i,n (d) ) is the third matrix. Note that the third matrix Z (d) Each element of Z i,n (d) represents the pitch feature when the pitch of a sound is expressed by chroma. Also, the third matrix Z (d) The 12-dimensional vectors represented by each column may be called, for example, "chroma vectors." Each element of the chroma vector represents a score (or a percentage) that indicates how much of the frequency of the class (note name) corresponding to that element is contained in the original sound.

[0044] The storage unit 205 stores the third matrix Z created in step S105. (d) The normal model ML (d) and stores it in the normal model storage unit 209 (step S106).

[0045] The first input unit 201 receives the acquired sound data a (d) is excluded from the normal data set D (step S107). That is, the first input unit 201 extracts D←D\{a (d)}

[0046] The first input unit 201 determines whether or not sound data exists in the normal data set D (step S108). That is, the first input unit 201 determines whether or not D=φ.

[0047] If it is determined in step S108 above that D=φ is not satisfied, steps S102 to S107 above are executed again.

[0048] On the other hand, if it is determined in step S108 that D=φ, the model creation process ends. (1) ,···,ML (|D|) is stored in the normal model storage unit 209.

[0049] <Abnormality Diagnosis Processing According to the First Embodiment> The abnormality diagnosis process according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the abnormality diagnosis process according to the first embodiment.

[0050] The first input unit 201 inputs given diagnostic object data (step S201). Hereinafter, the diagnostic object data is a=(a1, . . . , a T ) shall be expressed as

[0051] The first transformation unit 202 performs STFT on the diagnostic object data a acquired in the above step S201, similarly to step S103 in FIG. 3, and converts the diagnostic object data a into a first matrix X=(X m,n ) (step S202). The first matrix X is an M×N complex value matrix.

[0052] The second transform unit 203 transforms the first matrix X into the second matrix Y using a transform method that uses harmonic component separation (step S203). That is, the second transform unit 203 transforms, for example, each element X m,n Amplitude component of |X m,nBy applying a median filter to | in the row direction, the first matrix X is converted into the second matrix Y. Note that the second conversion unit 203 converts the first matrix X into the second matrix Y, for example, by a method similar to the above steps 1-1 to 1-3.

[0053] The third conversion unit 204 converts the second matrix Y into the third matrix Z (step S204). The third conversion unit 204 converts the second matrix Y into the third matrix Z in the same manner as in steps 2-1 to 2-5 above. The third matrix Z is a 12×N real-valued matrix. The third matrix Z may be called, for example, a "diagnosis target model."

[0054] The abnormality diagnosis unit 206 compares the third matrix Z obtained in step S204 with the normal model ML stored in the normal model storage unit 209. (1) ,···,ML (|D|) (Step S205) Specifically, the abnormality diagnosis unit 206 calculates the similarity through the following steps 3-1 to 3-3.

[0055] Step 3-1: The abnormality diagnosis unit 206 creates a 12N-dimensional vector V by vertically arranging each column of the third matrix Z in order from the first column.

[0056] Step 3-2: The abnormality diagnosis unit 206 calculates the normal model ML for each of d=1, . . . , |D|. (d) A 12N-dimensional vector V (d) Create a.

[0057] Step 3-3: The abnormality diagnosis unit 206 calculates V and V for each of d=1, , |D|. (d) The final similarity is calculated by taking the average value of the similarity between V and V. (d) The similarity between vectors can be determined by any scale that measures the similarity between vectors. For example, cosine similarity, similarity based on Euclidean distance, etc. can be used.

[0058] The abnormality diagnosis unit 206 determines whether the similarity calculated in the above step S205 is less than a predetermined threshold value (step S206).

[0059] If it is not determined in step S206 that the similarity is less than the threshold value, the abnormality diagnoser 206 diagnoses the condition as "normal" (step S207).

[0060] On the other hand, if it is determined in step S206 above that the similarity is less than the threshold value, the abnormality diagnosing unit 206 diagnoses an "abnormality" (step S208).

[0061] The output unit 207 outputs the result of the abnormality diagnosis in step S207 or step S208 to a predetermined output destination (step S209). Note that the output destination can be any output destination, and examples thereof include the display device 102 such as a display, a storage area of ​​the auxiliary storage device 107, a terminal used by an operator, and a device including an alarm that notifies an alert.

[0062] <Modification of the first embodiment> In the first embodiment, a median filter is applied in step S104 of FIG. 3 and step S203 of FIG. 4, but application of a median filter is not essential, and it is not necessary to apply a median filter in step S104 of FIG. 3 and step S203 of FIG. 4.

[0063] <Comparison between normal and abnormal conditions> A comparison example between the first matrix and the third matrix created from the sound data of the device in a normal state and the first matrix and the third matrix created from the sound data of the device in an abnormal state will be described with reference to Fig. 5. Fig. 5 is a diagram showing a comparison example between a normal state and an abnormal state.

[0064] Figure 5(A) shows an example of the visualization of the first matrix, Figure 5(B) shows an example of the visualization of the third matrix when the median filter is not applied, and Figure 5(C) shows an example of the visualization of the third matrix when the median filter is applied.

[0065] As shown by reference numeral 1000 in Fig. 5(B), even without applying a median filter, the high-pitched sound can be distinguished to some extent from other sounds. Furthermore, as shown by reference numeral 2000 in Fig. 5(C), when a median filter is applied, the high-pitched sound can be distinguished more clearly than in Fig. 5(B). This allows the fault diagnosis device 10 according to the first embodiment to detect high-pitched sounds that suggest that the device is operating abnormally, thereby achieving high fault diagnosis performance.

[0066] [Second embodiment] A second embodiment will be described below. In the second embodiment, an optimal conversion method is determined from various conversion methods using audio processing (e.g., a conversion method using harmonic component separation), and then a first matrix is ​​converted into a second matrix using the optimal conversion method in model creation and abnormality diagnosis.

[0067] In the second embodiment, differences from the first embodiment will be mainly described, and descriptions of components that are the same as those in the first embodiment will be omitted.

[0068] <Example of Hardware Configuration of Abnormality Diagnostic Device 10 According to Second Embodiment> The hardware configuration of the abnormality diagnostic device 10 according to the second embodiment may be the same as that of the first embodiment, and therefore a description thereof will be omitted.

[0069] <Example of functional configuration of the abnormality diagnosis device 10 according to the second embodiment> An example of the functional configuration of the abnormality diagnosis device 10 according to the second embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the functional configuration of the abnormality diagnosis device 10 according to the second embodiment.

[0070] As shown in Fig. 6, the abnormality diagnosis device 10 according to the second embodiment further includes a second input unit 210 and a conversion method determination unit 211. These units are realized, for example, by processing in which one or more programs installed in the abnormality diagnosis device 10 are executed by the processor 108 or the like. The abnormality diagnosis device 10 according to the second embodiment also includes a sound data storage unit 212. The sound data storage unit 212 is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, the sound data storage unit 212 may also be realized by a storage area of ​​a storage device (e.g., a storage device provided in a database server) or the like that is communicatively connected to the abnormality diagnosis device 10.

[0071] The second input unit 210 inputs a sound data set stored in the sound data storage unit 212. The sound data set is a set of sound data of a device. Note that the sound data included in the sound data set may be assigned a label indicating whether it is normal or abnormal.

[0072] The conversion method determination unit 211 determines the optimum conversion method from among a plurality of conversion methods that utilize voice processing, using the sound data set input by the second input unit 210. Hereinafter, the total number of conversion methods is K, and the kth (1≦k≦K) conversion method is referred to as P k Conversion method P k Examples of such conversion methods include "a conversion method that does not use audio processing (identity conversion)," "a conversion method that uses harmonic component separation," "a conversion method that uses percussive separation," "a conversion method that uses noise reduction," and "a conversion method that uses vocal separation." Note that harmonic component separation, percussive separation, and noise reduction are all examples of audio processing. However, it goes without saying that these are just examples.

[0073] The sound data storage unit 212 stores a sound data set.

[0074] <Conversion Method Determination Process According to the Second Embodiment> The conversion method determination process according to the second embodiment will be described with reference to Fig. 7. Fig. 7 is a flowchart showing an example of the conversion method determination process according to the second embodiment.

[0075] The second input unit 210 inputs a sound data set stored in the sound data storage unit 212 (step S301). Hereinafter, the sound data set is defined as D'={a (d) |d=1,...,|D'|}, where |D'| is the total number of sound data included in the sound data set D', a (d) is the d-th sound data included in the sound data set D'. Also, the d-th sound data a (d) is a (d) =(a1 (d) ,···,a T (d) ) where T is the dth sound data a (d) The length of the sound data a (d) The sound data a (d) A label y indicating whether the (d) may be given.

[0076] Below is the sound data a (d) ∈D' with label y (d) If no is given, all sound data a (d) ∈D' is normal data. On the other hand, sound data a (d) ∈D' with label y (d) If is given, then there is at least one label y (d) is a label that represents an anomaly.

[0077] The second input unit 210 receives one sound data a from the sound data set D′. (d) ∈D' is obtained (step S302).

[0078] The conversion method determination unit 211 converts the sound data a acquired in the above step S302 into a (d) ∈D', and the sound data a (d) into the first matrix X(d) (step S303).

[0079] The conversion method determination unit 211 determines the conversion method P k For each (1≦k≦K), the conversion method P k Using the first matrix X (d) into the second matrix Y k (d) (step S304).

[0080] As an example, P k is an identity transformation, the transformation method determination unit 211 determines the first matrix X (d) into the second matrix Y k (d) This can be done as follows.

[0081] Another example is P k is a conversion method using impact sound separation, the conversion method determination unit 211 may, for example, (d) In each column of m,n (d) Amplitude component of |X m,n (d) The first matrix X is obtained by applying a median filter row-wise to | (d) into the second matrix Y k (d) Just convert it to

[0082] By the above step S303, the conversion method P k For each (1≦k≦K), the conversion method P k The second matrix Y corresponding to k (d) (1≦k≦K) is obtained.

[0083] The conversion method determination unit 211 determines the second matrix Y k (d) For each (1≦k≦K), the second matrix Y k (d) into the third matrix Z k (d) (Step S305). That is, the conversion method determination unit 211 converts the second matrix Y k(d) For each (1≦k≦K), the second matrix Y k (d) into the third matrix Z k (d) This converts the conversion method P k For each (1≦k≦K), the conversion method P k The third matrix Z corresponding to k (d) is obtained.

[0084] The second input unit 210 receives the acquired sound data a (d) is excluded from the sound data set D' (step S306). That is, the second input unit 210 extracts D'←D'\{a (d)}

[0085] The second input unit 210 determines whether or not sound data exists in the sound data set D' (step S307). That is, the second input unit 210 determines whether or not D'=φ.

[0086] If it is determined in step S307 that D'=φ is not true, steps S302 to S306 are executed again. k For each (1≦k≦K), each sound data a (d) The third matrix Z for ∈D' k (d) is obtained.

[0087] On the other hand, if it is determined in step S307 that D'=φ, the second input unit 210 uses the sound data set D'={a (d) Whether abnormal data exists in |d=1, ,|D'| (i.e., whether each sound data a (d) ∈D' with label y (d) It is determined whether or not the item is assigned (step S308).

[0088] If it is determined in step S308 above that there is no abnormal data in the sound data set D' (i.e., sound data a (d)∈D' with label y (d) If the sound data a is not assigned to the sound data a, the conversion method determination unit 211 selects the sound data a that can be regarded as pseudo-abnormal. (d) After identifying ∈D', sound data a that can be considered as pseudo-anomalous (d) ∈D' has an anomaly label y (d) , other sound data a (d) ∈D' has a normal label y (d) This is because when determining the optimal conversion method in step S310 described later, it is necessary to evaluate an index value that indicates the prediction accuracy, and for this evaluation, the sound data a (d) For ∈D', label (d) As an example, the conversion method determination unit 211 selects sound data a that can be considered pseudo-abnormal by the following steps 4-1 to 4-4. (d) After identifying ∈D', sound data a that can be considered as pseudo-anomalous (d) For ∈D', the label y (d) , other sound data a (d) For ∈D', the label y (d) can be given.

[0089] Step 4-1: The conversion method determination unit 211 converts the sound data a (d) For each ∈D', the sound data a (d) ∈D' and other sound data a (d') ∈D'\{a (d)} and the distance Δ(a (d) ,a (d') ) and calculate the distance Δ(a (d) ,a (d') ) is used as the (d,d') component to create a distance matrix Δ of |D'|×|D'|.

[0090] Step 4-2: The conversion method determination unit 211 calculates the distance Δ(a (d) ,a (d') ) sum Δ (d) (d=1, ,|D'|) is calculated.

[0091] Step 4-3: The conversion method determination unit 211 determines Δ (d) κ Δ (d) Identify.

[0092] Step 4-4: The conversion method determination unit 211 determines the κ Δ (d) Sound data a corresponding to (d) Let ∈D' be sound data that can be assumed to be pseudo-anomalous. This makes it possible to assume that sound data that is far from other sound data is pseudo-anomalous based on the distance matrix Δ.

[0093] The conversion method determination unit 211 determines whether each conversion method P k (1≦k≦K) and then determine the optimal conversion method (step S310). As an example, the conversion method determination unit 211 determines the optimal conversion method P k All that is required is to evaluate (1≦k≦K) and determine the optimal conversion method.

[0094] Step 5-1: The conversion method determination unit 211 determines the conversion method P k For each (1≦k≦K), the third matrix Z k (d) (d=1, ,|D'|) and a predetermined prediction method (or a discrimination method, classification method, etc.) to obtain the third matrix Z k (d) Sound data a corresponding to (d) The label y assigned to (d) An index value representing the prediction accuracy of is calculated. Here, for example, an F1 score or the like may be used as the index value. The prediction method is not limited to a specific prediction method, but for example, a machine learning method (e.g., neural network, support vector machine, etc.) that realizes two-class classification using a linear or nonlinear method may be used. Note that the third matrix Z k (d) is a 12xN matrix, so it may be vectorized as needed when using a prediction method.

[0095] Step 5-2: The conversion method determination unit 211 selects the conversion method P with the best index value calculated in the above step 5-1 (i.e., the highest prediction accuracy). k is determined as the optimal conversion method P.

[0096] The conversion method determination unit 211 converts the sound data a (d) ∈D', the normal label y (d) Sound data a with (d) ∈D′ is stored in the normal data storage unit 208 (step S311). As a result, the normal data set D is stored in the normal data storage unit 208.

[0097] <Model Creation Process According to the Second Embodiment> The model creation process according to the second embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the model creation process according to the second embodiment.

[0098] Steps S401 to S403 may be similar to steps S101 to S103 in FIG. 3, respectively, and therefore a description thereof will be omitted.

[0099] Following step S404, the second conversion unit 203 converts the first matrix X (d) into the second matrix Y (d) (step S403).

[0100] The subsequent steps S404 to S408 may be similar to steps S104 to S108 in FIG. 3, respectively, and therefore a description thereof will be omitted.

[0101] <Abnormality Diagnosis Processing According to Second Embodiment> The abnormality diagnosis process according to the second embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing an example of the abnormality diagnosis process according to the second embodiment.

[0102] Steps S501 and S502 may be similar to steps S201 and S202 in FIG. 4, respectively, and therefore a description thereof will be omitted.

[0103] Following step S502, the second conversion unit 203 converts the first matrix X into the second matrix Y using the conversion method P determined by the conversion method determination unit 211 (step S503).

[0104] The subsequent steps S504 to S509 may be similar to steps S204 to S209 in FIG. 4, respectively, and therefore a description thereof will be omitted.

[0105] [summary] As described above, the abnormality diagnosis device 10 according to the first and second embodiments focuses on the fact that high-pitched sounds are likely to occur when equipment is operating abnormally, and performs abnormality diagnosis using pitch features, which are features related to the pitch of sounds. Moreover, the abnormality diagnosis device 10 according to the first and second embodiments uses a method called chroma, which classifies sounds into 12 classes based on their relative pitch differences, to classify the frequency components of the sounds into 12 classes, and uses these values ​​as pitch features. This is expected to enable highly accurate abnormality diagnosis.

[0106] In addition to the above, the abnormality diagnosis device 10 according to the second embodiment can determine the optimal conversion method from among various conversion methods using voice processing, and then convert the first matrix into the second matrix using this conversion method, which is expected to enable even more accurate abnormality diagnosis.

[0107] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims. [Explanation of symbols]

[0108] 10. Abnormality diagnosis device 101 Input Device 102 Display device 103 External I / F 103a Recording media 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage 108 processors 109 Bus 201 first input unit 202 First conversion unit 203 Second conversion unit 204 Third Conversion Unit 205 Preservation Department 206 Abnormality Diagnosis Unit 207 Output section 208 Normal data storage unit 209 Normal Model Memory 210 second input section 211 Conversion Method Decision Unit 212 Sound data storage unit

Claims

1. a first model creation unit that creates a normal model based on a normal sound representing a sound acquired from a device under normal conditions, the normal model being configured with feature quantities related to the pitch of the normal sound; a second model creation unit that creates a diagnosis target model based on a diagnosis target sound that represents a sound acquired from a target device that is the target of an abnormality diagnosis, the diagnosis target model being configured with feature quantities related to the pitch of the sound included in the diagnosis target sound; a diagnosis unit that diagnoses an abnormality in the target device by comparing the diagnosis target model with the normal model; and The feature amount related to the pitch of the sound is a score that indicates the degree to which sounds of frequencies corresponding to each of 12 classes defined based on chroma are included. Abnormality diagnosis device.

2. The first model creation unit The normal sound is subjected to a short-time Fourier transform to be converted into a Fourier spectrum for each window and each frequency component; creating the normal model based on a linear mapping that projects the frequency components onto frequencies corresponding to each of the 12 classes and on amplitude components included in the Fourier spectrum; The second model creation unit converting the diagnosis target sound into a Fourier spectrum for each window and each frequency component by performing the short-time Fourier transform; The abnormality diagnostic device according to claim 1 , wherein the diagnostic object model is created based on the linear mapping and an amplitude component included in the Fourier spectrum.

3. The first model creation unit applying a median filter to the amplitude component included in the Fourier spectrum; creating the normal model based on the linear mapping and the amplitude component after applying the median filter; The second model creation unit applying a median filter to the amplitude component included in the Fourier spectrum; The abnormality diagnostic device according to claim 2 , wherein the diagnostic object model is created based on the linear mapping and the amplitude component after the median filter is applied.

4. The diagnostic unit 4. The abnormality diagnosis device according to claim 1, wherein the similarity between the diagnosis target model and the normal model is compared, and if the similarity is less than a predetermined threshold, the device diagnoses an abnormality, and if the similarity is equal to or greater than the threshold, the device diagnoses a normality.

5. a determination unit that determines, from among a plurality of sound processes based on the sound acquired from the device, an optimal sound process that maximizes the accuracy of distinguishing between a sound acquired from the device in a normal state and a sound acquired from the device in an abnormal state; The first model creation unit creating the normal model based on the optimal speech processing; The second model creation unit The abnormality diagnosis device according to claim 1 , wherein the diagnosis target model is created based on the optimal voice processing.

6. The first model creation unit performing a short-time Fourier transform on the normal sound to convert it into a Fourier spectrum for each window and each frequency component, and then performing the optimal sound processing; creating the normal model based on a linear mapping that projects the frequency components onto frequencies corresponding to each of the 12 classes and on amplitude components included in the Fourier spectrum; The second model creation unit performing the short-time Fourier transform on the diagnosis target sound to convert it into a Fourier spectrum for each window and each frequency component, and then performing the optimal sound processing; The abnormality diagnostic device according to claim 5 , wherein the diagnostic object model is created based on the linear mapping and an amplitude component included in the Fourier spectrum.

7. The abnormality diagnosis device according to claim 5 or 6, wherein the plurality of audio processing processes include at least one of harmonic component separation, impact sound separation, noise reduction, and vocal separation.

8. a first model creation step of creating a normal model based on a normal sound representing a sound acquired from the device under normal conditions, the normal model being configured with feature quantities related to the pitch of the normal sound; a second model creation step of creating a diagnosis target model based on a diagnosis target sound obtained from a target device that is the target of an abnormality diagnosis, the diagnosis target model being configured with feature quantities related to the pitch of the sound included in the diagnosis target sound; a diagnostic procedure for diagnosing an abnormality in the target device by comparing the diagnosis target model with the normal model; The computer executes The feature amount related to the pitch of the sound is a score that indicates the degree to which sounds of frequencies corresponding to each of 12 classes defined based on chroma are included. Abnormality diagnosis method.

9. a first model creation step of creating a normal model based on a normal sound representing a sound acquired from the device under normal conditions, the normal model being configured with feature quantities related to the pitch of the normal sound; a second model creation step of creating a diagnosis target model based on a diagnosis target sound obtained from a target device that is the target of an abnormality diagnosis, the diagnosis target model being configured with feature quantities related to the pitch of the sound included in the diagnosis target sound; a diagnostic procedure for diagnosing an abnormality in the target device by comparing the diagnosis target model with the normal model; on the computer, The feature amount related to the pitch of the sound is a score that indicates the degree to which sounds of frequencies corresponding to each of 12 classes defined based on chroma are included. program.

Citation Information

Patent Citations

  • Apparatus and program for detecting / predicting failure

    JP2022144024A