Defect Detection Visualization Method and System
By generating time spectrum diagrams and using computer vision technology, the problem that speaker defect detection in the prior art relies on subjective evaluation of the human ear is solved, and more accurate and reliable defect detection is achieved, reducing the risk of occupational injury.
Patent Information
- Application Number
- CN202110495711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-04
- Filing Date
- 2021-05-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-05-07
AI Technical Summary
In the prior art, the detection of assembly defects of speakers depends on subjective assessment of the human ear, is susceptible to the listener's age, emotional and auditory fatigue, and has a risk of occupational injury.
By outputting the test audio signal to the device to be tested, receiving the response signal and performing signal processing, a time spectrum diagram is generated, and computer vision technology is used to visually determine whether the device to be tested complies with the pre-developed auditory standards.
It realizes defect detection that is more accurate than subjective judgment of the human ear, reduces the risk of occupational injury, and improves the reliability and efficiency of the detection.
Smart Images

Figure CN114519691B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a detection technology, and particularly to a method and system for visualizing defect detection. Background Art
[0002] A loudspeaker is a transducer that converts an electrical signal into an acoustic signal, which is widely used in devices such as audio systems and headphones, and its performance affects the use of these devices. Assembly defects of loudspeakers were previously detected by listeners with rich experience at the end of the production line. Such detection requires applying sine logarithmic swept chirps to the loudspeaker and using human hearing to detect and analyze whether the response signal is normal. However, the results detected by such human ear evaluation vary with subjective factors such as the age, mood changes, and auditory fatigue of the listener, and it is easy to cause occupational injuries to the listener. Summary of the Invention
[0003] The present invention provides a method and system for visualizing defect detection, which can use computer vision technology to detect whether a device under test has unacceptable defects in a pre-established auditory standard from a time-frequency spectrum diagram.
[0004] In an embodiment of the present invention, the above method includes the following steps. Output a test audio signal to the device under test, and receive the response signal of the device under test to the test audio signal to generate a received audio signal. Perform signal processing on the received audio signal to generate a time-frequency spectrum diagram, and visually determine whether the device under test has unacceptable defects in a pre-established auditory standard according to the time-frequency spectrum diagram.
[0005] In an embodiment of the present invention, the above system includes a signal output device, a microphone, an analog-to-digital converter, and a processing device. The signal output device is used to output a test audio signal to the device under test. The microphone is used to receive the response signal of the device under test to the test audio signal. The analog-to-digital converter is used to convert the response signal into a received audio signal. The processing device is used to perform signal processing on the received audio signal to generate a time-frequency spectrum diagram, and visually determine whether the device under test has unacceptable defects in a pre-established auditory standard according to the time-frequency spectrum diagram. Brief Description of the Drawings
[0006] Figure 1 It is a block diagram of a defect detection system shown according to an embodiment of the present invention.
[0007] Figure 2 It is a flowchart of a method for visualizing defect detection shown according to an embodiment of the present invention.
[0008] Figure 3 It is a schematic diagram of a time-frequency spectrum diagram shown according to an embodiment of the present invention.
[0009] Figure 4 It is a functional block flowchart for constructing a classifier as shown in an embodiment of the present invention.
[0010] Figure 5 It is a functional block flowchart for a method of obtaining spatial features as shown in an embodiment of the present invention.
[0011] Figure 6 It is a functional block flowchart for a method of visualizing defect detection as shown in an embodiment of the present invention.
[0012] Figure 7 It is a schematic diagram of a time-frequency spectrum as shown in an embodiment of the present invention.
[0013] Figure 8 It is a schematic diagram of a time-frequency spectrum with acceptable abnormal sounds and a time-frequency spectrum with unacceptable abnormal sounds as shown in an embodiment of the present invention.
[0014] Figure 9 It is a block diagram of a defect detection system as shown in an embodiment of the present invention.
[0015] Figure 10 It is a flowchart of a method of visualizing defect detection as shown in an embodiment of the present invention.
[0016] Figure 11 It is a functional diagram for converting a time-frequency spectrum into a projection curve as shown in an embodiment of the present invention.
[0017] Figure 12 It is a schematic diagram for dividing a projection curve as shown in an embodiment of the present invention.
[0018] Figure 13 It is a data diagram for hierarchical detection of abnormal sounds as shown in an embodiment of the present invention. Detailed implementation manners
[0019] Some embodiments of the present invention will be described in detail with reference to the accompanying drawings hereinafter. For the component symbols cited in the following descriptions, when the same component symbols appear in different drawings, they will be regarded as the same or similar components. These embodiments are only a part of the present invention and do not disclose all the implementable manners of the present invention. More precisely, these embodiments are only examples of the methods and systems in the claims of the present invention.
[0020] Figure 1 It is a block diagram of a defect detection system as shown in an embodiment of the present invention, but this is only for convenience of description and is not intended to limit the present invention. First Figure 1 introduce all the components and configuration relationships in the defect detection system, and the detailed functions will be disclosed together with Figure 2 simultaneously.
[0021] Please refer to Figure 1 , the defect detection system 100 includes a signal output device 110, a microphone 120, an analog-to-digital converter 130, and a processing device 140, which are used to detect whether there are defects in the device under test T.
[0022] The signal output device 110 is used to output a test audio signal to the device under test T. It can be, for example, an electronic device with a digital audio output interface, and output the test audio signal to the device under test T in a wireless or wired manner. The microphone 120 is used to pick up the response of the device under test T to the test audio signal. It can be set near the device under test T or at the optimal sound pickup position relative to the device under test T. The analog-to-digital converter 130 is connected to the microphone 120 and is used to convert the analog sound signal received by the microphone 120 into a digital sound signal.
[0023] The processing device 140 is connected to the analog-to-digital converter 130 and is used to process the digital sound signal received from the analog-to-digital converter 130 to detect whether there are defects in the device under test U. The processing device 140 includes a memory and a processor. The memory can be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk, or other similar devices, integrated circuits, and their combinations. The processor can be, for example, a central processing unit (CPU), an application processor (AP), or other programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), or other similar devices, integrated circuits, and their combinations.
[0024] It must be noted that in the embodiment, the signal output device 110, the microphone 120, the analog-to-digital converter 130, and the processing device 140 can belong to four independent devices. In the embodiment, the signal output device 110 and the processing device 140 can be integrated into the same device, and the processing device 140 can control the output of the signal output device 110. In the embodiment, the signal output device 110, the microphone 120, the analog-to-digital converter 130, and the processing device 140 can also be a single all-in-one computer system. The present invention does not impose any limitations on the integration of the signal output device 110, the microphone 120, the analog-to-digital converter 130, and the processing device 140. As long as the system includes these devices, it belongs to the scope of the defect detection system 100.
[0025] The following are examples to illustrate the detailed steps of the defect detection method performed by the defect detection system 100 on the device under test T. In the following examples, an electronic device with a speaker will be used as the device under test T for illustration, and the defect detected by the defect detection system 100 is the assembly abnormal sound (rub and buzz) of the device under test T.
[0026] Figure 2 It is a flowchart of the defect detection visualization method shown in the embodiments of the present invention. Figure 2 The process will be performed by Figure 1 the defect detection system 100.
[0027] Please also refer to Figure 1 and Figure 2 , the signal output device 110 will output a test audio signal to the device under test T (step S202), the microphone 120 will receive the response signal of the device under test T to the test audio signal (step S204), and the analog-to-digital converter 130 will convert the response signal into a received audio signal (step S206). Here, the audio range of the test audio signal can be 1K to 20 Hz, where the amplitude from 1K to 500 Hz is -25 dB, the amplitude from 500 Hz to 300 Hz is -15 dB, and the amplitude from 300 Hz to 20 Hz is -8 dB. However, since the assembly abnormal sound will resonate with specific frequency points of the test audio signal, and in order to avoid the resonance of non-assembly abnormal sounds affecting the detection of abnormal sounds (such as button resonance), the audio range and amplitude of the test audio signal will be adjusted according to the different device under test T. The device under test T will generate a response signal to the test audio signal, and the microphone 120 will receive the response signal from the device under test T. Then, the analog-to-digital converter 130 will perform analog-to-digital conversion on the analog response signal to generate a digital response signal (hereinafter referred to as "received audio signal").
[0028] The processing device 140 will perform signal processing on the received audio signal to generate a time-frequency spectrum diagram (step S208), and visually judge whether the device under test T has a defect according to the time-frequency spectrum diagram (step S210). The processing device 140 can perform a Fast Fourier Transform (FFT) on the received audio signal to generate a time-frequency spectrum diagram. The reason for converting the received audio signal into a time-frequency spectrum diagram here is that the abnormal sound does not have significant features in the received audio signal, but the abnormal sound has time continuity when resonating with the test audio signal. Therefore, if the time-domain signal is converted into a time-frequency spectrum diagram, the abnormal sound features will appear as time continuous and energy clustering in the time-frequency spectrum diagram, so as to use computer vision technology to achieve defect detection of the device under test.
[0029] Taking Figure 3 Taking the schematic diagram of the time-frequency spectrum shown in the embodiment of the present invention as an example, the time-frequency spectrum 310 corresponds to a sound signal without abnormal sound, and the time-frequency spectrum 320 corresponds to a sound signal with abnormal sound. It must be noted that those of ordinary skill in the art should understand that the time-frequency spectrum represents the distribution of signal intensity over time and frequency, and the time-frequency spectra 310 and 320 are only simply illustrated by curves to show the obvious signal intensity for explanation. Here, the sound signal with abnormal sound has the characteristics of time continuity and energy clustering in the time-frequency spectrum 320, such as the abnormal sound feature RB. Therefore, if the processing device 140 uses computer vision technology to analyze the time-frequency spectrum, it can determine whether the device under test T produces abnormal sound due to assembly defects.
[0030] In the following embodiments, an image classifier will be used for image recognition. Therefore, before the processing device 140 determines whether the device under test T has defects, a trained classifier will be obtained. The classifier here can be trained by the processing device 140 itself, or a trained classifier can be obtained from other processing devices. The present invention is not limited thereto.
[0031] Figure 4 FIG. is a functional block flow chart for constructing a classifier according to the embodiment of the present invention. In the following description, a classifier will be constructed in a manner similar to the processing device 140 (hereinafter referred to as the "training system").
[0032] Please refer to Figure 4 , first, the training system will collect a plurality of training data 402. The training data here can be N1 non-defective training objects and N2 defective training objects, which respectively generate N1 non-defective training sound samples and N2 defective training sound samples in a manner similar to steps S202 to S204, where these N1 + N2 training objects are the same as the device under test T, but have been pre-tested for defects.
[0033] Next, the training system will convert the training data into a time-frequency spectrum 404. To reduce the computational complexity and to avoid low-frequency noise and high-frequency noise images, the training system will select a preset frequency range, such as 3K to 15K Hz, as the detection area. Taking Figure 3 as an example, the area 315 is the detection area of the time-frequency spectrum 310, and the area 325 is the detection area of the time-frequency spectrum 320. For the convenience of explanation, hereinafter, the detection area in the time-frequency spectrum corresponding to the non-defective training sound sample will be called the "non-defective detection area image", and the detection area in the time-frequency spectrum corresponding to the defective training sound sample will be called the "defective detection area image".
[0034] After that, the training system will obtain the feature values corresponding to different regions in each defective detection area image and each non-defective detection area image, and obtain the texture correlation 406 between each defective detection area image and non-defective detection area image and the reference model 408, so as to use them as spatial features 410 to train the classifier 412, and then generate a classifier 414 for detecting whether the device under test T has defects.
[0035] Here, the training system will first perform image segmentation on all defective detection area images and non-defective detection area images to generate multiple sub-blocks of the same size (for example, a pixel size of 40×200). In this embodiment, if the size of the sub-block is too large, the proportion of abnormal sound features will be reduced, and if the size of the sub-block is too small, the abnormal sound features will not be covered, affecting the subsequent identification results. Therefore, the training system can, for example, Figure 5 Obtain the spatial features of each defective detection area image and non-defective detection area image according to the functional block flow chart shown in the embodiment of the present invention.
[0036] Please refer to Figure 5 , the training system will perform image pyramid processing H (image pyramid) on each defective detection area image and non-defective detection area image respectively to generate images of different scales. This embodiment will have two scales, namely the original image size and 1 / 4 of the original image size (reducing the length and width of the original image to 1 / 2 of the original respectively). Here, one defective detection area image will be used to Figure 5 Describe the process. Those of ordinary skill in the art can analogize the processing methods of other defective detection area images and non-defective detection area images. Assume that T1 is one of the defective detection area images, with a pixel size of 1000×800. T11 is one of the sub-blocks after image segmentation (hereinafter referred to as the "training sub-block"), with a pixel size of 40×200. On the other hand, T0 is the image generated after T1 undergoes image pyramid processing (shrinking processing), with a pixel size of 500×400. T01 is one of the training sub-blocks after image segmentation, and its pixel size will be the same as that of the training sub-block T11, that is, a pixel size of 40×200.
[0037] Next, the training system will perform feature extraction FE on each training sub-block segmented from the defect-free detection area images and defect detection area images of different scales. In this embodiment, the training system can, for example, calculate at least one of the standard deviation σ (standard deviation) and histogram skewness k (Kurtosis) of the pixel values of each training sub-block at each scale as the feature value of each training sub-block. However, the present invention is not limited thereto. In addition, in order to improve the difference between defect-free and defective, the training system can also generate a reference model associated with defect-free based on N1 defect-free detection area images. For example, the training system can average the pixel values of N1 defect-free detection area images of the same scale to obtain the reference model. Therefore, each scale has its corresponding reference model. In this embodiment, the training system will generate a reference model R1 corresponding to image T1 and a reference model R0 corresponding to image T0. Here, image T1 and reference model Rl have the same scale, so the training sub-blocks in image T1 will find corresponding sub-blocks (hereinafter referred to as "reference sub-blocks") in reference model R1. Similarly, T0 and reference model R0 have the same scale, so the training sub-blocks in image T0 will find corresponding reference sub-blocks in reference model R0.
[0038] Next, the training system will calculate the texture correlation between each sub-block at each scale and the reference sub-blocks in its corresponding reference model. Specifically, the training system will calculate the texture correlation between training sub-block T11 and reference sub-block R11 and calculate the texture correlation between sub-block T01 and reference sub-block R01. Here, the texture correlation can be the correlation coefficient coeff (coefficient) of the local binary pattern (LBP) between the sub-block and the reference sub-block.
[0039] Here, each sub-block will have its own feature vector f = {σ, k, coeff}, and each image will have its own image feature vector F = {f1, f2,..., f n}, where n is the number of sub-blocks. Taking Figure 5 as an example, the defect detection area image T1 will have an image feature vector where n1 is the number of training sub-blocks in the defect detection area image T1. Similarly, image T0 will have an image feature vector where n0 is the number of training sub-blocks in image T0. After that, the training system can concatenate the image feature vectors of the two scales into a feature vector to input to the classifier M.
[0040] After the training system inputs the feature vectors corresponding to all N1 + N2 pieces of training data into the classifier, the training of classifier M will start. Here, the classifier can be a support vector machines (SVM) classifier, and the training system will calculate the optimal separating hyperplane of the SVM classifier as the basis for distinguishing whether the device under test T has defects.
[0041] Figure 6 It is a functional block flowchart of the defect detection method shown in the embodiments of the present invention, and Figure 6 the process is applicable to the defect detection system 100. Before performing Figure 6 the process, the processing device 140 will pre-store Figure 5 the reference model and classifier mentioned above.
[0042] Please also refer to Figure 1 and Figure 6 . First, similar to steps S206 and S208, the processing device 140 will obtain the test data 602 (i.e., the received audio signal corresponding to the device under test T) and convert the test data into a time-frequency spectrogram 604. Here, the test data is the received audio signal in step S206.
[0043] Next, the processing device 140 will obtain multiple sub-blocks associated with the time-frequency spectrogram to obtain the spatial feature 610 therefrom and input it into the classifier 612. In this embodiment, the processing device 140 will also select a preset frequency range of, for example, 3K to 15K Hz as the detection area to generate a detection area image. In the embodiment, the processing device 140 can directly segment the detection area image to directly generate multiple sub-blocks of the same size. In another embodiment, the processing device 140 can perform image pyramid processing on the detection area image to generate multiple detection area images of different scales. Then, the processing device 140 segments the detection area images of different scales to generate multiple sub-blocks of the same size.
[0044] After that, the processing device 140 will obtain the eigenvalue of each sub-block and obtain the texture correlation 606 between each sub-block and the reference model 608 respectively. The eigenvalue here is, for example, at least one of the standard deviation of the pixel values of the sub-block and the histogram skewness, but it needs to meet the input requirements of the pre-stored classifier. The texture correlation here can be the correlation coefficient of the local binary pattern between the sub-block and the reference sub-block corresponding to the reference model. Then, the processing device 140 inputs the eigenvalue and texture correlation corresponding to each sub-block into the classifier 612 to generate an output result, and this output result will indicate whether the device under test T has a defect.
[0045] In this embodiment, in order to achieve more rigorous detection to avoid misjudging a device under test T with an actual defect as non-defective, the processing device 140 can further confirm based on the confidence level of the output result when the output result indicates that the device under test T does not have a defect. Specifically, taking the SVM classifier as an example, the processing device 140 can obtain the confidence value of the output result and determine whether the confidence value is greater than the preset confidence threshold 614, where the preset confidence threshold can be 0.75. If so, the processing device 140 will determine that the device under test T does not have a defect. Otherwise, the processing device 140 will determine that the device under test T has a defect.
[0046] In this embodiment, the defect detected by the defect detection system 100 is the assembly abnormal sound of the device under test T. Since different types of assembly abnormal sounds will generate resonance harmonics when playing a specific audio signal, the processing device 140 can further use the frequencies and harmonic frequency ranges of the abnormal sound in the time-frequency spectrum diagram to determine the components in the device under test T that cause the assembly abnormal sound. In another view, the processing device 140 will discriminate the components that cause the assembly abnormal sound according to a specific area of the time-frequency spectrum diagram.
[0047] For example, Figure 7 According to the schematic diagram of the time-frequency spectrum diagram shown in the embodiment of the present invention, only a partial area in the time-frequency spectrum diagram is schematically shown below. Both the time-frequency spectrum diagram 710 and the time-frequency spectrum diagram 720 have assembly abnormal sounds. Since the resonance frequency point when the screw is not tightened is a single-point resonance at 460 Hz, the processing device 140 can obtain from the time-frequency spectrum diagram 710 that the screw of the device under test T is not tightened. Since the resonance sound caused by iron filings in the speaker unit has resonance at 460 - 350 Hz, the processing device 140 can obtain from the time-frequency spectrum diagram 720 that there are iron filings in the device under test T.
[0048] In practice, to avoid overkill, when the device under test is determined to be defective due to assembly noise, the tester can judge whether the noise is acceptable or unacceptable based on the volume of the noise. When the noise is acceptable (difficult or impossible to be detected by the human ear), the device under test will be regarded as an "OK" device under test. When the noise is unacceptable, the device under test will be regarded as an "NG" device under test. Visually, Figure 8 FIG. 810 is a schematic diagram of a time-frequency spectrum of acceptable noise and FIG. 820 is a schematic diagram of a time-frequency spectrum of unacceptable noise shown according to an embodiment of the present invention, wherein the time-frequency spectrum 820 includes a distinct cluster with high brightness. In subsequent embodiments, an observation-based machine learning-based quantization mechanism will be proposed, which can classify the noise into acceptable and unacceptable to reduce the overkill rate.
[0049] Figure 9 FIG. 9 is a block diagram of a defect detection system shown according to an embodiment of the present invention, but this is only for convenience of description and is not intended to limit the present invention. First Figure 9 introduce all components and configuration relationships in the defect detection system, and the detailed functions will be disclosed in conjunction with Figure 10 together.
[0050] Please refer to Figure 9 , the defect detection system 900 includes a signal output device 910, a microphone 920, an analog-to-digital converter 930, and a processing device 940, where similar numbers prefixed with "9" are used to represent components similar to Figure 1 similar components. The defect detection system 900 is used to determine whether the device under test T has unacceptable defects in a pre-established auditory standard. Here, the pre-established auditory standard can be a range set according to human ear perception, a customized range according to customer requirements, a range defined by a third party, and so on. In the following embodiments, an electronic device with a speaker will also be used as the device under test T for illustration, and the defect detected by the defect detection system 900 is the assembly noise of the device under test T.
[0051] Figure 10 FIG. 10 is a flowchart of a defect detection visualization method shown according to an embodiment of the present invention, Figure 10 The process will be executed by the Figure 9 defect detection system 900.
[0052] Please also refer to Figure 9 and Figure 10, the signal output device 910 outputs a test audio signal to the device under test T (step S1002), and the microphone 920 receives the response signal of the device under test T to the test audio signal (step S1004). Then, the analog-to-digital converter 930 converts the response signal into a received audio signal (step S1006), and the processing device 940 performs signal processing on the received audio signal to generate a time-frequency spectrum diagram (step S1008). Here, for the details of steps S1002 to S1008, please refer to the relevant descriptions of steps S202 to S208 and will not be elaborated here. Then. The processing device 940 visually determines whether the device under test T has unacceptable defects in the pre-established auditory standard based on the time-frequency spectrum diagram (step S1010). In this embodiment, the processing device 940 can first determine whether the device under test T has defects in a manner similar to step S208 based on the time-frequency spectrum diagram. If so, the processing device 940 can further determine whether this defect is unacceptable in the pre-established auditory standard to avoid over-elimination.
[0053] Based on this, before the processing device 940 determines whether the device under test T has unacceptable defects, another classifier will be trained and constructed. The classifier here can be trained by the processing device 140 itself, or a pre-trained classifier can be obtained from other processing devices, and the present invention is not limited thereto. In the following embodiments, the classifier will be constructed in a manner similar to the processing device 940 (hereinafter referred to as the "training system"). First, the training system collects multiple pieces of training data, and these training data are training objects labeled as "acceptable defects" in the pre-established auditory standard. The training system will perform projection conversion and feature quantization processing on the time-frequency spectrum diagram corresponding to each training sound sample according to the time and spatial characteristics presented in the time-frequency spectrum diagram of the device under test with abnormal sounds.
[0054] Figure 11 It is a functional diagram for converting the time-frequency spectrum diagram shown in the embodiment of the present invention into a projection curve.
[0055] Please refer to Figure 11 , the training system extracts the region of interest 1115 from the time-frequency spectrum diagram 1110, where the region of interest 1115 can be, for example, Figure 7The regions that are discriminated as potentially representing abnormal sounds or the preset detection regions that usually have abnormal sounds. Then, the training system will divide the region of interest 1115 of the time-frequency spectrogram 1110 into multiple sub-time-frequency spectrograms (such as three regions R1 to R3 in this embodiment) relative to different frequency levels (i.e., horizontal division). The training system will convert the two-dimensional sub-time-frequency spectrograms R1 to R3 into one-dimensional projection curves R1 to R3 respectively. For example, such conversion can be to average the energy values (i.e., in the vertical direction) at each time in each sub-time-frequency spectrogram R1 to R3.
[0056] For the projection curves CR1 to CR3 where the horizontal axis represents time and the vertical axis represents energy, the projection values of the sub-time-frequency spectrograms with abnormal sound characteristics will be relatively high. When the projection values are continuously high over time, it is very likely to have serious abnormal sounds. In addition, the abnormal sound characteristics will be further classified into unacceptable (severe) and acceptable abnormal sound characteristics in a pre-established auditory standard. Assume that the pre-established auditory standard is set according to the range of human ear auditory perception. The human ear is more sensitive to specific frequencies. For example, when the abnormal sound characteristics only appear in the sub-time-frequency spectrogram R1 (frequencies are all approximately greater than 10K), then this abnormal sound may be acceptable. However, when the abnormal sound characteristics appear in all sub-time-frequency spectrograms R1 to R3, then this abnormal sound may be unacceptable. In other words, unacceptable (severe) abnormal sounds have the following characteristics: (1) higher projection energy, (2) longer duration, and (3) wider frequency coverage range. Then, the training system will perform feature quantization.
[0057] Specifically, in order to highlight local features, each projection curve will be further divided into multiple segments (i.e., vertical division) according to different time intervals. For example, Figure 11 the projection curve CR1 in Figure 12 can be further divided into five segments as and the sub-time-frequency spectrogram x i The eigenvalue of each segment corresponding to the j-th segment in the i-th region of the sub-time-frequency spectrogram will involve the statistical parameters of the data points in the corresponding segment and the weight assigned to the corresponding sub-time-frequency spectrogram, for example, according to Equation (1):
[0058]
[0059]
[0060] Here, H μ is the segment in which the value is greater than the segment average value The average value, and kσ is k times the standard deviation greater than the average value of this section. HH mean and HH mean are the average values of HH and HL respectively, HH size and HL size are the quantities of HH and HL respectively. Here, the H set composed of HH and HL represents the section in which the values are greater than the average value of the section and is the section the number of data points in is the weight of different sub-time spectrograms. L is the coefficient of each sub-time spectrogram, where the coefficient of the lower-frequency sub-time spectrogram is lower, and the abnormal sound in this interval is more important.
[0061] For the sake of clarity, assume then μ = 0.5 and H = {0.5, 0.9, 0.6, 0.7}. Assume kσ = 0, then H μ = 0.675. Assume HH = {0.9, 0.7} and HL = {0.5, 0.6}, then HH mean = 0.8, HL mean = 0.55, and Assume the weight of the low-frequency sub-time spectrogram is L = -1, then the section The quantization result can be expressed as
[0062] When the training system calculates the feature quantization results of each sub-time spectrogram of the training object with acceptable defects will be established and trained according to the known machine learning or deep learning model to judge the one-class support vector machine (OCSVM) classifier of acceptable abnormal sounds. After that, this classifier will be able to distinguish acceptable and unacceptable abnormal sounds.
[0063] Please go back to Figure 10, it should be understood that the details in S1010 regarding visually judging whether the device under test T has unacceptable defects according to the time spectrogram will correspond to the steps of training the OCSVM classifier. Specifically, when the processing device 940 receives the time spectrogram, it will extract the region of interest from the time spectrogram and divide the region of interest according to different frequency levels to generate multiple sub-time spectrograms. Then, the processing device 940 will convert all the sub-time spectrograms into one-dimensional projection curves respectively, and further divide all the projection curves into multiple sections according to different time intervals. The processing device 940 will calculate and input the feature quantization results of each sub-time spectrogram into the OCSVM classifier. The processing device 940 will obtain the confidence level of the output result (hereinafter referred to as "defect confidence level") and judge whether the defect confidence level is greater than the preset defect confidence level, where the preset defect confidence level can be 0. The preset defect confidence level here can be adjusted according to the actual application. When the judgment result is affirmative, the processing device 940 will determine that the device under test T has acceptable abnormal sounds. When the judgment result is negative, the processing device 940 will determine that the device under test T has unacceptable abnormal sounds.
[0064] For example, Figure 13 is a data graph of abnormal sound grading detection shown according to the embodiments of the present invention, where each data point represents a device under test. The defect confidence level of the device under test corresponding to point 1301 is 0.1, so it has acceptable abnormal sounds. In fact, all the devices under test corresponding to the points in cluster 1300 have acceptable abnormal sounds. The defect confidence level of the device under test corresponding to point 1303 is -0.38, and the defect confidence level of the device under test corresponding to point 1304 is -0.04, so they both have unacceptable abnormal sounds. The device under test corresponding to point 1302 is an outlier, so it will be retested.
[0065] Table 1 summarizes the use of no introduction (such as Figure 2 ) and introduction (such as Figure 10) Experimental results of the defect detection visualization method based on the machine learning quantization mechanism to distinguish acceptable and unacceptable abnormal sounds. Without introducing the quantization mechanism, among 1361 devices under test, 941 devices under test will be classified as "OK" devices under test (OK rate is 0.691), and 421 devices under test will be classified as "NG" devices under test (NG rate is 0.309). When the quantization mechanism is introduced, additional abnormal sound classification will be carried out according to the severity under the pre-established auditory standard. Among the 421 "NG" devices under test, 177 devices under test have acceptable abnormal sounds (overall OK rate is 0.821), and 244 devices under test have unacceptable abnormal sounds (overall NG rate is 0.179). Obviously, Table 1 shows that introducing the quantization mechanism can reduce the overall NG rate by 13%. In terms of product manufacturing and management, the costs of detection and retesting will be significantly reduced, and the lower over-rejection will improve the product yield. In specific applications, the NG devices under test can be graded according to the severity of the abnormal sounds for future product market planning.
[0066] OK Rate NG Rate Defect Detection with Imported Quantization Mechanism 0.691 0.309 Defect Detection without Imported Quantization Mechanism 0.821 0.179
[0067] Table 1
[0068] In summary, the defect detection method and system proposed by the present invention can use computer vision technology from the time-frequency spectrum diagram to detect whether the device under test has unacceptable defects in the pre-established auditory standard. In this way, the present invention can not only provide more accurate defect detection than the subjective judgment of the human ear, but also reduce related occupational injuries.
[0069] [Symbol Explanation]
[0070] 100, 900: Defect detection system
[0071] 110, 910: Signal output device
[0072] 120, 920: Microphone
[0073] 130, 930: Analog-to-digital converter
[0074] 140, 940: Processing device
[0075] T: Device under test
[0076] S202~S210, S1002~S1010: Steps
[0077] 310, 320, 710, 720, 810, 820, 1100: Time-frequency spectrum diagram
[0078] 315, 325: Detection area
[0079] RB: Abnormal sound feature
[0080] 402 - 414, 602 - 614: Process
[0081] H: Image pyramid
[0082] T1, T0: Image
[0083] T11, T01: Sub - block
[0084] R1, R0: Reference model
[0085] R11, R01: Reference sub - block
[0086] FE: Feature extraction
[0087] M: Classifier
[0088] 1115: Region of interest
[0089] R1 - R3: Sub - time frequency spectrum diagram
[0090] CR1 - CR3: Projection curve
[0091] 1300: Cluster
[0092] 1301 - 1304: Point
Claims
1. A method for visualizing defect detection, comprising: Output a test audio signal to the device under test; Receive a response signal of the device under test to the test audio signal to generate a received audio signal; Perform signal processing on the received audio signal to generate a time-frequency spectrum diagram; And According to the time-frequency spectrum diagram, visually determine whether the device under test has unacceptable defects in a pre-established auditory standard, wherein the step of visually determining whether the device under test has the unacceptable defects in the pre-established auditory standard according to the time-frequency spectrum diagram includes: Obtain a plurality of sub-time-frequency spectrum diagrams associated with the time-frequency spectrum diagram; Convert each of the sub-time-frequency spectrum diagrams into a projection curve respectively; Obtain a plurality of sections associated with each of the projection curves; Generate a feature quantization result corresponding to each of the sections of each of the projection curves; and According to the feature quantization result and a classifier, determine whether the device under test has the unacceptable defects.
2. The method according to claim 1, wherein the step of receiving the response signal of the device under test to the test audio signal to generate the received audio signal comprises: Use a microphone to receive the response signal; And Perform analog-to-digital conversion on the response signal to generate the received audio signal.
3. The method according to claim 2, wherein the step of performing signal processing on the received audio signal to generate the time spectrogram comprises: Perform a fast Fourier transform on the received audio signal to generate the time-frequency spectrum diagram.
4. The method according to claim 1, wherein the step of visually determining whether the device under test has the unacceptable defect in the pre-established auditory standard according to the time spectrogram comprises: According to the time-frequency spectrum diagram, visually determine whether the device under test has defects; And When the device under test has the defects, according to the time-frequency spectrum diagram, determine whether the defects are unacceptable in the pre-established auditory standard.
5. The method according to claim 1, wherein the step of obtaining the sub-time spectrograms associated with the time spectrogram comprises: Extract a region of interest from the time-frequency spectrum diagram, where the region of interest corresponds to a preset frequency range; And Divide the region of interest according to different frequency levels to generate the sub-time-frequency spectrum diagrams.
6. The method according to claim 1, wherein the step of respectively converting each of the sub-time spectrograms into the projection curve comprises: Calculate the average energy value at each time in each of the sub-time-frequency spectrum diagrams to generate the projection curve.
7. The method according to claim 1, wherein the step of obtaining the sections associated with each of the projection curves comprises: Divide each of the projection curves into the sections according to different time intervals.
8. The method according to claim 1, wherein the feature quantization result for each of the respective segments corresponding to the respective projection curves is a statistical parameter associated with a plurality of data points corresponding to the respective segment and a weight assigned to the corresponding sub-time spectrogram.
9. The method according to claim 1, wherein the step of determining whether the device under test has the unacceptable defect based on the feature quantization result and the classifier includes: Input the feature quantization results corresponding to all the sections of all the projection curves into the classifier; Receive the output result of the classifier; And According to the output result of the classifier, determine whether the device under test has the unacceptable defects.
10. The method according to claim 9, wherein the classifier is a support vector machine classifier and is constructed based on a plurality of defective training objects having acceptable defects in the pre-established auditory criteria.
11. The method according to claim 9, wherein the step of determining whether the device under test has the unacceptable defect based on the output result of the classifier includes: Obtain the defect confidence level of the output result; Judge whether the defect confidence level is greater than a preset defect confidence level; When the defect confidence level is greater than the preset defect confidence level, determine that the device under test has acceptable defects; And When the defect confidence level is not greater than the preset defect confidence level, determine that the device under test has the unacceptable defects.
12. The method according to claim 1, wherein the device under test is an electronic device having a speaker.
13. The method according to claim 1, wherein the defect is an assembly abnormal sound of the device under test.
14. A defect detection system, comprising: A signal output device for outputting a test audio signal to the device under test; A microphone for receiving a response signal of the device under test to the test audio signal; An analog-to-digital converter for converting the response signal into a received audio signal; And A processing device for performing signal processing on the received audio signal to generate a time-frequency spectrum diagram, and visually determining whether the device under test has unacceptable defects in a pre-established auditory standard according to the time-frequency spectrum diagram, The processing device further stores a classifier in advance, and the processing device obtains a plurality of sub-time spectrograms associated with the time spectrogram, converts each of the sub-time spectrograms into a projection curve respectively, obtains a plurality of segments associated with each of the projection curves, generates a feature quantization result corresponding to each of the segments of each of the projection curves, and determines whether the device under test has the unacceptable defect according to the feature quantization result and the classifier.
15. The system according to claim 14, wherein the processing device visually determines whether the device under test has a defect based on the time spectrogram, and when the device under test has the defect, determines whether the defect is unacceptable in the pre-established auditory criteria based on the time spectrogram.
16. The system according to claim 15, wherein the device under test is an electronic device having a speaker.
17. The system according to claim 15, wherein the defect is an assembly abnormal sound of the device under test.
18. A method for visualizing defect detection of a device under test, comprising: Output a test audio signal; Receive a response signal, which is generated with respect to the test audio signal; Convert the response signal into a received audio signal; Perform signal processing on the received audio signal to generate a time spectrogram; And Visually determine whether there is an unacceptable defect in a pre-established auditory standard according to the time spectrogram, wherein the step of visually determining whether there is an unacceptable defect in a pre-established auditory standard according to the time spectrogram includes: Obtain a plurality of sub-time spectrograms associated with the time spectrogram; Convert each of the sub-time spectrograms into a projection curve respectively; Obtain a plurality of segments associated with each of the projection curves; Generate a feature quantization result corresponding to each of the segments of each of the projection curves; and Determine whether the device under test has the unacceptable defect according to the feature quantization result and the classifier.
Citation Information
Patent Citations
Speaker abnormal sound detecting method based on short-time Fourier transformation
CN103546853A