Sound detection system using convolutional neural network technology

By designing a sound detection system using convolutional neural network technology, the problem of inaccurate identification in complex environments is solved, and higher recognition accuracy and reliability are achieved, which is suitable for real-time detection of industrial production lines.

CN120032631APending Publication Date: 2025-05-23BEIJING PHANTOM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411730979.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing sound recognition technology has limitations in practical applications. The traditional method is not robust, poorly adaptable, and requires professional debugging, making it difficult to accurately identify target sounds in complex environments.

Method used

A sound detection system using convolutional neural network (CNN) technology is designed, which includes an edge computing control module and sound acquisition sensor, using PaddlePaddle Image Recognition Kit or TensorFlow or Pytorch Development Kit, sound recognition is performed through spectral transformation and recognition models, and feedback is performed through buzzers and tri-color lights.

Benefits of technology

It improves the accuracy and reliability of sound recognition, reduces the requirements for sound acquisition conditions, simplifies the operation process, reduces dependence on professionals, has a wider applicability, and is small in size, which is suitable for real-time inspection on the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032631A_ABST
    Figure CN120032631A_ABST
Patent Text Reader

Abstract

The invention discloses a sound detection system and method using a convolutional neural network technology, and the system comprises an edge calculation control module and a sound collection sensor, and the sound collection sensor is disposed on the edge calculation control module. The beneficial effects of the invention are that the sound acquisition sensor is directly connected with the edge calculation control module, when other auxiliary devices are not accessed and sound analog signals are transmitted into the edge calculation module to produce a spectrogram, the built-in identification model can identify and judge the spectrogram, and meanwhile, the edge calculation control module is internally provided with the buzzer, so that the sound acquisition sensor can be directly connected with the edge calculation control module. Different buzzing frequencies can be given according to different sound recognition results, the system can directly communicate with a three-color lamp, direct acousto-optic feedback is achieved, collected sound information is judged from multiple aspects, recognition of multiple judgment factors is achieved through the use of a recognition model, and therefore the applicability is wider.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound detection, and in particular to a sound detection system using convolutional neural network technology. Background Art

[0002] In current industrial production, many inspection contents rely on sound. For example, the sound emitted during the engine test process can provide feedback on the engine's operating conditions. During steel production, the sound changes in the furnace can be used to determine whether the furnace is operating normally. During wine bottle production, the sound made by knocking on the bottle can be used to determine whether the bottle meets production standards.

[0003] In such scenarios, workers will be affected by long-term exposure to noise above 85dB. Therefore, based on the specific industrial needs of our company, we conducted a survey of the existing market and found that there is currently no general customizable sound detection equipment. Most of the existing sound recognition related equipment on the market use traditional sound recognition methods. However, these methods have obvious limitations in practical applications. Traditional methods mainly analyze based on some specific time domain and frequency domain characteristics of sound signals. It is necessary to adjust the thresholds in each step of the analysis. The adaptability of the scene is poor, and the on-site operation is complicated and affected by the surrounding environmental noise, making it difficult to accurately identify the target sound. Abroad, the research focus of sound recognition technology is mainly on extracting the characteristics of sound using deep learning technologies such as convolutional neural networks and recurrent neural networks. Through these advanced technical means, the potential information in the sound can be more deeply explored, thereby improving the accuracy and reliability of recognition. In contrast, domestic research in the field of sound recognition mainly focuses on the application field of technology. There is relatively little research on deep learning methods for sound recognition, and traditional machine learning technology still accounts for a large proportion. Although traditional machine learning technology can capture the differences in sound features and achieve classification tasks, each input is considered equally important, and the model cannot distinguish which inputs are most critical to the task. Therefore, it has certain deficiencies in processing complex sounds and extracting deep features. It also has high requirements for the collection of detection sounds and high investment costs.

[0004] As a supervised deep learning algorithm, Convolution Neural Network (CNN) has excellent scalability, generalization ability and diversity of input data, so it performs well in sound detection and has achieved many results in theory. Compared with unsupervised deep learning, this method can be quickly put into production line through a small number of pictures, and the cost is low. It can solve other interference factors generated by various production processes. We can distinguish different types of sound faults by expanding it by increasing the depth or width.

[0005] The time domain or frequency domain recognition used by traditional sound can achieve the recognition purpose, but the effect is not good, the robustness is not strong, and the requirements for sound collection conditions are high; the support vector machine method is currently the most commonly used method in theory, but the detection effect for the actual production environment is similar to that of traditional detection, and it requires professional personnel to debug. In view of this, in-depth research was conducted on the above-mentioned issues, which led to the creation of this case. Summary of the invention

[0006] The purpose of the present invention is to solve the above problems. A sound detection system using convolutional neural network technology is designed to solve the problem that the time domain or frequency domain recognition used in the existing sound can achieve the recognition purpose, but the effect is not good, the robustness is not strong, and the requirements for sound collection conditions are high; the support vector machine method is currently the most commonly used method in theory, but the detection effect for the actual production environment is similar to that of the traditional detection, and it requires professional personnel to debug.

[0007] The technical solution of the present invention to achieve the above-mentioned purpose is: a sound detection system using convolutional neural network technology, comprising: An edge computing control module and a sound collection sensor, wherein the sound collection sensor is installed on the edge computing control module; The edge computing control module includes a sound card module, a spectrogram conversion module, a recognition model and a prompt module; The recognition model uses the PaddleClas and PaddleDetection algorithms of the PaddlePaddle image recognition kit or the development kit of TensorFlow or Pytorch.

[0008] Preferably, the edge computing control module further includes: The body, power button, sensor interface, open interface and touch display; The power button and the sensor interface are respectively arranged on the two side walls of the body, the open interface is arranged on the side wall of the body and is located below the sensor interface, the touch display screen is arranged on the front end wall of the body, and the prompt module is arranged on the body.

[0009] Preferably, the prompt module includes: Buzzer and three-color light.

[0010] Preferably, an auxiliary device is further provided between the edge computing control module and the sound collection sensor, and the auxiliary device includes: Signal amplifiers, audio cards, analog-to-digital conversion equipment, filters, and pulse code modulation equipment.

[0011] Preferably, the sound collection sensor is used to collect a sound source, and the sound collection sensor converts the sound source into an analog electrical signal.

[0012] Preferably, the sound card module is used to convert the analog electrical signal into a digital signal, the spectrogram conversion module is used to convert the digital signal into a spectrogram, the recognition model is used to judge the spectrogram to obtain judgment information, and the prompt module is used to provide feedback on the judgment information.

[0013] Preferably, the number of the open interfaces is no less than 1, and the sensor interface and the open interface match USB, coaxial audio cable, optical fiber audio cable, TRS and RCA.

[0014] Preferably, the buzzer includes at least one buzzing frequency.

[0015] A sound detection method using convolutional neural network technology, comprising the following steps; Step 1: The acquisition sensor collects the sound to form an analog electrical signal; Step 2: The sound card module converts the analog electrical signal into a digital signal, and the spectrogram conversion module converts the digital signal into a spectrogram; Step 3: The spectrogram image is displayed on the touch screen, and the operator draws the BoundingBox on the touch screen; Step 4, training the spectrogram image drawn by BoundingBox to obtain a customized recognition algorithm; Step 5: When the detected object makes a sound during the production process, the edge computing control module determines the content in the spectrogram image.

[0016] The method for determining the spectrogram image by the edge computing control module described in step 5 includes the following steps: Step 5.1, after the spectrogram conversion module converts the digital signal into a spectrogram image, the recognition model randomly divides the spectrogram image into a number of grids; Step 5.2, each grid will predict the information and confidence of BoundingBox; Step 5.3: judging whether the detected target exists in the several grids randomly split in step 5.1 according to the confidence level; Step 5.4, remove target information with low probability according to the threshold; Step 5.5: Split and predict again to further predict the target to be detected; Step 5.6: Perform non-maximum suppression processing to obtain the target to be identified; Step 5.7, result processing, i.e., determining whether the recognition model is qualified or unqualified; Step 5.8: Output signal, that is, the three-color light or buzzer works.

[0017] In the sound detection system using the convolutional neural network technology produced by the technical solution of the present invention, the sound collection sensor is directly connected to the edge computing control module. When the other auxiliary devices are not connected, when the analog signal of the sound is transmitted to the edge computing module to produce the spectrogram, the built-in recognition model can recognize and judge the spectrogram. At the same time, the edge computing control module has a built-in buzzer, which can give different buzzing frequencies according to different sound recognition results. It can directly communicate with the three-color light to realize direct feedback of sound and light, and realize the judgment of the collected sound information from multiple aspects. The use of the recognition model also solves the recognition of multiple judgment factors, so it has a wider applicability and is more suitable for edge computing. The control module can directly change the judgment conditions of the sound without the intervention of computers or other equipment, and can be used completely offline; the judgment program can be changed without the operation of professionals, which is simple and fast, the equipment is small in size, and it is convenient and economical to install for production line upgrades. In the use of the sound detection system, the sound detection system can customize the collected sounds and can quickly adapt to different production product detection needs. The touch screen in the edge computing control module can display different spectrograms, which can intuitively reflect the detection content. There is no need for operators to adjust the code or parameters, only to make corresponding judgments on different pictures, which lowers the usage threshold and improves production efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a workflow diagram of a sound detection system using convolutional neural network technology described in the present invention.

[0019] Figure 2 This is an example of spectrogram code conversion of a sound detection system using convolutional neural network technology described in the present invention.

[0020] Figure 3 This is a connection diagram of a sound detection system using convolutional neural network technology described in the present invention when a windproof sound collecting cover is not used.

[0021] Figure 4 This is a schematic diagram of the operation of a sound detection system using convolutional neural network technology when detecting bottles according to the present invention.

[0022] Figure 5 This is a three-dimensional diagram from a left angle of view of an edge computing control module of a sound detection system using convolutional neural network technology as described in the present invention.

[0023] Figure 6 This is a three-dimensional diagram from the right angle of the edge computing control module of the sound detection system using convolutional neural network technology described in the present invention.

[0024] Figure 7 This is a workflow diagram of a sound detection method using convolutional neural network technology described in the present invention.

[0025] Figure 8 This is a workflow diagram of a method for determining a spectrogram image in a sound detection method using convolutional neural network technology described in the present invention.

[0026] Fig. 9 This is a spectrogram image in Example 4 of the present invention.

[0027] Fig.10 This is another spectrogram image in Example 4 of the present invention.

[0028] Fig.11 This is a spectrogram image in Example 5 of the present invention.

[0029] In the figure: 1. Edge computing control module, 101. Machine body, 102. Power button, 103. Sensor interface, 104. Open interface, 105. Touch screen, 106. Buzzer, 107. Three-color light, 2. Sound collection sensor, 3. Auxiliary equipment, 4. Bottle. DETAILED DESCRIPTION

[0030] The present invention will be described in detail below in conjunction with the accompanying drawings. Figure 1-11 As shown, a sound detection system and method using convolutional neural network technology.

[0031] Through the personnel in this field, all the electrical components in this case are connected to their corresponding power supplies through wires, and a suitable controller should be selected according to the actual situation to meet the control requirements. The specific connection and control sequence should refer to the following working principle, and the electrical connection between the electrical components is completed in the order of working in sequence. The detailed connection means are well-known technologies in this field. The following mainly introduces the working principle and process, and does not explain the electrical control.

[0032] Embodiment 1: A sound detection system using convolutional neural network technology, comprising: An edge computing control module 1 and a sound collection sensor 2, wherein the sound collection sensor 2 is installed on the edge computing control module 1; The edge computing control module 1 includes a sound card module, a spectrogram conversion module, a recognition model and a prompt module; The recognition model uses the PaddleClas and PaddleDetection algorithms of the PaddlePaddle image recognition suite or the TensorFlow or Pytorch development suite; In the specific implementation process, the sound collection sensor 2 is used to collect the sound source, and the sound collection sensor 2 converts the sound source into an analog electrical signal; In the specific implementation process, the sound card module is used to convert the analog electrical signal into a digital signal, the spectrogram conversion module is used to convert the digital signal into a spectrogram, the recognition model is used to judge the spectrogram to obtain judgment information, and the prompt module is used to provide feedback on the judgment information; It should be noted that the sound collection sensor 2 includes but is not limited to array microphones, microphones based on micro-electromechanical system technology, voice recorders and other devices. The edge computing control module 1 contains a set of software programs for converting analog signals into spectrograms. The software program covers commonly used methods such as Mel-frequency cepstrum coefficients, Fourier transform, and linear prediction. The edge computing control module 1 is also provided with a sound card module. The sound collection sensor 2 is directly connected to the edge computing control module 1. When the auxiliary device 3 is not connected, the sound source collected by the sound collection sensor 2 is converted into an analog electrical signal, and the analog electrical signal is converted into a digital signal through the sound card module. The digital signal is then converted into a spectrogram through the spectrogram conversion module, and the recognition model recognizes and judges the spectrogram. In the specific implementation process, the edge computing control module 1 also includes: a body 101, a power button 102, a sensor interface 103, an open interface 104 and a touch display screen 105; the power button 102 and the sensor interface 103 are respectively arranged on the two side walls of the body 101, the open interface 104 is arranged on the side wall of the body 101 and is located below the sensor interface 103, the touch display screen 105 is arranged on the front end wall of the body 101, and the prompt module is arranged on the body 101; It should be noted that the power button 102 is used to start or shut down the edge computing control module 1, the sensor interface 103 and the open interface 104 are used to plug in the data cable, and the touch screen 105 provides an intuitive operation interface for the user, through which the user can view the operating status, configuration parameters, and receive prompt information of the edge computing control module 1; In the specific implementation process, the prompt module includes: a buzzer 106 and a three-color light 107; In a specific implementation, the buzzer 106 includes at least one buzzing frequency It should be noted that after the recognition model recognizes and judges the spectrogram, the buzzer 106 or the three-color light 107 is used to feedback the judgment result. If it is unqualified, an alarm or a red light will be turned on, and if it is qualified, a green light will be turned on; In the specific implementation process, the sound collection sensor 2 can also be equipped with a windproof sound collecting cover; It should be noted that when there is wind, the windproof sound-gathering cover is used to improve the initial sound quality of the collection end, and the windproof sound-gathering cover can collect the sound more concentratedly or emit most of the sound in a specific direction; It should be noted that when the windproof sound-gathering cover is needed, the sound collection sensor 2 can be inserted into the windproof sound-gathering cover; In the specific implementation process, an auxiliary device 3 is also provided between the edge computing control module 1 and the sound collection sensor 2, and the auxiliary device 3 includes: signal amplifiers, audio cards, analog-to-digital conversion equipment, filters, and pulse code modulation equipment; It should be noted that the sound collection sensor 2 and the auxiliary device 3 are usually connected by a 3.5mm audio cable, an XLR cable and a USB cable. The signal amplifier amplifies the weak sound signal captured by the sound collection sensor 2 to a sufficient level for subsequent information processing. At the same time, it can also reduce noise and distortion in the signal, thereby improving the quality of the audio signal; the audio card is used for the input and output of some special sound collection sensors 2 to the edge computing control module 1, which plays a certain role in improving sound quality and signal conversion; the analog-to-digital conversion device is used as a professional analog signal to digital device for signal conversion; the filter can eliminate noise and enhance some feature points of the sound during the collection process, which is helpful for the final algorithm recognition; the pulse code modulation device is a professional signal conversion device, which is used to quantize and encode the analog sound signal into a high-quality digital signal; the above are all related equipment with mature market, and are used to assist in the collection and signal conversion of sound; In a specific implementation process, the number of the open interface 104 is no less than one, and the sensor interface 103 and the open interface 104 match USB, coaxial audio cable, optical fiber audio cable, TRS and RCA.

[0033] Embodiment 2: A sound detection method using convolutional neural network technology, comprising the following steps; Step 1: The acquisition sensor 2 collects the sound to form an analog electrical signal; Step 2: The sound card module converts the analog electrical signal into a digital signal, and the spectrogram conversion module converts the digital signal into a spectrogram; Step 3: The spectrogram image is displayed on the touch screen 105 , and the operator draws a BoundingBox on the touch screen 105 ; Step 4: After training the spectrogram image with the BoundingBox drawn, a customized recognition algorithm is obtained; Step 5: During the production process, after the object under detection makes a sound, the edge computing control module 1 determines the content in the spectrogram image.

[0034] It should be noted that when continuously detecting the sounds of products of the same category (such as glass bottles), through a custom recognition algorithm, it can be recognized that the sounds are generated by the same product, and the operator does not need to draw a BoundingBox (bounding box) on the touch display screen 105. When switching to continuously detecting the sounds of another product (such as switching from a glass bottle to a metal bottle), the operator draws a BoundingBox (bounding box) on the touch display screen 105, that is, selects the spectrogram image of this product (metal bottle) on the touch display screen 105 for subsequent continuous sound detection.

[0035] Embodiment 3: The method for the edge computing control module 1 to determine the spectrogram image described in Step 5 includes the following operating steps; Step 5.1: After the spectrogram conversion module converts the digital signal into a spectrogram image, the recognition model randomly divides the spectrogram image into several grids (Grid cell); Step 5.2: Each grid will predict the information of the BoundingBox and the confidence; Step 5.3: According to the confidence, it is judged whether the detected target exists in the several grids (Grid cell) randomly split in Step 5.1; Step 5.4: Remove the target information with lower possibility according to the threshold; Step 5.5: Further splitting and prediction can further predict the target to be detected; Step 5.6: Perform non-maximum suppression processing to obtain the target to be recognized; Step 5.7: Result processing, that is, the recognition model determines whether it is qualified or unqualified; Step 5.8: Output signal, that is, the three-color lamp or the buzzer works.

[0036] Embodiment 4: Quality inspection of glass wine bottles During the production process, the bottle 4 manufacturing enterprise needs to manually distinguish the sound made by knocking on the bottle 4 to judge the forming quality of the bottle 4. For example, if there are air bubbles, cracks and foreign objects in the bottle 4, the sound will be different.

[0037] The sound collection sensor 2 can use a bone conduction sensor. The sound collection sensor 2 is placed close to the bottle 4 and knocks the bottle 4. The recognition model in the edge computing control module 1 can make recognition judgments on the trend information of the spectrogram after collection and learning. In the field experiment, the signal of the sound collection sensor 2 was tested. It was found that at an appropriate installation distance, the signal collected by the sound collection sensor 2 was small, and the noise produced by the machine was too large. Therefore, an auxiliary device 3, such as a filter, was connected to improve the quality of sound collection. The whole device was installed on the assembly line, and an automatic knocking device was set to achieve real-time detection in the production process.

[0038] Example 5: Mixed sound module detection The implementation background is that during the production process of this product, there may be input errors or input failures when writing sound information to the sound unit, resulting in mixed loading or no sound of products of different specifications; since the product itself has a sound function and the sound is relatively loud, installing this product on the production line can realize automatic production and testing of the product without the need for auxiliary equipment 3.

[0039] Embodiment 6: Engine abnormal noise detection The user needs to detect the vibration of the engine cavity and determine its vibration amplitude. Since the engine itself is noisy and easily interferes with conventional sound collection equipment, a bone conduction microphone (MEMS) based on micro-electromechanical system technology is used, such as a bone conduction microphone, which is placed in contact with a vibrating object. The built-in vibration sensor can also convert the vibration frequency information into a digital signal and transmit it to the edge computing control module 1 for identification.

[0040] The sound collection sensor 2 is brought into contact with the engine casing. When the engine is started, the sound is directly in contact with the engine, so the source of the sound is more likely to be generated by the vibration of the engine cavity. Due to the high requirement for test accuracy, auxiliary equipment 3, such as an audio card and a filter, is provided.

[0041] The above technical solutions only reflect the preferred technical solutions of the technical solutions of the present invention. Some changes that may be made to certain parts thereof by technicians in this technical field all reflect the principles of the present invention and fall within the protection scope of the present invention.

Claims

1. A sound detection system using convolutional neural network technology, characterized in that: include: An edge computing control module (1) and a sound collection sensor (2), wherein the sound collection sensor (2) is installed on the edge computing control module (1); The edge computing control module (1) comprises a sound card module, a spectrogram conversion module, a recognition model and a prompt module; The recognition model uses the PaddleClas and PaddleDetection algorithms of the PaddlePaddle image recognition kit or the development kit of TensorFlow or Pytorch.

2. The sound detection system using convolutional neural network technology according to claim 1, characterized in that: The edge computing control module (1) further includes: A machine body (101), a power button (102), a sensor interface (103), an open interface (104), and a touch display screen (105); The power-on button (102) and the sensor interface (103) are respectively arranged on the side walls of the body (101); the open interface (104) is arranged on the side wall of the body (101) and is located below the sensor interface (103); the touch display screen (105) is arranged on the front end wall of the body (101); and the prompt module is arranged on the body (101).

3. The sound detection system using convolutional neural network technology according to claim 2, characterized in that: The prompt module comprises: A buzzer (106) and a three-color light (107).

4. The sound detection system using convolutional neural network technology according to claim 1, characterized in that: An auxiliary device (3) is also provided between the edge computing control module (1) and the sound collection sensor (2), and the auxiliary device (3) comprises: Signal amplifiers, audio cards, analog-to-digital conversion equipment, filters, and pulse code modulation equipment.

5. The sound detection system using convolutional neural network technology according to claim 1, characterized in that: The sound collection sensor (2) is used to collect a sound source, and the sound collection sensor (2) converts the sound source into an analog electrical signal.

6. The sound detection system using convolutional neural network technology according to claim 5, characterized in that: The sound card module is used to convert the analog electrical signal into a digital signal, the spectrogram conversion module is used to convert the digital signal into a spectrogram, the recognition model is used to judge the spectrogram to obtain judgment information, and the prompt module is used to provide feedback on the judgment information.

7. The sound detection system using convolutional neural network technology according to claim 2, characterized in that: The number of the open interfaces (104) is no less than one, and the sensor interface (103) and the open interface (104) are compatible with USB, coaxial audio cable, optical fiber audio cable, TRS and RCA.

8. The sound detection system using convolutional neural network technology according to claim 3, characterized in that: The buzzer (106) includes at least one buzzing frequency.

9. A sound detection method using convolutional neural network technology, applied to the system according to any one of claims 1 to 8, characterized in that: The following steps are included: Step 1: The acquisition sensor (2) acquires the sound and forms an analog electrical signal; Step 2: The sound card module converts the analog electrical signal into a digital signal, and the spectrogram conversion module converts the digital signal into a spectrogram; Step 3: The spectrogram image is displayed on the touch screen (105), and the operator draws a BoundingBox on the touch screen (105); Step 4, training the spectrogram image drawn by BoundingBox to obtain a customized recognition algorithm; Step 5: When the detected object makes a sound during the production process, the edge computing control module (1) determines the content in the spectrogram image.

10. The sound detection method using convolutional neural network technology according to claim 9, characterized in that: The method for determining the spectrogram image by the edge computing control module (1) in step 5 comprises the following steps: Step 5.1, after the spectrogram conversion module converts the digital signal into a spectrogram image, the recognition model randomly divides the spectrogram image into a number of grids; Step 5.2, each grid will predict the information and confidence of BoundingBox; Step 5.3: judging whether the detected target exists in the several grids randomly split in step 5.1 according to the confidence level; Step 5.4, remove target information with low probability according to the threshold; Step 5.5: Split and predict again to further predict the target to be detected; Step 5.6: Perform non-maximum suppression processing to obtain the target to be identified; Step 5.7, result processing, i.e., determining whether the recognition model is qualified or unqualified; Step 5.8: Output signal, that is, the three-color light or buzzer works.