Noise event identification method, system and device based on main sound source detection, and medium
Through continuous slice processing and main energy detection algorithm on the noisy audio, the type of noise source that accounts for the largest proportion in the noisy audio is identified, which solves the problem of low accuracy in the prior art and realizes accurate identification in the case of aliasing of multiple noise sources.
Patent Information
- Application Number
- CN202510287251.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing noise source recognition methods cannot accurately identify noise sources to specific features under the aliasing of multiple noise sources, resulting in low recognition accuracy.
The noise audio is obtained through the noise sensor for continuous slice processing, the time frequency domain soundprint features are extracted and the deep learning model is input to the noise source classification, the energy value and confidence of the audio slice are calculated, and the main energy detection algorithm is used to identify the noise source type that accounts for the largest proportion of the noise audio.
Accurately identifying major noise sources in complex multi-source noise environments improves the accuracy of noise source identification.
Smart Images

Figure CN120279941A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of noise recognition technology, and in particular, to a method, system, device, and medium for noise event recognition based on main sound source detection. Background Art
[0002] Noise source recognition technology is widely used in noise pollution prevention and control and environmental monitoring. It can analyze the time-frequency characteristics of noise signals, identify different types of noise events, and thus achieve intelligent monitoring, classification, and prediction of noise sources. Existing noise source recognition methods generally use simple time-frequency analysis or Fourier transform. However, these methods lack pertinence in the control of noise pollution sources. Due to the complex actual noise environment, especially in the case of overlapping multiple noise sources, it is impossible to accurately identify the specific characteristics of the noise source and match the sound source type, resulting in low accuracy of noise source recognition. Summary of the Invention
[0003] Embodiments of this application provide a method, system, device, and medium for noise event recognition based on main sound source detection, which can identify the main noise source and determine an accurate noise source recognition result based on the main noise source.
[0004] In a first aspect, embodiments of this application provide a method for noise event recognition based on main sound source detection, including:
[0005] Obtain noisy audio through a noise sensor, and perform continuous slicing processing on the noisy audio to obtain multiple audio slices with the same duration;
[0006] Extract the time-frequency domain voiceprint features of each audio slice, and input each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and recognition to obtain the recognition result of each slice. The slice recognition result includes the noise source classification result corresponding to the audio slice and the confidence of the noise source classification result;
[0007] Calculate the energy value of each audio slice, and associate each energy value with the noise source classification result and the confidence of the corresponding audio slice, where the energy value is the time-domain energy value or the frequency-domain energy value of the audio slice;
[0008] Based on each energy value, the corresponding noise source classification result and confidence, and a preset main energy detection algorithm, calculate the main energy type corresponding to the noisy audio, where the main energy type is used to indicate the type of the noise source with the largest proportion in the noisy audio;
[0009] Perform noise event recognition based on the noise source classification result corresponding to the main energy type.
[0010] In some embodiments, based on each of the energy values, the corresponding noise source classification result, the confidence level, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio includes:
[0011] Modifying the noise source type corresponding to the noise source classification result corresponding to the confidence level less than the confidence level threshold to a reference noise source type;
[0012] Adding up the energy values corresponding to the noise source classification results of the same type to obtain a plurality of first reference energy values;
[0013] Determining the noise source type corresponding to the noise source classification result corresponding to the first reference energy value with the largest value as the main energy type.
[0014] In some embodiments, based on each of the energy values, the corresponding noise source classification result, the confidence level, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio includes:
[0015] Performing multiplicative weighting processing on each of the energy values and the corresponding confidence level to obtain each second reference energy value;
[0016] Determining the noise source type corresponding to the noise source classification result corresponding to the second reference energy value with the largest value as the main energy type.
[0017] In some embodiments, based on each of the energy values, the corresponding noise source classification result, the confidence level, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio includes:
[0018] Calculating the leq value of the corresponding audio slice based on each of the energy values;
[0019] Modifying the noise source classification result corresponding to the confidence level less than the confidence level threshold to a reference noise source type, adding up the leq values corresponding to the noise source classification results of the same type to obtain a plurality of first reference leq values, and determining the noise source type corresponding to the noise source classification result corresponding to the first reference leq value with the largest value as the main energy type;
[0020] Or,
[0021] Performing multiplicative weighting processing on each of the leq values and the confidence level of the corresponding audio slice to obtain each second reference leq value, and determining the noise source type corresponding to the noise source classification result corresponding to the second reference leq value with the largest value as the main energy type.
[0022] In some embodiments, when applied to a noise recognition system, after calculating the main energy type corresponding to the noisy audio based on each of the energy values, the corresponding noise source classification results, the confidence level, and a preset main energy detection algorithm, the method further includes:
[0023] Determine the number of noise source types in the noise source classification results corresponding to all the audio slices corresponding to the noisy audio;
[0024] Determine the current available system resource amount of the noise recognition system;
[0025] When the number of noise source types is greater than a preset number threshold and the available system resource amount is greater than a preset resource amount threshold, calculate the energy proportion of the main energy sound source corresponding to the main energy type in each of the audio slices to obtain a plurality of first proportions, and sum all the first proportions to obtain a second proportion, where the second proportion is used to indicate the energy proportion of the main energy sound source corresponding to the main energy type in the noisy audio.
[0026] In some embodiments, when applied to a noise recognition system, the noise recognition system includes the noise sensor, the noise recognition system is associated with a target account, and performing noise event recognition based on the noise source classification result corresponding to the main energy type includes:
[0027] Obtain meteorological data, sound source localization data, the physical location of the noise sensor, an over-standard classification threshold, and sound source investigation situation information during the time period when the noisy audio is located;
[0028] Obtain the minute-level leq monitoring data corresponding to the noisy audio;
[0029] Based on a preset AI model, integrate data of the recognition results of each slice corresponding to the noisy audio, the main energy type, the meteorological data, the sound source localization data, the physical location of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data to obtain a noise event recognition data unit;
[0030] Send the noise event recognition data unit to the target account, and obtain a target noise event recognition result for the noisy audio based on the feedback data of the target account.
[0031] In some embodiments, when applied to a noise recognition system, the noise recognition system includes the noise sensor, the noise recognition system is associated with a target account, and performing noise event recognition based on the noise source classification result corresponding to the main energy type includes:
[0032] Obtain meteorological data, sound source localization data, the physical location of the noise sensor, the over-standard classification threshold, and sound source investigation situation information during the time period when the noise audio is located;
[0033] Obtain the minute-level leq monitoring data corresponding to the noise audio;
[0034] Based on a preset AI model, integrate the recognition results of each slice corresponding to the noise audio, the main energy type, the second ratio, the meteorological data, the sound source localization data, the physical location of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data to obtain a noise event recognition data unit;
[0035] Send the noise event recognition data unit to the target account, and obtain a target noise event recognition result for the noise audio based on the feedback data of the target account.
[0036] In a second aspect, an embodiment of the present application provides a control device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the noise event recognition method based on main sound source detection as described in the first aspect.
[0037] In a third aspect, an embodiment of the present application further provides an electronic device, including the control device in the second aspect.
[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, storing computer-executable instructions for executing the noise event recognition method based on main sound source detection as described in the first aspect.
[0039] The embodiments of the present application provide a method, system, device, and medium for noise event recognition based on main sound source detection. The method includes: obtaining a noise audio through a noise sensor, performing continuous slicing processing on the noise audio to obtain a plurality of audio slices with the same duration; extracting the time-frequency domain voiceprint features of each audio slice, and inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and recognition to obtain the recognition result of each slice, where the slice recognition result includes the noise source classification result corresponding to the audio slice and the confidence level of the noise source classification result; calculating the energy value of each audio slice, and associating each energy value with the noise source classification result and the confidence level of the corresponding audio slice, where the energy value is the time-domain energy value or the frequency-domain energy value of the audio slice; based on each energy value, the corresponding noise source classification result and confidence level, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noise audio, where the main energy type is used to indicate the type of the noise source with the largest proportion in the noise audio; and performing noise event recognition based on the noise source classification result corresponding to the main energy type. According to the solution provided by the embodiments of the present application, the main noise source is identified in a complex multi-source noise environment through audio slices and the main energy detection algorithm, and the noise source recognition result is determined based on the main noise source, effectively ensuring the accuracy of the noise source recognition result. Description of the Drawings
[0040] Figure 1 is a flowchart of the steps of a method for noise event recognition based on main sound source detection provided by an embodiment of the present application;
[0041] Figure 2 is a structural diagram of a control device provided by another embodiment of the present application. Detailed Embodiments
[0042] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] It can be understood that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims, or the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0044] The noise source identification technology is widely used in noise pollution prevention and control and environmental monitoring. It can analyze the time-frequency characteristics of noise signals, identify different types of noise events, and thus realize the intelligent monitoring, classification, and prediction of noise sources. The existing noise source identification methods generally use simple time-frequency analysis or Fourier transform. However, these methods lack pertinence in the control of noise pollution sources. Due to the complex actual noise environment, especially in the case of multi-noise source aliasing, it is impossible to accurately identify the specific characteristics of the noise source and match the sound source type, resulting in a low accuracy of noise source identification.
[0045] To solve the above problems, the embodiments of the present application provide a noise event identification method, system, device, and medium based on main sound source detection. The method includes: obtaining a noise audio through a noise sensor, performing continuous slicing processing on the noise audio to obtain a plurality of audio slices with the same duration; extracting the time-frequency domain voiceprint features of each audio slice, and inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and identification to obtain the identification result of each slice. The slice identification result includes the noise source classification result corresponding to the audio slice and the confidence level of the noise source classification result; calculating the energy value of each audio slice, and associating each energy value with the noise source classification result and the confidence level of the corresponding audio slice, where the energy value is the time-domain energy value or the frequency-domain energy value of the audio slice; based on each energy value, the corresponding noise source classification result, the confidence level, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noise audio, where the main energy type is used to indicate the type of the noise source with the largest proportion in the noise audio; and performing noise event identification based on the noise source classification result corresponding to the main energy type. According to the solution provided by the embodiments of the present application, the main noise source is identified in a complex multi-source noise environment through audio slices and the main energy detection algorithm, and the noise source identification result is determined based on the main noise source, effectively ensuring the accuracy of the noise source identification result.
[0046] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings.
[0047] Reference Figure 1 , Figure 1 is a flowchart of the steps of a noise event identification method based on main sound source detection provided by an embodiment of the present application. The embodiments of the present application provide a noise event identification method based on main sound source detection. The method includes but is not limited to the following steps:
[0048] Step S10: Obtain a noise audio through a noise sensor, and perform continuous slicing processing on the noise audio to obtain a plurality of audio slices with the same duration.
[0049] Specifically, the duration range of the audio slice corresponding to this embodiment is between 2 seconds and 6 seconds, which meets the audio duration condition for identifying most types of noise sources and provides an effective data basis for obtaining the slice recognition result subsequently. And it should be noted that this embodiment does not limit the specific duration of the noisy audio. The noisy audio can be a complete 1-minute audio. When the audio slice duration is 2 seconds, there are 30 audio slices per minute.
[0050] Step S20: Extract the time-frequency domain voiceprint features of each audio slice, and input each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and recognition to obtain each slice recognition result. The slice recognition result includes the noise source classification result corresponding to the audio slice and the confidence level of the noise source classification result.
[0051] Specifically, this embodiment does not involve the improvement of the model. The deep learning model can be a DenseNet model, a ResNet model, an RNN model, etc., and those skilled in the art can select according to actual needs.
[0052] It can be understood that since the current noise source is unknown, there is a possibility that the noise source changes dynamically in time (for example, there are different noise sources in different time periods), and some noise sources show obvious characteristics in a short time (such as transient noise). In this case, slicing a complete long-duration noisy audio can capture these changes, thereby improving the recognition accuracy. Moreover, the audio slice may have a more single noise source, reducing the interference of multiple noise sources existing simultaneously, making it easier for the model to focus on local features. At the same time, slicing can be regarded as a data augmentation method, which can provide more diverse training samples for the model, thereby improving the generalization ability. That is to say, in this case, continuously slicing a noisy audio and separately identifying the noise source for each audio slice has a higher recognition result accuracy than identifying the noise source of an unsliced complete noisy audio. And the obtained slice recognition results corresponding to each audio slice (that is, the noise source classification results corresponding to each audio slice and the confidence levels of the noise source classification results) can provide an effective data basis for determining the main noise source subsequently.
[0053] Step S30: Calculate the energy value of each audio slice, and associate each energy value with the noise source classification result and confidence level of the corresponding audio slice, where the energy value is the time-domain energy value or the frequency-domain energy value of the audio slice.
[0054] Step S40: Based on each energy value, the corresponding noise source classification result and confidence level, and a preset main energy detection algorithm, calculate the main energy type corresponding to the noisy audio, where the main energy type is used to indicate the type of the noise source with the largest proportion in the noisy audio.
[0055] Specifically, in this embodiment, the energy value of the audio slice is calculated as the time-domain energy value or the frequency-domain energy value of the audio slice, that is, it is obtained by calculating the power spectrum or the time-domain envelope integral. The energy value of the audio slice is used to indicate the intensity or loudness of the audio signal corresponding to the audio slice. In the field of noise recognition, silent segments or noise segments can be detected through the energy value. The parts with higher energy values usually correspond to human voice segments, and the parts with lower energy values correspond to silence or environmental noise.
[0056] It can be understood that by calculating the energy values of the audio slices and associating each energy value with the noise source classification result and confidence level corresponding to the corresponding audio slice, it can help determine the noise intensity of the noise source type corresponding to the noise source recognition result of the audio slice, providing an effective data basis for determining the main noise source type subsequently.
[0057] It can be understood that after calculating the main energy type of the noisy audio based on each energy value, the corresponding noise source classification result and confidence level, and a preset main energy detection algorithm, it can provide effective support for accurately determining the noise event.
[0058] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0059] Step S41, modifying the noise source type corresponding to the noise source classification result with a confidence level less than the confidence level threshold to the reference noise source type;
[0060] Step S42, adding the energy values corresponding to the noise source classification results of the same type to obtain multiple first reference energy values;
[0061] Step S43, determining the noise source type corresponding to the noise source classification result with the largest first reference energy value as the main energy type.
[0062] Specifically, the confidence level threshold in this embodiment is 85%, and it can also be 90%. Those skilled in the art can adjust it according to the actual situation.
[0063] Specifically, the reference noise source type in this embodiment is the environmental noise type.
[0064] It can be understood that in this embodiment, by setting a confidence threshold and correcting the noise source type corresponding to the classification result of the noise source with a confidence lower than the confidence threshold to the reference noise source type, in the current data, each energy value corresponding to an audio slice corresponds to a highly credible noise source type. In this case, the energy values corresponding to the classification results of the same type of noise source are added together to obtain multiple first reference energy values. At this time, different first reference energy values correspond to different noise source types. Further, the values of all the first reference energy values are sorted from largest to smallest, and the noise source type corresponding to the classification result of the first reference energy value with the largest value is determined as the main energy type. In this way, a main energy type with a relatively high confidence can be determined.
[0065] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0066] Step S44, performing multiplicative weighting on each energy value and the corresponding confidence to obtain each second reference energy value;
[0067] Step S45, determining the noise source type corresponding to the classification result of the second reference energy value with the largest value as the main energy type.
[0068] It can be understood that when the slice recognition results and energy values of each audio slice are obtained in Step S20 and Step S30, the energy values are 100% accurate, but the noise source classification results corresponding to the slice recognition results are not completely accurate. Therefore, to ensure the accuracy of the main energy type, in this embodiment, all the energy values are multiplicatively weighted with the confidence of their associated noise source classification results to obtain each second reference energy value. Further, the values of all the second reference energy values are sorted from largest to smallest, and the noise source type corresponding to the classification result of the second reference energy value with the largest value is determined as the main energy type. In this way, a main energy type with a relatively high confidence can be determined.
[0069] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0070] Step S46, calculating the leq value of the corresponding audio slice based on each energy value;
[0071] Step S47, correcting the noise source classification result corresponding to the confidence lower than the confidence threshold to the reference noise source type, adding the leq values corresponding to the classification results of the same type of noise source to obtain multiple first reference leq values, and determining the noise source type corresponding to the classification result of the first reference leq value with the largest value as the main energy type;
[0072] Or,
[0073] Step S48: Multiply each leq value with the confidence level of the corresponding audio slice for weighted processing to obtain each second reference leq value, and determine the noise source type corresponding to the classification result of the noise source with the largest second reference leq value as the main energy type.
[0074] Specifically, the leq value of the audio slice corresponds to the data obtained by a noise monitor or a sound level meter. The operation of calculating the leq value of the corresponding audio slice based on each energy value in this embodiment is to quantify the energy value, improving the accuracy of subsequent determination of the main energy type. The leq value of the audio slice in this embodiment is obtained by converting the energy value to a second-level leq value sequence and calculating the leq value of the slice time. Taking a 2-second slice as an example, the 2-second slice leq value is calculated according to the exponential weighting algorithm in HJ3096 and HJ906 standards.
[0075] It can be understood that after obtaining the leq value of the audio slice, by correcting the classification result of the noise source corresponding to the confidence level less than the confidence level threshold to the reference noise source type, adding the leq values corresponding to the classification results of the noise sources of the same type to obtain multiple first reference leq values, and determining the noise source type corresponding to the classification result of the noise source with the largest first reference leq value as the main energy type; or by multiplying each leq value with the confidence level of the corresponding audio slice for weighted processing to obtain each second reference leq value, and determining the noise source type corresponding to the classification result of the noise source with the largest second reference leq value as the main energy type, a main energy type with higher accuracy can be obtained. It should be noted that the technical principle of determining the main energy type in this embodiment can refer to the content of the embodiments of steps S41 to S43 and steps S44 to S45 above. The difference is that the data source for judgment and sorting in this embodiment is the leq value of the audio slice, which will not be elaborated here. Since the leq value is a quantified value and the data is more accurate, the credibility of the main energy type obtained in this embodiment is higher.
[0076] In addition, in some embodiments, the noise event recognition method based on main sound source detection in this embodiment is applied to a noise recognition system. After executing step S40, the noise event recognition method based on main sound source detection further includes but is not limited to the following steps:
[0077] Step S61: Determine the number of noise source types in the classification results of the noise sources corresponding to all audio slices corresponding to the noisy audio;
[0078] Step S62: Determine the current available system resource amount of the noise recognition system;
[0079] Step S63: When the number of noise source types is greater than a preset number threshold and the available system resource amount is greater than a preset resource amount threshold, calculate the energy proportion of the main energy sound source corresponding to the main energy type in each audio slice to obtain a plurality of first proportions, and sum all the first proportions to obtain a second proportion, where the second proportion is used to indicate the energy proportion of the main energy sound source corresponding to the main energy type in the noisy audio.
[0080] In addition, it should be noted that this embodiment also provides a solution for calculating on demand the energy proportion (i.e., the second proportion) of the main energy sound source corresponding to the main energy type in the noisy audio, so that while determining the main energy type with the largest proportion in the noisy audio, the corresponding proportion value of this main energy type can also be determined, thereby providing a more effective data basis for obtaining an accurate noise event recognition result subsequently.
[0081] It can be understood that calculating the second proportion requires first calculating the energy proportion of the main energy sound source corresponding to the main energy type in each audio slice to obtain each first proportion, and then summing all the first proportions to obtain the second proportion, which requires consuming more system computing resources, and calculating the second proportion is more necessary when there are more noise source types and sufficient system computing resources. Therefore, before calculating the second proportion in this embodiment, first determine the number of noise source types in the noise source classification results corresponding to all audio slices of the noisy audio, and determine the available system resource amount of the noise recognition system currently. When the number of noise source types is greater than a preset number threshold and the available system resource amount is greater than a preset resource amount threshold, the calculation of the second proportion is started.
[0082] Step S50: Perform noise event recognition based on the noise source classification result corresponding to the main energy type.
[0083] Specifically, in some embodiments, the noise event recognition method based on main sound source detection in this embodiment is applied to a noise recognition system, the noise recognition system includes a noise sensor, and the noise recognition system is associated with a target account. Figure 1 Step S50 includes but is not limited to the following steps:
[0084] Step S51: Obtain meteorological data, sound source localization data, the physical location of the noise sensor, the over-standard grading threshold, and sound source investigation situation information during the time period when the noisy audio is located;
[0085] Step S52: Obtain the minute-level leq monitoring data corresponding to the noisy audio.
[0086] Step S53: Integrate the recognition results of each slice corresponding to the noisy audio, the main energy type, meteorological data, sound source localization data, the physical location of the noise sensor, the over-standard grading threshold, and the sound source investigation information based on a preset AI model to obtain a noise event recognition data unit;
[0087] Step S54: Send the noise event recognition data unit to the target account, and obtain the target noise event recognition result for the noisy audio based on the feedback data of the target account.
[0088] Specifically, in some embodiments, the noise event recognition method based on main sound source detection of this embodiment is applied to a noise recognition system. The noise recognition system includes a noise sensor, and the noise recognition system is associated with a target account. Figure 1 Step S50 includes but is not limited to the following steps:
[0089] Step S55: Obtain meteorological data, sound source localization data, the physical location of the noise sensor, the over-standard grading threshold, and the sound source investigation information during the time period when the noisy audio is located;
[0090] Step S56: Obtain the minute-level leq monitoring data corresponding to the noisy audio;
[0091] Step S57: Integrate the recognition results of each slice corresponding to the noisy audio, the main energy type, the second ratio, meteorological data, sound source localization data, the physical location of the noise sensor, the over-standard grading threshold, the sound source investigation information, and the minute-level leq monitoring data based on a preset AI model to obtain a noise event recognition data unit;
[0092] Step S58: Send the noise event recognition data unit to the target account, and obtain the target noise event recognition result for the noisy audio based on the feedback data of the target account.
[0093] Specifically, the target account of this embodiment corresponds to the system administrator, who can make a decision on the final noise event recognition result based on the content in the noise event recognition data unit.
[0094] Specifically, the physical location of the noise sensor in this embodiment refers to the physical location of the detection instrument that obtains the noisy audio, and the over-standard grading threshold refers to the over-standard limit value of the functional area where the noise sensor is located or the over-standard factory boundary emission limit value.
[0095] Specifically, the output format of the noise event recognition data unit in this embodiment is determined according to actual requirements. The output format of the noise event recognition data unit in this embodiment is as follows: {time information of the noise audio, physical location of the noise sensor, slice recognition result, main energy type, meteorological data, sound source localization data, minute-level leq monitoring data, over-limit classification threshold, sound source investigation situation information...}, or, {time information of the noise audio, physical location of the noise sensor, slice recognition result, main energy type, second proportion, meteorological data, sound source localization data, minute-level leq monitoring data, over-limit classification threshold, sound source investigation situation information...}.
[0096] It can be understood that after determining the main energy type corresponding to the noise audio, in this embodiment, by obtaining the meteorological data, sound source localization data, physical location of the noise sensor, over-limit classification threshold, and sound source investigation situation information within the time period corresponding to the noise audio, and obtaining the minute-level leq monitoring data corresponding to the noise audio, and integrating the data of each slice recognition result, main energy type, meteorological data, sound source localization data, physical location of the noise sensor, over-limit classification threshold, sound source investigation situation information, and minute-level leq monitoring data corresponding to the noise audio through a preset AI model, or integrating the data by adding the second proportion as needed on this basis, a noise event recognition data unit is obtained. In addition to the main energy type, the noise event recognition data unit also includes meteorological data within the time period corresponding to the noise audio, as well as data such as sound source localization and physical location of the detection point, which can be used by the administrator corresponding to the target account to make an effective and accurate judgment in combination, whether the main energy type is correct, whether it is affected by environmental factors or meteorological factors, or how to perform subsequent noise reduction measures.
[0097] As Figure 2 shown, Figure 2 is the structural diagram of a control device provided by an embodiment of the present application. The present invention also provides a control device 200, including:
[0098] A processor 210, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0099] The memory 220 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 220 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 220 and are called by the processor 210 to execute the noise event recognition method based on primary sound source detection in the embodiments of this application;
[0100] The input / output interface 230 is used to implement information input and output;
[0101] The communication interface 240 is used to implement communication and interaction between this device and other devices. Communication can be achieved through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0102] The bus 250 transmits information between various components of the device (such as the processor 210, the memory 220, the input / output interface 230, and the communication interface 240);
[0103] Among them, the processor 210, the memory 220, the input / output interface 230, and the communication interface 240 achieve communication connections with each other inside the device through the bus 250.
[0104] In addition, the embodiments of this application also provide an electronic device, including the control device 200 in the above embodiments.
[0105] In addition, the embodiments of this application also provide a storage medium. The storage medium is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned noise event recognition method based on primary sound source detection.
[0106] A memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0107] Those of ordinary skill in the art can understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disc (DVD), or other optical disc storage, magnetic cartridges, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0108] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A noise event recognition method based on main sound source detection, characterized in that Including: Obtaining noisy audio through a noise sensor, performing continuous slicing processing on the noisy audio to obtain a plurality of audio slices with the same duration; Extracting the time-frequency domain voiceprint features of each of the audio slices, and inputting each of the time-frequency domain voiceprint features into a preset deep learning model for noise source classification and recognition to obtain the recognition result of each slice, where the slice recognition result includes the noise source classification result corresponding to the audio slice and the confidence of the noise source classification result; Calculating the energy value of each of the audio slices, and associating each of the energy values with the noise source classification result and the confidence of the corresponding audio slice, where the energy value is the time-domain energy value or the frequency-domain energy value of the audio slice; Based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio, where the main energy type is used to indicate the type of the noise source with the largest proportion in the noisy audio; Performing noise event recognition based on the noise source classification result corresponding to the main energy type.
2. The noise event recognition method based on main sound source detection according to claim 1, wherein Based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio, including: Modifying the noise source type corresponding to the noise source classification result corresponding to the confidence less than the confidence threshold to the reference noise source type; Adding the energy values corresponding to the noise source classification results of the same type to obtain a plurality of first reference energy values; Determining the noise source type corresponding to the noise source classification result corresponding to the largest first reference energy value as the main energy type.
3. The noise event recognition method based on main sound source detection according to claim 1, characterized in that Based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio, including: Performing multiplicative weighting processing on each of the energy values and the corresponding confidence to obtain each second reference energy value; Determining the noise source type corresponding to the noise source classification result corresponding to the largest second reference energy value as the main energy type.
4. The method for identifying a noise event based on primary sound source detection according to claim 1, wherein, Based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, calculating the main energy type corresponding to the noisy audio, including: Calculating the leq value of the corresponding audio slice based on each of the energy values; Modifying the noise source classification result corresponding to the confidence less than the confidence threshold to the reference noise source type, adding the leq values corresponding to the noise source classification results of the same type to obtain a plurality of first reference leq values, and determining the noise source type corresponding to the noise source classification result corresponding to the largest first reference leq value as the main energy type; Or, Performing multiplicative weighting processing on each of the leq values and the confidence of the corresponding audio slice to obtain each second reference leq value, and determining the noise source type corresponding to the noise source classification result corresponding to the largest second reference leq value as the main energy type.
5. The method for identifying a noise event based on main sound source detection according to claim 1, characterized in that, Applied to a noise recognition system, after calculating the main energy type corresponding to the noisy audio based on each of the energy values, the corresponding noise source classification results, the confidence level, and a preset main energy detection algorithm, the method further includes: Determining the number of noise source types in the noise source classification results corresponding to all the audio slices corresponding to the noisy audio; Determining the current available system resource amount of the noise recognition system; When the number of noise source types is greater than a preset number threshold and the available system resource amount is greater than a preset resource amount threshold, calculating the energy proportion of the main energy sound source corresponding to the main energy type in each of the audio slices to obtain a plurality of first proportions, and summing all the first proportions to obtain a second proportion, where the second proportion is used to indicate the energy proportion of the main energy sound source corresponding to the main energy type in the noisy audio.
6. The noise event recognition method based on main sound source detection according to claim 1, characterized in that Applied to a noise recognition system, the noise recognition system includes the noise sensor, the noise recognition system is associated with a target account, and performing noise event recognition based on the noise source classification result corresponding to the main energy type includes: Obtaining meteorological data, sound source localization data, the physical location of the noise sensor, an over-standard classification threshold, and sound source investigation situation information during the time period when the noisy audio is located; Obtaining the minute-level leq monitoring data corresponding to the noisy audio; Based on a preset AI model, integrating data of each slice recognition result, main energy type, the meteorological data, the sound source localization data, the physical location of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data corresponding to the noisy audio to obtain a noise event recognition data unit; Sending the noise event recognition data unit to the target account, and obtaining a target noise event recognition result for the noisy audio based on the feedback data of the target account.
7. The method for identifying a noise event based on main sound source detection according to claim 5, characterized in that Applied to a noise recognition system, the noise recognition system includes the noise sensor, the noise recognition system is associated with a target account, and performing noise event recognition based on the noise source classification result corresponding to the main energy type includes: Obtaining meteorological data, sound source localization data, the physical location of the noise sensor, an over-standard classification threshold, and sound source investigation situation information during the time period when the noisy audio is located; Obtaining the minute-level leq monitoring data corresponding to the noisy audio; Based on a preset AI model, integrating data of each slice recognition result, main energy type, the second proportion, the meteorological data, the sound source localization data, the physical location of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data corresponding to the noisy audio to obtain a noise event recognition data unit; Sending the noise event recognition data unit to the target account, and obtaining a target noise event recognition result for the noisy audio based on the feedback data of the target account.
8. A control device, characterized in that, Comprising at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the method for identifying a noise event based on main sound source detection according to any one of claims 1 to 7.
9. An electronic device, characterized in that, Comprising the control device according to claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method for identifying a noise event based on main sound source detection according to any one of claims 1 to 7.
Citation Information
Patent Citations
Device and method for audio classification and audio processing
CN109616142A
Method and system for intelligently identifying environmental noise
CN115662464A
Environmental noise detection method and device, electronic equipment and storage medium
CN117352001A
Method and system for monitoring and analyzing regional noise
CN119334461A
Call recording
US20180324293A1