Noise event recognition method, system, apparatus, and medium based on primary sound source detection
By performing continuous slicing processing on noisy audio and using the main energy detection algorithm, the type of noise source with the largest proportion in the noisy audio is identified, which solves the problem of accuracy in noise source identification under multiple noise source aliasing and realizes high-precision noise source identification in complex environments.
Patent Information
- Application Number
- CN202510287251.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing noise source identification methods cannot accurately identify the specific features of noise sources when multiple noise sources are mixed, resulting in low identification accuracy.
Noise audio is acquired by a noise sensor, and continuous slicing is performed to extract time-frequency domain voiceprint features. These features are then input into a deep learning model for classification and recognition. Energy values and confidence levels are calculated, and the main energy detection algorithm is used to identify the noise source type that accounts for the largest proportion in the noise audio.
Accurately identifying the main noise sources in complex multi-source noise environments improves the accuracy of noise source identification results.
Smart Images

Figure CN120279941B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of noise recognition, and in particular to a noise event recognition method, system, device and medium based on main sound source detection. BACKGROUND
[0002] Noise source recognition technology is widely used in noise pollution prevention and environmental monitoring, and can analyze the time-frequency characteristics of noise signals, identify different types of noise events, and thus realize intelligent monitoring, classification and prediction of noise sources. The existing noise source recognition methods generally use simple time-frequency analysis or Fourier transform, but these methods lack pertinence for noise pollution source control. Due to the complexity of the actual noise environment, especially in the case of multiple noise source mixing, it is impossible to accurately identify the noise source to the specific feature and match the sound source type, thereby resulting in low accuracy of noise source recognition. SUMMARY
[0003] The embodiments of the present application provide a noise event recognition method, system, device and medium based on main sound source detection, which identifies the main noise source and determines the accurate noise source recognition result based on the main noise source.
[0004] In a first aspect, the embodiments of the present application provide a noise event recognition method based on main sound source detection, comprising:
[0005] obtaining a noise audio through a noise sensor, performing continuous slicing processing on the noise audio to obtain a plurality of audio slices with the same time length;
[0006] extracting time-frequency domain voiceprint features of each audio slice, and inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and recognition to obtain a slice recognition result, the slice recognition result including a noise source classification result corresponding to the audio slice and a confidence degree of the noise source classification result;
[0007] calculating energy values of each audio slice, and associating each energy value to the noise source classification result and the confidence degree of the corresponding audio slice, wherein the energy value is a time domain energy value or a frequency domain energy value of the audio slice;
[0008] based on each energy value, the corresponding noise source classification result and the confidence degree, and a preset main energy detection algorithm, calculating a main energy type corresponding to the noise audio, wherein the main energy type is used to indicate the noise source type with the largest proportion in the noise audio;
[0009] performing noise event recognition based on the noise source classification result corresponding to the main energy type.
[0010] In some embodiments, the main energy type corresponding to the noise audio is calculated based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, including:
[0011] The noise source type corresponding to the noise source classification result corresponding to the confidence less than the confidence threshold is corrected to a reference noise source type;
[0012] The energy values corresponding to the noise source classification results of the same type are added to obtain a plurality of first reference energy values;
[0013] The noise source type corresponding to the noise source classification result corresponding to the first reference energy value with the maximum value is determined as the main energy type.
[0014] In some embodiments, the main energy type corresponding to the noise audio is calculated based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, including:
[0015] Each of the energy values is multiplied and weighted with the corresponding confidence to obtain each second reference energy value;
[0016] The noise source type corresponding to the noise source classification result corresponding to the second reference energy value with the maximum value is determined as the main energy type.
[0017] In some embodiments, the main energy type corresponding to the noise audio is calculated based on each of the energy values, the corresponding noise source classification result and the confidence, and a preset main energy detection algorithm, including:
[0018] The leq value of the corresponding audio slice is calculated based on each of the energy values;
[0019] The noise source classification result corresponding to the confidence less than the confidence threshold is corrected to a reference noise source type, the leq values corresponding to the noise source classification results of the same type are added to obtain a plurality of first reference leq values, and the noise source type corresponding to the noise source classification result corresponding to the first reference leq value with the maximum value is determined as the main energy type;
[0020] Or,
[0021] Each of the leq values is multiplied and weighted with the confidence of the corresponding audio slice to obtain each second reference leq value, and the noise source type corresponding to the noise source classification result corresponding to the second reference leq value with the maximum value is determined as the main energy type.
[0022] In some embodiments, applied to a noise identification system, after calculating the main energy type corresponding to the noise audio based on each of the energy values, the corresponding noise source classification result, the confidence, and a preset main energy detection algorithm, the method further comprises:
[0023] determining the number of noise source types in the noise source classification result corresponding to all of the audio slices of the noise audio;
[0024] determining the amount of available system resources of the noise identification system at present;
[0025] when the number of noise source types is greater than a preset number threshold, and the amount of available system resources is greater than a preset resource amount threshold, calculating the energy proportion of the main energy source corresponding to the main energy type in each of the audio slices to obtain a plurality of first proportions, and summing all of the first proportions to obtain a second proportion, wherein the second proportion is used to indicate the energy proportion of the main energy source corresponding to the main energy type in the noise audio.
[0026] In some embodiments, applied to a noise identification system, the noise identification system comprises the noise sensor, and the noise identification system is associated with a target account, and the noise event identification based on the noise source classification result corresponding to the main energy type comprises:
[0027] obtaining meteorological data, sound source positioning data, a physical position of the noise sensor, an over-standard classification threshold, and sound source investigation situation information in a time period in which the noise audio is located;
[0028] obtaining minute-level leq monitoring data corresponding to the noise audio;
[0029] integrating each of the slice identification results, the main energy type, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data corresponding to the noise audio based on a preset AI model to obtain a noise event identification data unit;
[0030] sending the noise event identification data unit to the target account, and obtaining a target noise event identification result for the noise audio based on feedback data of the target account.
[0031] In some embodiments, applied to a noise identification system, the noise identification system comprises the noise sensor, and the noise identification system is associated with a target account, and the noise event identification based on the noise source classification result corresponding to the main energy type comprises:
[0032] acquire meteorological data, sound source positioning data, physical position of the noise sensor, over-standard classification threshold and sound source investigation information in a time period in which the noise audio is located;
[0033] acquire minute-level leq monitoring data corresponding to the noise audio;
[0034] integrate the noise event recognition data unit based on the preset AI model, the slice recognition result corresponding to the noise audio, the main energy type, the second proportion, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the over-standard classification threshold, the sound source investigation information and the minute-level leq monitoring data.
[0035] send the noise event recognition data unit to the target account, and obtain a target noise event recognition result for the noise audio based on feedback data of the target account.
[0036] In a second aspect, an embodiment of the present application provides a control device, including at least one control processor and a memory in communication connection with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the noise event recognition method based on main sound source detection according to the first aspect.
[0037] In a third aspect, an embodiment of the present application further provides an electronic device including the control device of the second aspect.
[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer executable instructions for executing the noise event recognition method based on main sound source detection according to the first aspect.
[0039] The embodiment of the present application provides a noise event identification method, system and device based on main sound source detection, and a medium. The method comprises the following steps: acquiring a noise audio through a noise sensor, performing continuous slicing processing on the noise audio to obtain a plurality of audio slices with the same time length; extracting time-frequency domain voiceprint features of each audio slice, inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and identification to obtain a slice identification result, wherein the slice identification result comprises a noise source classification result corresponding to the audio slice and a confidence degree of the noise source classification result; calculating energy values of each audio slice, and associating each energy value with the noise source classification result and the confidence degree of the corresponding audio slice, wherein the energy value is a time domain energy value or a frequency domain energy value of the audio slice; based on each energy value, the noise source classification result and the confidence degree corresponding to the energy value, and a preset main energy detection algorithm, a main energy type corresponding to the noise audio is calculated, wherein the main energy type is used to indicate a noise source type with the largest proportion in the noise audio; and performing noise event identification based on the noise source classification result corresponding to the main energy type. According to the scheme provided by the embodiment of the present application, the main noise source is identified in a complex multi-source noise environment through the audio slice and the main energy detection algorithm, and the noise source identification result is determined based on the main noise source, so that the accuracy of the noise source identification result is effectively ensured. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 is a step flow chart of a noise event identification method based on main sound source detection provided by an embodiment of the present application;
[0041] Figure 2 is a structural diagram of a control device provided by another embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0043] It can be understood that, although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flow chart, in some cases, the steps shown or described can be executed in a manner different from the module division in the device or the order in the flow chart. The terms "first", "second", etc. in the specification, claims or above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0044] Noise source identification technology is widely used in noise pollution prevention and environmental monitoring, which can analyze the time-frequency characteristics of noise signals, identify different types of noise events, and thus realize intelligent monitoring, classification and prediction of noise sources. The existing noise source identification methods generally use simple time-frequency analysis or Fourier transform, but these methods lack pertinence for noise pollution source control. Due to the complexity of the actual noise environment, especially in the case of multiple noise source mixing, it is difficult to accurately identify the noise source to the specific feature and match to the sound source type, resulting in low accuracy of noise source identification.
[0045] To solve the above problems, the embodiments of the present application provide a noise event identification method, system, device and medium based on main sound source detection. The method comprises: acquiring noise audio through a noise sensor, continuously slicing the noise audio to obtain a plurality of audio slices with the same time length; extracting the time-frequency domain voiceprint features of each audio slice and inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and identification to obtain a slice identification result, which includes a noise source classification result corresponding to the audio slice and a confidence degree of the noise source classification result; calculating the energy value of each audio slice and associating each energy value with the noise source classification result and the confidence degree of the corresponding audio slice, wherein the energy value is the time energy value or the frequency energy value of the audio slice; based on each energy value, the corresponding noise source classification result and the confidence degree, and a preset main energy detection algorithm, the main energy type corresponding to the noise audio is calculated, wherein the main energy type is used to indicate the noise source type with the largest proportion in the noise audio; and performing noise event identification based on the noise source classification result corresponding to the main energy type. According to the scheme provided by the embodiments of the present application, the main noise source is identified in a complex multi-source noise environment through audio slicing and main energy detection algorithm, and the noise source identification result is determined based on the main noise source, which effectively ensures the accuracy of the noise source identification result.
[0046] The embodiments of the present application will be further described below with reference to the accompanying drawings.
[0047] Reference Figure 1 , Figure 1 is a step flowchart of a noise event identification method based on main sound source detection provided by an embodiment of the present application. The embodiment of the present application provides a noise event identification method based on main sound source detection, which comprises but is not limited to the following steps:
[0048] Step S10, acquiring noise audio through a noise sensor, continuously slicing the noise audio to obtain a plurality of audio slices with the same time length.
[0049] Specifically, the time length range of the slice time length corresponding to the audio slice of the embodiment is between 2 seconds and 6 seconds, which meets the audio time length condition of most types of noise source identification, and provides an effective data basis for subsequent acquisition of slice identification results. It should be noted that the embodiment does not limit the specific time length of the noise audio, and the noise audio can be a complete audio of 1 minute. When the audio slice time length is 2 seconds, there are 30 audio slices corresponding to each minute.
[0050] In step S20, the time-frequency domain voiceprint features of each audio slice are extracted, and each time-frequency domain voiceprint feature is input into a preset deep learning model for noise source classification and identification to obtain each slice identification result. The slice identification result includes the noise source classification result corresponding to the audio slice and the confidence of the noise source classification result.
[0051] Specifically, the embodiment does not involve improvement of the model, and the deep learning model can be a DenseNet model, a ResNet model or an RNN model, etc. Those skilled in the art can select according to actual needs.
[0052] It can be understood that, since the current noise source is unknown, there is a possibility that the noise source changes dynamically in time (for example, different noise sources in different time periods), and some noise sources show obvious characteristics in a short time (such as transient noise). In this case, slicing a complete time length of noise audio can capture these changes, thereby improving the recognition accuracy. In addition, the audio slice can have a more single noise source, which reduces the interference of multiple noise sources existing at the same time, so that the model can focus on local features more easily. At the same time, the slice can be regarded as a kind of data enhancement method, which can provide more diversified training samples for the model, thereby improving the generalization ability. That is to say, in this case, a continuous slice is performed on a piece of noise audio, and noise source identification is performed on each audio slice respectively. The accuracy of the identification result is higher than that of the noise source identification on the complete piece of noise audio without slicing. In addition, the slice identification result corresponding to each audio slice (i.e., the noise source classification result corresponding to each audio slice and the confidence of the noise source classification result) can provide an effective data basis for subsequent determination of the main noise source.
[0053] In step S30, the energy values of each audio slice are calculated, and each energy value is associated with the noise source classification result and the confidence of the corresponding audio slice. The energy value is a time energy value or a frequency energy value of the audio slice.
[0054] In step S40, based on each energy value, the corresponding noise source classification result and confidence, and a preset main energy detection algorithm, the main energy type corresponding to the noise audio is calculated. The main energy type is used to indicate the noise source type with the largest proportion in the noise audio.
[0055] Specifically, the embodiment calculates the energy value of the audio slice as the time domain energy value or the frequency domain energy value of the audio slice, i.e., obtained by calculating the power spectrum or the time domain envelope integral, the energy value of the audio slice is used to indicate the intensity or loudness of the audio signal corresponding to the audio slice, in the noise recognition field, the energy value can be used to detect the silent section or the noise section, and the part with higher energy value usually corresponds to the human voice section, and the part with lower energy value corresponds to the silence or environmental noise.
[0056] It can be understood that, by calculating the energy value of the audio slice and associating each energy value with the noise source classification result and the confidence of the corresponding audio slice, the noise intensity of the noise source type corresponding to the noise source recognition result of the audio slice can be determined, thereby providing an effective data basis for subsequent determination of the main noise source type.
[0057] It can be understood that, after calculating the main energy type of the noise audio based on the energy value, the corresponding noise source classification result and the confidence, and the preset main energy detection algorithm, effective support can be provided for accurately determining the noise event.
[0058] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0059] Step S41, correcting the noise source type corresponding to the noise source classification result corresponding to the confidence less than the confidence threshold to the reference noise source type;
[0060] Step S42, adding the energy values corresponding to the noise source classification results of the same type to obtain a plurality of first reference energy values;
[0061] Step S43, determining the noise source type corresponding to the noise source classification result corresponding to the first reference energy value with the largest value as the main energy type.
[0062] Specifically, the confidence threshold of the embodiment is 85%, and can also be 90%, which can be adjusted by those skilled in the art according to the actual situation.
[0063] Specifically, the reference noise source type of the embodiment is the environmental noise type.
[0064] It can be understood that, after the noise source classification result corresponding to the confidence less than the confidence threshold is corrected to the reference noise source type, each audio slice in the current data corresponds to a noise source type with high confidence. In this case, the energy values corresponding to the noise source classification results of the same type are added to obtain a plurality of first reference energy values. At this time, different first reference energy values correspond to different noise source types. Further, the noise source classification result corresponding to the noise source type with the largest first reference energy value is determined as the main energy type. In this way, a main energy type with high confidence can be determined.
[0065] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0066] Step S44, multiplying each energy value by the corresponding confidence to obtain a plurality of second reference energy values;
[0067] Step S45, determining the noise source classification result corresponding to the noise source type with the largest second reference energy value as the main energy type.
[0068] It can be understood that, after the noise source classification result corresponding to the confidence less than the confidence threshold is corrected to the reference noise source type, each audio slice in the current data corresponds to a noise source type with high confidence. In this case, the energy values corresponding to the noise source classification results of the same type are added to obtain a plurality of first reference energy values. At this time, different first reference energy values correspond to different noise source types. Further, the noise source classification result corresponding to the noise source type with the largest first reference energy value is determined as the main energy type. In this way, a main energy type with high confidence can be determined.
[0069] Specifically, in some embodiments, Figure 1 Step S40 includes but is not limited to the following steps:
[0070] Step S46, calculating the leq value of each audio slice based on the energy value of the audio slice;
[0071] Step S47, correcting the noise source classification result corresponding to the confidence less than the confidence threshold to the reference noise source type, adding the leq values corresponding to the noise source classification results of the same type to obtain a plurality of first reference leq values, and determining the noise source classification result corresponding to the noise source type with the largest first reference leq value as the main energy type.
[0072] Or,
[0073] Step S48, multiply each leq value with the confidence of the corresponding audio slice to obtain each second reference leq value, and determine the noise source type corresponding to the noise source classification result corresponding to the second reference leq value with the largest value as the main energy type.
[0074] Specifically, the leq value corresponding to the audio slice is the data obtained by the noise monitor or the sound level meter, and the operation of calculating the leq value of the corresponding audio slice based on each energy value in the embodiment is to quantify the energy value, which improves the accuracy of subsequent determination of the main energy type. The leq value of the audio slice in the embodiment is obtained by converting the energy value into a second-level leq value sequence, and calculating the leq value of the slice time. Taking a 2-second slice as an example, the 2-second slice leq value is calculated according to the exponential weighting algorithm in HJ3096 and HJ906 standards.
[0075] It can be understood that after obtaining the leq value of the audio slice, the noise source classification result corresponding to the confidence less than the confidence threshold is corrected to the reference noise source type, the leq values corresponding to the noise source classification results of the same type are added to obtain a plurality of first reference leq values, and the noise source type corresponding to the noise source classification result corresponding to the first reference leq value with the largest value is determined as the main energy type; or by multiplying each leq value with the confidence of the corresponding audio slice to obtain each second reference leq value, and determining the noise source type corresponding to the noise source classification result corresponding to the second reference leq value with the largest value as the main energy type, the main energy type with higher accuracy can be obtained. It should be noted that the technical principle of determining the main energy type in the embodiment can refer to the contents of the embodiments of steps S41 to S43 and steps S44 to S45, and the difference between the embodiments is that the data source for judgment and sorting in the embodiment is the leq value of the audio slice, which will not be described in detail here. Since the leq value is a quantized value, the data is more accurate, and therefore the main energy type obtained in the embodiment has higher reliability.
[0076] In addition, in some embodiments, the noise event recognition method based on main sound source detection of the embodiment is applied to a noise recognition system, and after step S40, the noise event recognition method based on main sound source detection further includes but is not limited to the following steps:
[0077] Step S61, determining the number of noise source types in the noise source classification results corresponding to all audio slices corresponding to the noise audio;
[0078] Step S62, determining the amount of available system resources of the noise recognition system at present;
[0079] Step S63, when the number of noise source types is greater than the preset number threshold, and the amount of available system resources is greater than the preset resource amount threshold, the energy proportion of the main energy type corresponding to the main energy sound source in each audio slice is calculated to obtain a plurality of first proportions, and the second proportion is obtained by summing all the first proportions, wherein the second proportion is used to indicate the energy proportion of the main energy type corresponding to the main energy sound source in the noise audio.
[0080] In addition, it should be noted that the embodiment also provides a scheme for calculating the energy proportion (i.e., the second proportion) of the main energy type corresponding to the main energy sound source in the noise audio on demand, so that the main energy type with the highest proportion in the noise audio can be determined, and the proportion value corresponding to the main energy type is also determined, thereby providing a more effective data basis for obtaining an accurate noise event recognition result subsequently.
[0081] It can be understood that calculating the second proportion requires first calculating the energy proportion of the main energy type corresponding to the main energy sound source in each audio slice to obtain each first proportion, and then summing all the first proportions to obtain the second proportion, which consumes more system computing resources, and it is more necessary to calculate the second proportion in the case that there are more noise source types and the system computing resources are sufficient. Therefore, before calculating the second proportion, the embodiment first determines the number of noise source types in the noise source classification result corresponding to all audio slices of the noise audio, and determines the amount of available system resources of the noise recognition system, and when the number of noise source types is greater than the preset number threshold, and the amount of available system resources is greater than the preset resource amount threshold, the calculation of the second proportion is started.
[0082] Step S50, noise event recognition based on the noise source classification result corresponding to the main energy type.
[0083] Specifically, in some embodiments, the noise event recognition method based on the main sound source detection of the embodiment is applied to a noise recognition system, the noise recognition system includes a noise sensor, the noise recognition system is associated with a target account, Figure 1 Step S50 includes but is not limited to the following steps:
[0084] Step S51, obtaining meteorological data, sound source positioning data, physical position of the noise sensor, over-standard classification threshold, and sound source investigation situation information in the time period in which the noise audio is located;
[0085] Step S52, obtaining the minute-level leq monitoring data corresponding to the noise audio;
[0086] Step S53, based on the preset AI model, the data integration of the slice recognition result corresponding to the noise audio, the main energy type, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the exceeding classification threshold and the sound source investigation information, to obtain the noise event recognition data unit;
[0087] Step S54, sending the noise event recognition data unit to the target account, and obtaining the target noise event recognition result for the noise audio based on the feedback data of the target account.
[0088] Specifically, in some embodiments, the noise event recognition method based on the main sound source detection of the present embodiment is applied to a noise recognition system, the noise recognition system comprising a noise sensor, the noise recognition system being associated with a target account, Figure 1 Step S50 includes but is not limited to the following steps:
[0089] Step S55, obtaining the meteorological data, the sound source positioning data, the physical position of the noise sensor, the exceeding classification threshold and the sound source investigation information in the time period in which the noise audio is located;
[0090] Step S56, obtaining the minute-level leq monitoring data corresponding to the noise audio;
[0091] Step S57, based on the preset AI model, the data integration of the slice recognition result corresponding to the noise audio, the main energy type, the second proportion, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the exceeding classification threshold, the sound source investigation information and the minute-level leq monitoring data, to obtain the noise event recognition data unit;
[0092] Step S58, sending the noise event recognition data unit to the target account, and obtaining the target noise event recognition result for the noise audio based on the feedback data of the target account.
[0093] Specifically, the target account of the present embodiment corresponds to a system administrator, which can make a final decision on the noise event recognition result based on the content in the noise event recognition data unit.
[0094] Specifically, the physical position of the noise sensor of the present embodiment refers to the physical position of the detection instrument for obtaining the noise audio, and the exceeding classification threshold refers to the exceeding limit value of the functional area where the noise sensor is located or the exceeding plant boundary emission limit value.
[0095] Specifically, the output style of the noise event identification data unit in this embodiment is determined according to actual needs. The output style of the noise event identification data unit in this embodiment is as follows: {time information of noise audio, physical location of noise sensor, slice identification result, main energy type, meteorological data, sound source localization data, minute-level LEQ monitoring data, exceedance classification threshold, sound source investigation information...}, or, {time information of noise audio, physical location of noise sensor, slice identification result, main energy type, second proportion, meteorological data, sound source localization data, minute-level LEQ monitoring data, exceedance classification threshold, sound source investigation information...}.
[0096] Understandably, after determining the main energy type corresponding to the noise audio, this embodiment acquires meteorological data, sound source localization data, the physical location of the noise sensor, the exceedance threshold, and sound source investigation information for the time period corresponding to the noise audio, as well as minute-level LEQ monitoring data corresponding to the noise audio. A preset AI model integrates the identification results of each slice of the noise audio, the main energy type, meteorological data, sound source localization data, the physical location of the noise sensor, the exceedance threshold, the sound source investigation information, and the minute-level LEQ monitoring data. Alternatively, a second proportion may be added as needed for data integration to obtain a noise event identification data unit. In addition to the main energy type, the noise event identification data unit also contains meteorological data, sound source localization, and physical location of the detection point for the time period corresponding to the noise audio. This data can be used by the administrator of the target account to make an effective and accurate judgment on whether the main energy type is correct, whether it is affected by environmental or meteorological factors, or how to implement subsequent noise reduction measures.
[0097] like Figure 2 As shown, Figure 2 This is a structural diagram of a control device provided in one embodiment of this application. The present invention also provides a control device 200, comprising:
[0098] The processor 210 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0099] The memory 220 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 220 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 220 and are called and executed by the processor 210 to implement the noise event recognition method based on main sound source detection according to the embodiments of the present application;
[0100] The input / output interface 230 is configured to realize information input and output.
[0101] The communication interface 240 is configured to realize the communication interaction between the device and other devices, and can realize the communication through a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).
[0102] The bus 250 is configured to transmit information between various components (for example, the processor 210, the memory 220, the input / output interface 230, and the communication interface 240) of the device.
[0103] The processor 210, the memory 220, the input / output interface 230, and the communication interface 240 are connected to each other through the bus 250 to realize the communication connection between the device.
[0104] In addition, the embodiments of the present application also provide an electronic device including the control device 200 of the above embodiments.
[0105] In addition, the embodiments of the present application also provide a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program. When the computer program is executed by a processor, the noise event recognition method based on main sound source detection is realized.
[0106] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, which can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The above-described device embodiments are only illustrative, and units described as separate components can or can not be physically separated, implemented in one place, or distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.
[0107] Those of ordinary skill in the art can understand that all or some steps in the above disclosed method and system can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as known to those of ordinary skill in the art, communication media typically includes computer readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery medium.
[0108] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the above-described embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application. These equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A method of noise event identification based on dominant source detection, characterized in that, The method comprises the following steps: obtaining a noise audio through a noise sensor, continuously slicing the noise audio to obtain a plurality of audio slices with the same time length; extracting a time-frequency domain voiceprint feature of each audio slice, inputting each time-frequency domain voiceprint feature into a preset deep learning model for noise source classification and recognition to obtain a slice recognition result, the slice recognition result comprising a noise source classification result corresponding to the audio slice and a confidence degree of the noise source classification result; calculating an energy value of each audio slice, and associating each energy value with the noise source classification result and the confidence degree of the corresponding audio slice, wherein the energy value is a time domain energy value or a frequency domain energy value of the audio slice; based on each energy value, the noise source classification result and the confidence degree corresponding thereto, and a preset main energy detection algorithm, calculating a main energy type corresponding to the noise audio, wherein the main energy type is used to indicate a noise source type with the largest proportion in the noise audio; based on the noise source classification result corresponding to the main energy type, performing noise event recognition; based on each energy value, the noise source classification result and the confidence degree corresponding thereto, and a preset main energy detection algorithm, calculating a main energy type corresponding to the noise audio, comprising: calculating a leq value of each audio slice based on each energy value; modifying the noise source type corresponding to the noise source classification result corresponding to the confidence degree less than a confidence degree threshold to a reference noise source type, adding the leq values corresponding to the noise source classification results of the same type to obtain a plurality of first reference leq values, and determining the noise source type corresponding to the noise source classification result corresponding to the largest first reference leq value as the main energy type; or, performing multiplication weighting processing on each leq value and the confidence degree corresponding to the audio slice to obtain a plurality of second reference leq values, and determining the noise source type corresponding to the noise source classification result corresponding to the largest second reference leq value as the main energy type.
2. The method of claim 1, wherein, based on each energy value, the noise source classification result and the confidence degree corresponding thereto, and a preset main energy detection algorithm, calculating a main energy type corresponding to the noise audio, comprising: modifying the noise source type corresponding to the noise source classification result corresponding to the confidence degree less than a confidence degree threshold to a reference noise source type; adding the energy values corresponding to the noise source classification results of the same type to obtain a plurality of first reference energy values; determining the noise source type corresponding to the noise source classification result corresponding to the largest first reference energy value as the main energy type.
3. The method of claim 1, wherein, based on each energy value, the noise source classification result and the confidence degree corresponding thereto, and a preset main energy detection algorithm, calculating a main energy type corresponding to the noise audio, comprising: performing multiplication weighting processing on each energy value and the confidence degree corresponding thereto to obtain a plurality of second reference energy values; determining the noise source type corresponding to the noise source classification result corresponding to the largest second reference energy value as the main energy type.
4. The method of claim 1, wherein, The method is applied to a noise recognition system, and after a main energy type corresponding to the noise audio is calculated based on each energy value, a corresponding noise source classification result, the confidence, and a preset main energy detection algorithm, the method further comprises: determining a number of noise source types in the noise source classification result corresponding to all of the audio slices of the noise audio; determining an amount of available system resources of the noise recognition system at present; when the number of noise source types is greater than a preset number threshold, and the amount of available system resources is greater than a preset resource amount threshold, calculating an energy proportion of a main energy sound source corresponding to the main energy type in each of the audio slices to obtain a plurality of first proportions, and summing all of the first proportions to obtain a second proportion, wherein the second proportion is used to indicate an energy proportion of the main energy sound source corresponding to the main energy type in the noise audio.
5. The method of claim 1, wherein, The noise recognition system comprises the noise sensor, and the noise recognition system is associated with a target account. Noise event recognition is performed based on the noise source classification result corresponding to the main energy type, comprising: obtaining meteorological data, sound source positioning data, a physical position of the noise sensor, an over-standard classification threshold, and sound source investigation situation information in a time period in which the noise audio is located; obtaining minute-level leq monitoring data corresponding to the noise audio; integrating each slice recognition result, the main energy type, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data corresponding to the noise audio based on a preset AI model to obtain a noise event recognition data unit; sending the noise event recognition data unit to the target account, and obtaining a target noise event recognition result for the noise audio based on feedback data of the target account.
6. The noise event recognition method based on primary sound source detection according to claim 4, characterized in that, The noise recognition system comprises the noise sensor, and the noise recognition system is associated with a target account. Noise event recognition is performed based on the noise source classification result corresponding to the main energy type, comprising: obtaining meteorological data, sound source positioning data, a physical position of the noise sensor, an over-standard classification threshold, and sound source investigation situation information in a time period in which the noise audio is located; obtaining minute-level leq monitoring data corresponding to the noise audio; integrating each slice recognition result, the main energy type, the second proportion, the meteorological data, the sound source positioning data, the physical position of the noise sensor, the over-standard classification threshold, the sound source investigation situation information, and the minute-level leq monitoring data corresponding to the noise audio based on a preset AI model to obtain a noise event recognition data unit; sending the noise event recognition data unit to the target account, and obtaining a target noise event recognition result for the noise audio based on feedback data of the target account.
7. A control device characterized by comprising: comprising at least one control processor and a memory communicatively connected to the at least one control processor; the memory storing instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the method of identifying a noise event based on primary sound source detection according to any one of claims 1 to 6.
8. An electronic device, comprising: comprising the control device of claim 7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions for causing a computer to perform the method of identifying a noise event based on primary sound source detection according to any one of claims 1 to 6.
Citation Information
Patent Citations
Speech signal processing method, apparatus and device, and storage medium
WO2022134833A1
Training method for speech enhancement network, speech enhancement method, and electronic device
WO2025035975A1