Urban environment noise source automatic identification method, medium and system

By building a parallel CBAM-DCRNN model, combining CNN, CBAM and GRU networks, the efficiency and accuracy of online automatic identification of urban environmental noise sources in the existing technology are solved, and efficient identification and real-time supervision of 20 noise sources are achieved.

CN120108418APending Publication Date: 2025-06-06SHANGHAI ACADEMY OF ENVIRONMENTAL SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510112204.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient and accurate automatic identification of urban environmental noise sources online, and there is a lack of intelligent identification methods for specific environmental noise sources.

Method used

An automatic noise source recognition model for urban environment based on parallel CBAM-DCRNN was constructed. By integrating convolutional neural network (CNN), integrated convolutional block attention module (CBAM) and gated recurrent unit network (GRU), combined with transfer learning methods, an efficient and accurate noise source recognition model was developed.

Benefits of technology

It has achieved efficient identification of 20 representative noise sources in urban environments, with good generalization and scalability, and can be quickly deployed on noise online monitoring equipment at low cost, real-time online automatic identification and supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108418A_ABST
    Figure CN120108418A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of environmental noise control, and particularly relates to an urban environmental noise source automatic identification method, medium and system, and the method comprises the following steps: S1, constructing a data set; s2, a convolutional neural network is adopted to integrate a convolutional block attention module in parallel, and a parallel CBAM-DCRNN model is constructed in combination with a gated cycle unit network; s3, extracting time-frequency features of the audio samples in the data set constructed in the step S1 by adopting Mel-scale short-time Fourier transform, and inputting the time-frequency features into the model constructed in the step S2 for training; and S4, inputting real city environment sound into the model trained in the step S3 to carry out noise source identification. Compared with the prior art, the method solves the problem that an intelligent recognition method for specific environmental noise sources such as urban environmental noise sources is lacked in the prior art. According to the scheme, the urban environment noise source automatic identification model based on the parallel CBAM-DCRNN is constructed, and the noise source in the urban environment can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of environmental noise control, and in particular relates to an automatic identification method, medium and system for urban environmental noise sources. Background Art

[0002] With the rapid development of urbanization, in recent years, the scale of cities in my country has continued to expand and the population density has continued to rise. The dense traffic network, numerous commercial buildings, frequent urban construction, etc. have generated various environmental noise sources. These noise sources are of various types and change frequently. Based on this, noise pollution has become a serious environmental problem that affects human health and quality of life.

[0003] In order to strengthen the precision and efficiency of noise pollution prevention and control, it is crucial to obtain accurate information on noise sources in a timely manner. However, existing online noise monitoring technologies generally still remain at the level of monitoring physical indicators such as sound level intensity. Regulators usually need to use offline methods such as on-site surveys, listening back to recordings, or analyzing spectrograms to determine the noise source of a certain historical period. Obviously, these noise source tracing methods are time-consuming and labor-intensive, with high labor costs, and are seriously lacking in timeliness, making them even more difficult to use for long-term noise source monitoring.

[0004] In 2023, the Ministry of Environmental Protection's "Noise Pollution Prevention and Control Action Plan" clearly proposed to encourage cities with conditions to rely on emerging means such as noise source tracing to strengthen the precise control of noise pollution prevention and control. Existing research related to intelligent identification of environmental noise was investigated, such as: CN115662464A, Guangzhou Yunjing Information Technology Co., Ltd. proposed "a method and system for intelligent identification of environmental noise", CN118262746A, Beijing Institute of Ecological and Environmental Protection Science proposed "intelligent noise identification method", these two inventions mainly proposed technical methods related to environmental noise signal feature processing and analysis; CN115954019A, Guangzhou Shengbo Acoustic Technology Co., Ltd. proposed "a method and system for environmental noise identification that integrates self-attention and convolution operations", this invention mainly proposed a concept of an environmental noise identification model that integrates self-attention and convolution operations, but there is no real reliance on a specific environmental noise data set to build a model. In fact, data sets with different characteristics have very different requirements for the model's network architecture, weights, and training hyperparameters.

[0005] According to the comprehensive survey results, there is no relevant achievement in developing intelligent identification models for specific urban environmental noise source data sets. Therefore, in the context of big data and intelligence, an efficient and accurate online automatic identification technology for urban environmental noise sources is a top priority. Summary of the invention

[0006] The purpose of the present invention is to provide a method, medium and system for automatically identifying urban environmental noise sources in order to solve at least one of the above problems, so as to solve the lack of intelligent identification methods for specific environmental noise sources, such as urban environmental noise sources, in the prior art. This solution constructs an efficient and high-accuracy automatic identification model of urban environmental noise sources based on parallel CBAM-DCRNN, which can accurately identify noise sources in urban environments.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] The first aspect of the present invention discloses a method for automatically identifying urban environmental noise sources, comprising the following steps:

[0009] S1: Build a dataset including representative large-scale urban environmental noise sources;

[0010] S2: A parallel CBAM-DCRNN model is constructed by using a convolutional neural network (CNN) in parallel with a convolutional block attention module (CBAM) and a gated recurrent unit network (GRU).

[0011] S3: using Mel-scale short-time Fourier transform (Mel-STFT) to extract the time-frequency features of the audio samples in the data set constructed in step S1, and inputting the extracted time-frequency feature map into the model constructed in step S2 for training;

[0012] S4: Input the real urban environment sound into the model trained in step S3 to identify the noise source.

[0013] Preferably, in step S1, the representative large-scale urban environmental noise sources include: traffic noise, HVAC equipment operation noise, mechanical operation noise, human activity noise, animal sounds and natural sounds.

[0014] Preferably, in step S1, the audio samples in the data set are obtained by resampling the original audio after time-length normalization.

[0015] Preferably, in step S2, the convolutional neural network parallel integrated convolutional block attention module is obtained by integrating a convolutional block attention module in parallel on each convolution (Conv) layer of the convolutional neural network.

[0016] Preferably, in step S2, the gated recurrent unit network is connected in series between the pooling layer and the fully connected layer of the convolutional neural network.

[0017] Preferably, in step S3, the Mel-scale short-time Fourier transform is: converting the audio sample into time-frequency domain by short-time Fourier transform, and then passing through a Mel filter bank to obtain a Mel-STFT time-frequency feature map.

[0018] Preferably, in step S3, the training is performed by transfer learning.

[0019] Preferably, the transfer learning is as follows: 1) randomly collecting n categories from the extracted time-frequency feature graph to construct a feature subset, and obtaining an initial weight by pre-training a model with the feature subset; 2) based on the initial weight, obtaining a final weight by training a model with all the extracted time-frequency feature graphs;

[0020] Where n is a positive integer and does not exceed the number of categories in the dataset.

[0021] A second aspect of the present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the automatic identification methods described above.

[0022] The third aspect of the present invention discloses an automatic identification system for urban environmental noise sources, the system comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements any of the automatic identification methods described above when executing the computer program.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1) Problem-oriented, a large-scale representative urban environmental noise source data set was constructed, and based on this data set, an intelligent identification method for urban environmental noise sources was developed: The present invention integrates existing environmental noise audio and supplements field collection to construct a data set covering 20 representative urban environmental noise sources, which basically covers all characteristic frequency bands of environmental noise. Furthermore, based on this data set and integrated intelligent algorithms, an efficient and automatic identification method for urban environmental noise sources was developed. Different from the existing technology that remains at the conceptual research of identification methods, the present invention is problem-oriented and is based on solving the problem of identifying real urban environmental noise sources. The proposed method can efficiently identify at least 20 representative urban environmental noise sources.

[0025] 2) Developed an intelligent identification model of urban environmental noise sources based on parallel CBAM-DCRNN: The present invention integrates CBAM with the traditional CNN algorithm, designs an integration strategy of parallel connection of CBAM module and Conv block, and then develops a parallel CBAM-CNN integrated network. On this basis, the CBAM-CNN network is combined with the GRU network, the network architecture of the deep model is designed, and finally a parallel CBAM-DCRNN deep intelligent identification model is developed. Therefore, the parallel CBAM-DCRNN model developed by the present invention integrates the powerful ability of the convolutional network to extract high-dimensional spectrum features and the powerful ability of the recurrent network to learn feature time series information. On this basis, the application of the CBAM module increases the ability of the network to focus on important features and suppress unnecessary features. Moreover, the deep network structure further enhances the model's ability to learn data features. In addition, the GRU and CBAM used are both lightweight structures, which achieve maximum savings in computing costs while improving model accuracy. Therefore, the parallel CBAM-DCRNN model proposed in the present invention can handle complex environmental noise source identification problems with high performance.

[0026] 3) Promising generalization and scalability: The data set constructed by the present invention covers various traffic noises such as roads, subways, airplanes, etc., which are mainly in the medium and low frequency bands, and the operating noises of HVAC equipment such as air conditioners and heat pumps, various high-frequency bird sounds, insect sounds, and various frequency bands of noises generated by other human activities. Moreover, the collected audios are all from real environments, and they are generally superimposed with various background sounds such as traffic and human activities. Therefore, the data set constructed by the present invention basically covers the characteristics of each frequency band of environmental noise. Therefore, the recognition model developed based on this data set has excellent generalization ability. It is not limited to the 20 types of noise sources of the present invention, but can also be expanded to more environmental noise source categories and scenarios. Furthermore, the present invention introduces a transfer learning method to optimize the model training process, which not only greatly improves the training efficiency and reduces the computational cost, but also further improves the scalability of the model.

[0027] 4) The automatic identification application developed based on the automatic identification method can be quickly deployed at low cost on any online noise monitoring device to achieve real-time online automatic identification and supervision of urban environmental noise sources: The present invention has developed an automatic identification application based on the parallel CBAM-DCRNN model, which is an independently executable program that can be efficiently deployed or ported to any online sound level monitoring device at low cost without the need to configure additional hardware and software. Therefore, the invention provides a feasible solution for online automatic supervision, source tracing and early warning of regional environmental noise in the city, enabling noise management departments to take precise control measures in a timely manner, and noise law enforcement departments to grasp the basis for law enforcement in real time and investigate and deal with noise violations in a timely manner. Furthermore, long-term online noise monitoring and identification data can be used to evaluate the sound environment in the area, providing a data basis for accurate and long-term management of the regional sound environment.

[0028] In summary, the present invention provides technical support for efficient and accurate control and management of urban environmental noise pollution. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A flow chart of the method for automatically identifying urban environmental noise sources;

[0030] Figure 2 The architecture diagram of the parallel CBAM-DCRNN intelligent recognition model;

[0031] Figure 3 This is the integrated network structure diagram of the CBAM-Conv layer;

[0032] Figure 4 This is the operation flow chart of the automatic identification method of urban environmental noise sources based on the CBAM-DCRNN model;

[0033] Figure 5(1) and Figure 5(2) are the identification result diagrams of the automatic identification method of urban environmental noise sources in Application Examples 1 to 3. DETAILED DESCRIPTION

[0034] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should belong to the scope of protection of the present invention.

[0035] The following unfinished matters can all adopt existing technologies.

[0036] Example

[0037] like Figure 1-Figure 3As shown, the present invention develops a high-performance automatic identification method for urban environmental noise sources based on a parallel CBAM-DCRNN intelligent model, which is not limited to the identification of the 20 types of urban environmental noise sources proposed in this embodiment, and can be flexibly extended to more categories and scenarios.

[0038] The main implementation steps include: by integrating representative categories of urban environmental noise audio in existing public datasets and supplementary collection of on-site noise, a large dataset of representative large-scale urban environmental noise sources is constructed; by integrating CNN models, CBAM modules and GRU networks, a high-performance recognition model is designed, and based on this dataset, a parallel CBAM-DCRNN-based intelligent recognition model for urban environmental noise sources is trained; finally, by packaging the model architecture and its weights, an application with the CBAM-DCRNN model as the core is developed for automatic identification of urban environmental noise sources, which realizes efficient and intelligent recognition of each input noise audio.

[0039] like Figure 1 As shown, the present invention proposes a method for efficiently and automatically identifying urban environmental noise sources, comprising the following steps:

[0040] Step 1: Collect audio from representative urban environmental noise sources and build a large urban environmental noise source dataset;

[0041] Specifically include:

[0042] Step 11: Investigate the types and noise scenarios of environmental noise that are prevalent in the city and have a significant impact on residents, and select representative types of urban environmental noise sources to fully cover the characteristics of all frequency bands of urban environmental noise;

[0043] Step 12: By collecting urban environmental noise audio from various public environmental sound, scene sound and sound effect datasets, and further supplementing the data through on-site recording, an audio dataset covering 20 environmental noise sources was constructed, including road traffic, engine, subway station, airplane, vehicle horn, whistle, air conditioner, heat pump, drilling, crusher, children playing, street music, grocery store, restaurant, dog barking, cat meowing, bird singing, cricket singing, rain, thunderstorm, etc. These 20 noise sources belong to 6 major categories: transportation, HVAC equipment, mechanical operation, human activities, animals and nature; the constructed dataset basically covers the characteristics of all frequency bands of environmental noise;

[0044] Step 13: Due to the diverse sources of samples, the length and sampling rate of the audio vary. Therefore, based on the Python platform, all audio samples are re-standardized to 4s in length by means of truncation, time shift, splicing, and trimming silence, and resampled at a sampling rate of 22050Hz. The classID and class of each audio sample are re-labeled, and a meta file is created. Finally, a large urban environmental noise source dataset is constructed, which covers 20 noise source categories (including 6 major noise source categories) and a total of 13,654 labeled audio samples.

[0045] Step 2: Based on the deep learning algorithm, the CNN model is used to integrate the CBAM attention module in parallel, and combined with the GRU network, a high-performance integration strategy is designed to develop an intelligent identification model for urban environmental noise sources based on parallel CBAM-DCRNN;

[0046] Specifically include:

[0047] Step 21: Integrate the CBAM module into the CNN network and design the network structure of the CBAM-Conv integration to build a parallel CBAM-CNN network;

[0048] Step 22: Combine the CBAM-CNN network and the GRU network to design a joint deep network architecture to build a parallel CBAM-DCRNN deep intelligent recognition model.

[0049] Step 3: Based on the Python platform, using the urban environmental noise source dataset constructed in Step 1, the transfer learning method is introduced to train the CBAM-DCRNN deep model proposed in Step 2.

[0050] Specifically include:

[0051] Step 31: Use STFT time-frequency conversion combined with Mel filter bank to extract the time-frequency features of all audio in the data set, and construct the Mel-STFT time-frequency feature set of urban environmental noise source (Mel-STFT time-frequency representation of noise signal) as model input;

[0052] Step 32: In the large (complete) Mel-STFT time-frequency feature set, randomly collect n categories to construct a small Mel-STFT feature subset, where n is a positive integer and n does not exceed the total number of categories in the data set.

[0053] Step 33: Since the model has a deep network layer and a large number of parameters, a transfer learning method is used to optimize the model training process to improve the model training efficiency and save computing costs, thereby improving the flexibility of the model to be transplanted to a larger data set: First, a feature subset is used as input, with 80% of the training set and 20% of the validation set to pre-train the CBAM-DCRNN model weights. Then, the pre-trained weights are used as the initial weights again, and a large feature set is used as input. Similarly, 80% of the training set and 20% of the validation set are used to train the CBAM-DCRNN model again, and finally the optimal weights of the model are obtained.

[0054] Specifically,

[0055] The main method for extracting Mel-STFT time-frequency features is: first use STFT to convert the collected audio into time-frequency domain, that is, to obtain the STFT time-frequency features of the linear scale. On this basis, the Mel filter group is used to map the linear STFT spectrum diagram to the Mel nonlinear frequency scale, and then a Mel-STFT time-frequency feature diagram that is closer to the response characteristics of the human auditory system is obtained. This time-frequency representation can more effectively capture the important characteristics of non-stationary and dynamic environmental sound signals, and is very suitable for processing complex environmental noise signals.

[0056] The training hyperparameters are designed as follows: the Adam optimizer is used to train the deep network, and the learning rate, batch size, and training rounds are set to 10 respectively. -3 , 128 and 1024. The cross entropy loss function is used as the loss function. In order to avoid overfitting and save computing resources, the early stopping method is used for training, and its tolerance is set to 50.

[0057] Step 4: Apply the CBAM-DCRNN model to develop an automatic urban environmental noise source identification application to achieve portable, online, efficient and intelligent identification of each input noise sample.

[0058] Specifically include:

[0059] Step 41: Package the trained CBAM-DCRNN model architecture and its weights, as well as the Mel-STFT time-frequency feature extraction algorithm, into an independent executable application;

[0060] Step 42: For each or each batch of input noise audio, the program automatically calls the built-in Mel-STFT algorithm to extract the time-frequency feature map of the audio, and then calls the CBAM-DCRNN model to efficiently identify the noise source category of each input sample and automatically output the identification result.

[0061] like Figure 2-Figure 3As shown in the figure, it is the architecture diagram of the parallel CBAM-DCRNN intelligent recognition model designed by the present invention, and the integrated network structure diagram of the CBAM-Conv layer. The specific model architectures are designed as follows:

[0062] 1. The architecture of the parallel CBAM-DCRNN intelligent recognition model mainly includes 4 CBAM-Conv blocks, 4 pooling layers, 1 GRU block, and two fully connected layers. The specific network structure and parameter design are as follows:

[0063] Each CBAM-Conv block is designed as two CBAM-Conv layers connected together. Each CBAM-Conv layer uses convolution kernels to extract local features of the spectrogram image, and uses the CBAM enhancement network to focus on important features and suppress unnecessary features. Among them, CBAM-Conv1 and CBAM-Conv2 layers are designed with 32 convolution kernels of size 3*5, CBAM-Conv3 and CBAM-Conv4 layers are designed with 64 convolution kernels of size 3*1, CBAM-Conv5 and BAM-Conv6 are designed with 128 convolution kernels of size 1*5, and CBAM-Conv7 and CBAM-Conv8 layers are designed with 256 3*3 kernels. Each CBAM-Conv layer uses ReLU as the activation function. After each CBAM-Conv block, there is a maximum pooling layer to reduce the feature size in the frequency domain or time domain. After the four CBAM-Conv blocks, a GRU block consisting of two GRU units is connected. Each GRU unit uses the tanh activation function and a dropout rate of 0.5 is used to regularize the GRU network. Finally, the entire network is connected to two fully connected layers through a flattening layer. The flattening layer is used to flatten the features extracted by attention-convolution, pooling, loop and other operations into a one-dimensional structure, which is then sent to the last two fully connected layers. The first fully connected layer is designed with 128 neurons and uses the ReLU activation function. The last fully connected layer is the output layer, and its number of neurons is the same as the number of noise source categories covered in the constructed urban environmental noise dataset, that is, 20 neurons. The output layer uses the SoftMax function to output the probability of each sample being classified into the corresponding category.

[0064] 2. Each CBAM-Conv layer is an integration of a CBAM module and a 2D Conv layer in parallel. The structure of the integrated network is specifically designed as follows:

[0065] A CBAM attention module is designed on each 2D Conv layer of the CNN network to optimize the ability of the network layer to extract effective features. Its integration method is designed to be parallel, that is, a CBAM-Conv layer is constructed. For the input Mel-STFT time-frequency feature map, this layer integrates the spectral features extracted by the 2D Conv layer and the features extracted by the CBAM module, which together constitute the input of the next CBAM-Conv layer. Among them, CBAM is designed to be composed of two sub-modules, CAM and SAM, connected in series. First, CAM generates the input feature map required by SAM after a series of pooling, multi-layer perceptron, summation, multiplication and other operations on the input time-frequency feature map. Then, SAM applies pooling, convolution, multiplication and other operations to the feature map output by CAM along the channel axis to generate CBAM features.

[0066] The specific network structure and operation principle of the two sub-modules of CAM and SAM are as follows:

[0067] CAM module: First, the input H×W×C time-frequency feature map is subjected to global maximum pooling and global average pooling respectively to obtain two 1×1×C feature maps. Then, they are sent to a two-layer perceptron respectively. The number of neurons in the first layer is C / r (r is the reduction rate), the activation function is ReLU, and the number of neurons in the second layer is C. The two-layer neural network is shared. Then, the output features of the two-layer perceptron are combined using element-wise summation. After the sigmoid activation operation, the CAM feature, Mc, is generated. Finally, Mc and the original input feature map are element-wise multiplied to generate the input feature map required by SAM.

[0068] SAM module: The features output by CAM are used as the input feature maps of this module. First, global maximum pooling and global average pooling are applied along the channel axis to obtain two H×W×1 feature maps. Then, these two feature maps are concatenated. Next, after a 7×7 convolution operation, the feature map is reduced to 1 channel, i.e., H×W×1. Then, a sigmoid operation is performed to generate the SAM feature, M_s. Finally, the feature M_s is multiplied by the input feature of this module to generate the CBAM feature.

[0069] like Figure 4 As shown, the operation process of the urban environmental noise source automatic identification application based on the parallel CBAM-DCRNN model of the present invention is as follows:

[0070] 1) Run the program to read the noise audio to be identified;

[0071] 2) Automatically call the built-in Mel-STFT algorithm to extract the time-frequency feature map of the input audio;

[0072] 3) Automatically call the built-in CBAM-DCRNN intelligent recognition model to efficiently identify the noise source category of each input sample;

[0073] 4) Output and save the recognition results.

[0074] Application Example 1

[0075] In concentrated residential areas along urban traffic routes, residents are generally affected by road traffic noise, engine noise, horn noise, operating noise of air conditioners and heat pumps, entertainment music and instrument noise, children playing noise, dogs and cats meowing noise, electrical noise generated by decoration, and mechanical gravel noise generated by construction.

[0076] The urban environmental noise source automatic identification application based on the parallel CBAM-DCRNN model constructed in the embodiment of the present invention is deployed on any noise online automatic monitoring equipment, and applied to the online automatic supervision and early warning of environmental noise sources in noise-sensitive areas. By setting a sound level threshold for the noise online monitoring system (generally, the environmental noise limit value published in the "Sound Environment Quality Standard" can be used), when the sound level is monitored to exceed the set threshold, the automatic identification application deployed in the background is triggered to automatically identify the current main noise source in real time, and save the identification results in real time. The noise management department takes precise control measures in a timely manner based on this online monitoring information; the noise law enforcement department uses this as the basis for law enforcement, traces the source of noise in real time, and investigates and punishes noise violations in a timely manner; the long-term identification data accumulated by the online automatic monitoring system is used to evaluate the sound environment in the area, to clarify the key noise impact sources in the area, and to provide a data basis for the long-term governance of the regional sound environment.

[0077] Application Example 2

[0078] In urban areas with better ecology, the sound level in the area often exceeds the environmental noise limit due to bird and insect chirping.

[0079] Existing online monitoring equipment that only focuses on sound levels often causes noise management departments to misjudge the quality of the regional acoustic environment. Therefore, the urban environmental noise source automatic identification application developed by the present invention is deployed in the noise online automatic monitoring equipment, which can distinguish natural sounds such as bird calls, cricket calls, rain, thunderstorms, etc. in real time, so that the management department can more truly understand the acoustic environment of the area, and then formulate an acoustic environment supervision plan that meets the characteristics of the region, formulate more refined acoustic environment assessment standards, and implement more targeted noise supervision measures.

[0080] Application Example 3

[0081] For residential areas near aircraft routes, the urban environmental noise source automatic identification application of the present invention is deployed in online automatic monitoring equipment, which can be used to automatically distinguish aircraft noise and other noises in real time, providing a basis for the corresponding management departments to timely allocate and implement supervision responsibilities.

[0082] As shown in Figure 5 (1) and Figure 5 (2), the automatic recognition application of the present invention corresponds to the recognition results of the 15 types of environmental noise sources involved in use cases 1-3. Three noise audios were randomly collected for each category, and a total of 45 recognition results were obtained. It can be seen from the test results that all the recognitions are accurate, indicating that the recognition accuracy of the automatic recognition method is high.

[0083] In summary, the present invention focuses on the task of urban environmental noise source identification. First, a large dataset covering 20 representative urban environmental noise sources is constructed. Based on this, a deep intelligent recognition model integrating convolutional neural network (CNN), lightweight convolutional block attention module (CBAM), and gated recurrent unit (GRU) is proposed, and a parallel integration strategy and a deep network architecture are designed. Then, the constructed dataset is used to train the model, and the transfer learning method is introduced to optimize the model training process. Finally, an efficient and high-accuracy urban environmental noise source automatic identification model based on parallel CBAM-DCRNN is constructed.

[0084] The above description of the embodiments is to facilitate the understanding and use of the invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the present invention should be within the scope of protection of the present invention.

Claims

1. A method for automatically identifying urban environmental noise sources, characterized in that: The steps include: S1: Build a dataset including representative large-scale urban environmental noise sources; S2: A parallel CBAM-DCRNN model is constructed by integrating the convolutional block attention module with the convolutional neural network in parallel and combining it with the gated recurrent unit network. S3: using Mel-scale short-time Fourier transform to extract the time-frequency features of the audio samples in the data set constructed in step S1, and inputting the extracted time-frequency feature map into the model constructed in step S2 for training; S4: Input the real urban environment sound into the model trained in step S3 to identify the noise source.

2. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S1, the representative large-scale urban environmental noise sources include: traffic noise, HVAC equipment operation noise, mechanical operation noise, human activity noise, animal noise and natural sound.

3. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S1, the audio samples in the data set are obtained by resampling the original audio after time-length normalization.

4. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S2, the convolutional neural network parallel integrated convolutional block attention module is obtained by integrating a convolutional block attention module in parallel on each convolutional layer of the convolutional neural network.

5. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S2, the gated recurrent unit network is connected in series between the pooling layer and the fully connected layer of the convolutional neural network.

6. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S3, the Mel-scale short-time Fourier transform is: converting the audio sample into the time-frequency domain through the short-time Fourier transform, and then passing through the Mel filter bank to obtain the Mel-STFT time-frequency feature map.

7. The method for automatically identifying urban environmental noise sources according to claim 1, characterized in that: In step S3, the training is performed by transfer learning.

8. The method for automatically identifying urban environmental noise sources according to claim 7, characterized in that: The transfer learning is as follows: 1) randomly collecting n categories from the extracted time-frequency feature graph to construct a feature subset, and obtaining an initial weight through a feature subset pre-training model; 2) Based on the initial weights, the model is trained by extracting all time-frequency feature maps to obtain the final weights; Where n is a positive integer and does not exceed the number of categories in the dataset.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the automatic identification method as claimed in any one of claims 1 to 8.

10. An automatic identification system for urban environmental noise sources, characterized in that: The system comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the automatic identification method as claimed in any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Method and system for intelligently identifying environmental noise

    CN115662464A

  • Environmental noise identification method and system fusing self-attention and convolution operation

    CN115954019A

  • Intelligent noise identification method

    CN118262746A