Environmental sound classification method and device, computer equipment and storage medium
By using preset classification models in the ambient sound classification method, combined with the generation of synthetic sound samples generated by the adversarial network and the diffusion model network, the problems of low data processing efficiency and insufficient classification accuracy in the prior art are solved, and more efficient and accurate environmental sound classification is achieved.
Patent Information
- Application Number
- CN202510276854.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-03
AI Technical Summary
Among the existing ambient sound classification methods, data processing efficiency and classification accuracy are insufficient, especially in noisy environments.
By obtaining real-time ambient sound data for preprocessing, the sound samples to be analyzed are obtained, and data analysis is performed using preset classification models. This preset classification model is trained from the initial sound sample and the synthetic sound sample. The synthetic sound sample is generated by generating adversarial networks and diffusion model networks, reducing the dependence on field recording and expert labeling data.
It improves data processing efficiency and sample quality of model training, enhances the generalization ability and robustness of the model, and improves the accuracy and stability of environmental sound classification.
Smart Images

Figure CN120089131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, computer device and storage medium for classifying environmental sounds in the fields of fintech and medical and health. Background Art
[0002] In the fields of fintech and medical and health, the detection and classification of environmental sounds play an important role. On the one hand, in the financial insurance industry, financial institutions can monitor the noise level of the surrounding environment through environmental sound classification technology to identify potential safety hazards, such as construction noise, crowd gathering, etc. On the other hand, in the wards of the medical and health industry, environmental sound classification technology can not only be used to monitor the breathing sounds and heart sounds of patients in real time, but also be used to automatically identify the alarm sounds that may be emitted by medical devices such as ventilators and monitors during operation, and notify medical staff in time for diagnosis and treatment.
[0003] In the prior art, common environmental sound classification methods include traditional passive acoustic monitoring methods and deep learning-based acoustic classification methods, and these methods have obvious deficiencies and defects. On the one hand, the passive acoustic monitoring method relies on a large number of on-site recording devices for data collection and requires manual regular maintenance, which limits the scale and efficiency of monitoring and classification. On the other hand, deep learning methods require a large amount of labeled data for model training. The acquisition of these data is time-consuming and laborious, and at the same time, the applicable range of the trained sound classification model is limited, and it often performs poorly in noisy environments. For example, the background noise in a wind farm will seriously affect the accuracy of sound classification.
[0004] Therefore, it is urgent to solve the problems of low data processing efficiency and insufficient classification accuracy in environmental sound classification. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, computer device and storage medium for classifying environmental sounds to solve the problems of low data processing efficiency and poor classification accuracy in environmental sound classification for the above technical problems.
[0006] An environmental sound classification method includes: Obtain real-time environmental sound data, preprocess the real-time environmental sound data to obtain a sound sample to be analyzed; Perform data analysis and processing on the sound sample to be analyzed through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthetic sound sample, and the synthetic sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0007] An environmental sound classification device, comprising: A sound preprocessing module, configured to obtain real-time environmental sound data, preprocess the real-time environmental sound data, and obtain a sound sample to be analyzed; A sound classification module, configured to perform data analysis and processing on the sound sample to be analyzed through a preset classification model, and obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthetic sound sample, the synthetic sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0008] A computer device, comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein when the processor executes the computer-readable instructions, the above-mentioned environmental sound classification method is implemented.
[0009] A computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the environmental sound classification method as described above.
[0010] In the above-mentioned environmental sound classification method, device, computer device, and storage medium, the environmental sound classification method obtains real-time environmental sound data, preprocesses the real-time environmental sound data, and obtains a sound sample to be analyzed; performs data analysis and processing on the sound sample to be analyzed through a preset classification model, and obtains a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthetic sound sample, the synthetic sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network. The present invention obtains a preset synthesis model based on a generative adversarial network and a diffusion model network, performs data synthesis processing on an initial sound sample through the preset synthesis model to synthesize high-quality environmental sound data as a synthetic sound sample, fully considers a complex noise environment, thereby reducing the dependence on field recording and expert-annotated data, and helping to improve the data processing efficiency and the sample quality of model training. The present invention trains a preset classification model based on an initial sound sample and a synthetic sound sample, improves the generalization ability and robustness of the model, and at the same time uses the preset classification model to perform data analysis and processing on the sound sample to be analyzed, further improving the accuracy and stability of environmental sound classification. Description of the Drawings
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 is a schematic diagram of an application environment of an environmental sound classification method in an embodiment of the present invention; Figure 2 is a schematic flowchart of an environmental sound classification method in an embodiment of the present invention; Figure 3 is a schematic structural diagram of an environmental sound classification device in an embodiment of the present invention; Figure 4 is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0014] The environmental sound classification method provided in this embodiment can be applied to an application environment such as Figure 1 where the client is communicatively connected to the server. Among them, the client has a sound collection function, including but not limited to various personal computers, laptop computers, smart phones, pickups, and recording devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The client sends the collected real-time environmental sound data to the server, requests the server to classify the real-time environmental sound data, and triggers the operation process of the environmental sound classification method. The server obtains the real-time environmental sound data, preprocesses the real-time environmental sound data, and performs data analysis and processing on the to-be-analyzed sound samples obtained by preprocessing through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data.
[0015] The environmental sound classification method of this embodiment can be applied to the property insurance business in the field of fintech. Through the data interaction between the client and the server, the real-time environmental sound data of the property insurance object (such as wind farm equipment) is monitored and classified, so as to accurately identify the risk weather or human damage sounds that may damage the equipment in a complex outdoor environment, and can more effectively evaluate and manage the environmental-related risk factors, providing more accurate and suitable property insurance services for customers.
[0016] The environmental sound classification method of this embodiment can also be applied to the medical service monitoring business in the field of medical and health. Through the data interaction between the client and the server, the real-time environmental sound data during the operation of medical devices (such as ventilators and monitors) is monitored and classified. In the noisy environment of the ward, the alarm sounds emitted by the devices are accurately identified, and medical staff are notified in time for diagnosis and treatment.
[0017] In one embodiment, as Figure 2 shown, an environmental sound classification method is provided, including the following steps S10 - S20: S10. Obtain real-time environmental sound data, preprocess the real-time environmental sound data to obtain a sound sample to be analyzed.
[0018] Understandably, the real-time environmental sound data refers to the natural sounds and background noises in the surrounding environment captured by the sound collection device in real time. The real-time environmental sound data is continuous and unprocessed raw audio signals. Therefore, it is necessary to go through the preprocessing steps to improve the accuracy of subsequent data analysis. The preprocessing process may include sound segmentation (dividing continuous audio into smaller segments for easy analysis), feature extraction (extracting key features helpful for classification from the audio, such as frequency, volume, duration, etc.), and format conversion (converting the linear spectrogram to the Mel spectrogram). The sound sample to be analyzed refers to the sound data to be classified that has been preprocessed and is easy to be understood and processed by the preset classification model.
[0019] S20. Perform data analysis and processing on the sound sample to be analyzed through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training the initial classification model with the initial sound sample and the synthetic sound sample, and the synthetic sound sample is obtained by performing data synthesis processing on the initial sound sample by a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0020] Understandably, the sound sample to be analyzed is input into a preset classification model for data analysis and processing, and the preset classification model outputs a sound classification result corresponding to the real-time environmental sound data. The preset classification model is a pre-trained neural network model that can predict different sound categories based on the characteristics of the sound sample. The sound classification result refers to the category label corresponding to the characteristics such as the source or use of the real-time environmental sound determined according to the classification rules. The sound classification result can include both the probability distribution of different category labels to which the sound sample to be analyzed belongs, or only the most likely category label, reflecting the characteristics of the real-time environmental sound data. The preset classification model can learn the classification rules from a large number of labeled sound samples and apply these classification rules to new samples. For example, the category labels include bird chirping, running water sound, thunderstorm sound, vehicle passing noise, conversation sound, machine running sound, etc.
[0021] The model training process of the preset classification model requires a large amount of training sample data, and these samples include initial sound samples and synthetic sound samples. The initial sound samples are the sound data in the real environment collected in the past period of time, reflecting the real characteristics of various environmental sounds. The synthetic sound samples are the sound samples that simulate the real environment obtained by performing data synthesis processing on the initial sound samples. The preset synthesis model is a pre-trained neural network model that can generate more diverse category sounds based on the initial sound samples. The synthetic sound samples can increase the diversity and quantity of the training sample data set, thereby improving the generalization ability of the preset classification model.
[0022] The preset synthesis model combines a Generative Adversarial Network (GAN) and a Diffusion Model network. The Generative Adversarial Network is a network composed of a generator and a discriminator, which learns the distribution of sound data through the mutual competition between the generator and the discriminator and generates new samples that are indistinguishable from the real data. For example, the generator receives a noise vector containing category information and outputs a fake spectrogram, while the discriminator tries to distinguish between real and fake spectrograms and their categories. The Diffusion Model network generates high-quality spectrograms as new sound samples by gradually removing noise. The Diffusion Model network is a fully convolutional neural network with a U-shaped structure, which can ensure the accuracy of segmentation while using fewer training pictures.
[0023] In this embodiment, by obtaining real-time environmental sound data, preprocessing the real-time environmental sound data to obtain a sound sample to be analyzed, performing data analysis and processing on the sound sample to be analyzed through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data. The preset classification model is obtained by training an initial classification model with an initial sound sample and a synthesized sound sample, and the synthesized sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model. This embodiment is based on a generative adversarial network and a diffusion model network to obtain a preset synthesis model, and synthesizes high-quality environmental sound data as a synthesized sound sample by performing data synthesis processing on the initial sound sample through the preset synthesis model, fully considering complex noise environments, thereby reducing the dependence on field recordings and expert-annotated data, and helping to improve data processing efficiency and the sample quality of model training. This embodiment trains a preset classification model based on the initial sound sample and the synthesized sound sample, improves the generalization ability and robustness of the model, and at the same time uses the preset classification model to perform data analysis and processing on the sound sample to be analyzed, further improving the accuracy and stability of environmental sound classification.
[0024] In one embodiment, in step S10, that is, preprocessing the real-time environmental sound data to obtain a sound sample to be analyzed, includes: S101. Perform Fourier transform processing on the real-time environmental sound data to obtain a linear spectrogram; S102. Convert the linear spectrogram into a Mel spectrogram; S103. Perform normalization processing on the Mel spectrogram to obtain a sound sample to be analyzed.
[0025] Understandably, in the process of preprocessing real-time environmental sound data, Fourier transform processing, spectrogram conversion processing, and normalization processing are required. The Fourier transform is a mathematical tool that can transform a signal from the time domain (time domain) to the frequency domain (frequency domain), thereby revealing the frequency components of the signal. The spectrogram conversion processing can be achieved through a Mel filter bank. The Mel filter bank is a set of triangular filter banks distributed according to the Mel scale, which is used to simulate the perception characteristics of the human ear for different frequencies. Specifically, first, the server performs operations such as digitization, pre-filtering, framing, and windowing on the real-time environmental sound data, making the characteristics of the sound signal more obvious, removing redundant data, and preparing for the subsequent Fourier transform processing step. Then, the fast Fourier transform is used to transform the time-domain signal into a frequency-domain signal to obtain a linear spectrogram, which shows the change of the intensity of the sound signal at different frequencies over time. Next, the linear spectrogram is smoothed through the Mel filter bank, and the effect of harmonics is eliminated to highlight the formants of the original sound. After the Mel filter bank processing of the sound signal, mathematical operations such as logarithmic operation and discrete cosine transform are also required to transform the frequency-domain signal into cepstral coefficients to obtain a Mel spectrogram. Finally, the Mel spectrogram is normalized to obtain the sound sample to be analyzed. The process of normalization processing is to normalize the pixel values so that the pixel value range of the Mel spectrogram is between 0 and 1, which is convenient for subsequent model analysis.
[0026] In this embodiment, the Fourier transform processing, spectrogram conversion processing, and normalization processing of the real-time environmental sound data help to extract important features in the sound signal, reduce noise and interference at the same time, make the sound sample to be analyzed more suitable for subsequent model analysis tasks, and improve the accuracy and robustness of the analysis.
[0027] In one embodiment, in step S20, that is, before the data analysis processing of the sound sample to be analyzed by the preset classification model, it includes: S201. Obtain historical environmental sound data, preprocess the historical environmental sound data to obtain an initial sound sample; S202. Perform data synthesis processing on the initial sound sample through a preset synthesis model to obtain a synthesized sound sample; S203. Obtain an initial classification model, and train the initial classification model based on the initial sound sample and the synthesized sound sample to obtain a preset classification model.
[0028] Understandably, before analyzing the sound sample to be analyzed using the preset classification model, it is necessary to first obtain the preset classification model through model training. The training sample data of the preset classification model includes the initial sound sample and the synthetic sound sample. First, obtain the historical environmental sound data, perform Fourier transform processing, spectrogram conversion processing, and normalization processing on the historical environmental sound data, and label the samples according to the category labels to obtain the initial sound sample. The historical environmental sound data refers to the environmental sound obtained within the preset acquisition period. The preset acquisition period refers to the past time period used for collecting environmental sound, for example, it can be the past month. Then, perform data synthesis processing on the initial sound sample through the preset synthesis model, and label the samples according to the category labels to obtain the synthetic sound sample. Finally, train the initial classification model based on the initial sound sample and the synthetic sound sample to obtain the preset classification model. The initial classification model is a pre-established basic classification model that has not been trained yet.
[0029] In one embodiment, in order to meet the environmental sound classification requirements of the wind farm in property insurance business, obtain the historical environmental sound data of the wind farm in the past month for preprocessing to obtain the initial sound sample. Perform data synthesis processing on the initial sound sample through the preset synthesis model to obtain the synthetic sound sample. Train the initial classification model based on the initial sound sample and the synthetic sound sample to obtain the preset classification model applicable to the environmental sound classification task of the wind farm.
[0030] In addition to obtaining the initial sound sample based on the real historical environmental sound data, this embodiment also generates high-quality synthetic sound samples through the preset synthesis model, reducing the cost and time of data collection and enhancing the diversity of the training sample dataset. Moreover, this embodiment trains the preset classification model based on the initial sound sample and the synthetic sound sample, which helps to improve the generalization ability and robustness of the model, making the preset classification model more reliable and effective in practical applications.
[0031] In one embodiment, in step S202, that is, performing data synthesis processing on the initial sound sample through the preset synthesis model to obtain the synthetic sound sample, includes: S2021. Perform spectrogram synthesis processing on the initial sound sample through the generative adversarial network to obtain the initial synthetic sample; S2022. Perform denoising processing on the initial synthetic sample through the diffusion model network to obtain the optimized synthetic sample; S2023. Perform post-processing on the optimized synthetic sample to obtain the synthetic sound sample.
[0032] Understandably, traditional sample data augmentation methods (such as rotation, flipping, etc.) are not applicable to the spectrograms of sound samples and may damage the time and frequency structures of audio. Therefore, in this embodiment, a preset synthesis model is used to generate spectrograms similar to real data to achieve the purpose of sample data augmentation.
[0033] Before performing data synthesis processing on the initial sound sample through the preset synthesis model, it is also necessary to train the preset synthesis model first. Specifically, environmental sound data of multiple categories are obtained and labeled with category labels to obtain a synthetic training sample data set corresponding to the preset synthesis model. The synthetic training sample data set is used to train the generative adversarial network and the diffusion model network. By optimizing the model parameters and introducing the noise guidance mechanism of the diffusion model into the generator, the synthetic sound samples generated by the generator are made closer to the distribution of real spectrograms. The training process involves the cooperation among the generator, discriminator, and diffusion model network. First, the generator of the generative adversarial network receives a random noise vector and generates a preliminary fake image. Then, the diffusion model network denoises the generated fake image to improve the image quality. Next, the discriminator of the generative adversarial network outputs the true and false probabilities. Finally, the discriminator loss is calculated and the parameters of the discriminator are updated. At the same time, the generator loss is calculated and the parameters of the generator and the denoising module are updated until the preset loss condition is met and the training is completed to obtain the preset synthesis model.
[0034] The server inputs the initial sound sample into the preset synthesis model. The generative adversarial network of the preset synthesis model performs spectrogram synthesis processing on the initial sound sample to obtain an initial synthetic sample. The initial synthetic sample refers to the environmental sound spectrogram containing noise synthesized by the generative adversarial network. The diffusion model network of the preset synthesis model denoises the initial synthetic sample to obtain an optimized synthetic sample. The optimized synthetic sample refers to the environmental sound spectrogram after removing the noise in the initial synthetic sample by the diffusion model network. In addition, post-processing such as cropping and normalization needs to be performed on the optimized synthetic sample, and sample labeling is performed according to the category labels to obtain synthetic sound samples to ensure the consistency and quality of the sample data.
[0035] In one embodiment, to meet the environmental sound classification requirements of the ward in medical monitoring services, historical environmental sound data of the hospital ward in the past month are obtained for preprocessing to obtain the initial sound sample. However, the initial sound sample only includes the ward environmental sounds under sunny weather conditions and is difficult to cover the ward environmental sounds under all weather conditions. The server inputs the initial sound sample into the preset synthesis model, and the preset synthesis model performs synthesis processing on the initial sound sample to obtain synthetic sound samples. The synthetic sound samples include the ward environmental sounds under rainy weather conditions, windy weather conditions, and thunderstorm weather conditions.
[0036] This embodiment utilizes the combination of a generative adversarial network and a diffusion model network to generate high-quality spectrograms with different class labels, thereby expanding the training dataset of a preset classification model using synthetic sound samples and improving the generalization ability of the model. The synthetic sound samples can be used as additional training data to enhance the training of the initial classification model, helping the model learn more diverse sound features, especially when the number of real sound samples is limited. At the same time, the synthetic sound samples can also simulate sound variations under different environmental conditions or sound scenarios, thereby improving the recognition ability and classification accuracy of the preset classification model for different environmental sounds.
[0037] In one embodiment, in step S203, that is, before obtaining the initial classification model, it includes: S2031. Determine a convolutional base network according to the first pre-trained network, the second pre-trained network, and the third pre-trained network; S2032. Construct a classification layer network based on a linear layer; S2033. Obtain an initial classification model according to the convolutional base network and the classification layer network.
[0038] Understandably, before obtaining the initial classification model, it is necessary to first construct the initial classification model. The initial classification model is a hybrid convolutional network integrated from multiple pre-trained networks, including an integrated convolutional base network and a classification layer network. A pre-trained network refers to a convolutional neural network that has been pre-trained on a large dataset for classification tasks. A convolutional neural network consists of two parts: a convolutional base and a classifier. The convolutional base of the model refers to a series of pooling layers and convolutional layers, and the classifier refers to the final densely connected classifier of the model. In this embodiment, the model reuse of the pre-trained network is achieved by means of model fine-tuning. The old convolutional base part in each pre-trained network is frozen and remains unchanged, the old classifier part in each pre-trained network is deleted, a new classifier part is added as a fully connected classifier, and each pre-trained network is jointly trained.
[0039] First, extract the respective convolutional base networks in the first pre-trained network, the second pre-trained network, and the third pre-trained network for freezing processing and weight initialization processing, and empty the classifier to obtain an integrated convolutional base network. In addition to the first pre-trained network, the second pre-trained network, and the third pre-trained network, other custom networks can also be added for integration as needed. Among them, the first pre-trained network can be a mobile convolutional network introducing the inverted residuals structure and the linear bottleneck layer, the second pre-trained network can be a residual convolutional network, and the third pre-trained network can be a visual geometry group convolutional network including multiple small-sized 3x3 convolutional kernels. Then, use a linear layer to connect the outputs of the integrated convolutional base network as a fully connected layer to obtain an integrated classification layer network. Finally, obtain an initial classification model based on the integrated convolutional base network and the classification layer network.
[0040] This embodiment is not limited to a single classification model, but obtains an initial classification model by integrating multiple different pre-trained classification models and linear layers, which can synthesize the advantages of each pre-trained model, reduce the deviation of a single model, and help improve the accuracy of the preset classification model.
[0041] In one embodiment, in step S203, that is, training the initial classification model based on the initial sound sample and the synthesized sound sample to obtain a preset classification model, includes: S2034. Determine the training sample data of the initial classification model according to the initial sound sample and the synthesized sound sample; S2035. Perform transfer learning training on the initial classification model through the training sample data to obtain an index training value; S2036. Determine whether the index training value reaches a preset index threshold; S2037. If the index training value has not reached the preset index threshold, adjust the parameters of the initial classification model, and perform transfer learning training on the optimized classification model obtained by parameter adjustment through the training sample data until the index training value of the optimized classification model reaches the preset index threshold, and determine the optimized classification model as the preset classification model.
[0042] Understandably, first, all the initial sound samples and synthesized sound samples are divided according to a preset division ratio to obtain a training set and a test set. The training sample data refers to the sound sample data in the training set, and each training sample data corresponds to a true class label of a sound sample. That is, the training sample data of the initial classification model is determined according to the initial sound samples and synthesized sound samples. The preset division ratio can be set and adjusted as needed, such as 5:5 or 8:2. Then, the training sample data is input into the initial classification model for transfer learning training to obtain the training classification results output by the initial classification model. The training classification results refer to the predicted class labels corresponding to the training sample data predicted by the initial classification model during the training process. Calculate the true class label and the predicted class label corresponding to the training sample data according to a preset index to obtain the index training value. The preset index refers to the standard used to evaluate the performance of the initial classification model after training, such as accuracy, precision, recall, etc. The index training value refers to the index value calculated after the latest model training. For example, when the index is accuracy, as the training progresses in each round, the iteration count gradually accumulates, and the index training value also changes continuously. Next, determine whether the index training value reaches the preset index threshold. The preset index threshold is the maximum critical value of the index preset for determining that the index of model training meets the requirement of stopping iteration. If the index training value reaches the preset index threshold and the training iteration count reaches the preset count threshold, it indicates that the model training is completed. The preset count threshold is the maximum critical value of the count preset for determining that the index of model training meets the requirement of stopping iteration. If the index training value has not reached the preset index threshold, adjust the parameters of the initial classification model (such as adjusting hyperparameters, optimizing the network structure, etc.), and perform transfer learning training on the optimized classification model obtained by parameter adjustment through the training sample data until the index training value of the optimized classification model reaches the preset index threshold and the training iteration count reaches the preset count threshold, and determine the optimized classification model as the preset classification model In this embodiment, the index training value is introduced during the transfer learning training process, and it is determined whether the training is completed by judging the relationship between the index training value and the preset index threshold, which improves the efficiency of model iteration optimization and the accuracy of stopping training. At the same time, by continuously adjusting and training the model until the performance reaches the preset requirements, the accuracy of the model for classifying environmental sounds is improved, and the generalization ability is also enhanced.
[0043] In one embodiment, in step S20, that is, after the sound classification result corresponding to the real-time environmental sound data is obtained by performing data analysis and processing on the sound sample to be analyzed through the preset classification model, it includes: S204. Generate a new sound sample according to the real-time environmental sound data and the sound classification result; S205. Update and train the preset classification model according to the newly added sound sample to obtain an updated preset classification model.
[0044] Understandably, after obtaining the sound classification result corresponding to the real-time environmental sound data through the preset classification model, in order to realize the continuous learning and update of the preset classification model, it is necessary to integrate the real-time environmental sound data and the sound classification result to form a newly added sound sample containing the sound data and the corresponding category label (i.e., the classification result). That is, generate a newly added sound sample according to the real-time environmental sound data and the sound classification result. The newly added sound sample refers to the sound sample data used to update the already trained preset classification model. Input the newly added sound sample as training data into the preset classification model, update the model parameters, and obtain an updated preset classification model. In addition, after the update training is completed, use the test set data to evaluate the metrics of the updated preset classification model. If the evaluation result shows that the model performance has improved, save the updated preset classification model for use in the subsequent operation process of the environmental sound classification method.
[0045] In this embodiment, the new environmental sound data is used as training samples to continuously update and optimize the model, ensuring that the preset classification model can adapt to environmental sound changes and improving the reliability of the model and the accuracy of sound classification.
[0046] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0047] In one embodiment, an environmental sound classification device is provided, and the environmental sound classification device corresponds one-to-one with the environmental sound classification method in the above embodiment. As Figure 3 shown, the environmental sound classification device includes a sound preprocessing module 10 and a sound classification module 20. The detailed description of each functional module is as follows: The sound preprocessing module 10 is used to obtain real-time environmental sound data and preprocess the real-time environmental sound data to obtain a sound sample to be analyzed; The sound classification module 20 is used to perform data analysis and processing on the sound sample to be analyzed through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthetic sound sample, and the synthetic sound sample is obtained by performing data synthesis processing on the initial sound sample by a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0048] In one embodiment, the sound preprocessing module 10 includes: A Fourier transform processing unit for performing Fourier transform processing on the real-time environmental sound data to obtain a linear spectrogram; A spectrogram conversion processing unit for converting the linear spectrogram into a Mel spectrogram; A normalization processing unit for performing normalization processing on the Mel spectrogram to obtain a sound sample to be analyzed.
[0049] In one embodiment, the sound classification module 20 includes: An initial sound sample determination unit for acquiring historical environmental sound data and performing preprocessing on the historical environmental sound data to obtain an initial sound sample; A synthesized sound sample determination unit for performing data synthesis processing on the initial sound sample through a preset synthesis model to obtain a synthesized sound sample; A classification model training unit for acquiring an initial classification model and training the initial classification model based on the initial sound sample and the synthesized sound sample to obtain a preset classification model.
[0050] In one embodiment, the sound classification module 20 further includes: A spectrogram synthesis processing unit for performing spectrogram synthesis processing on the initial sound sample through the generative adversarial network to obtain an initial synthesized sample; A denoising processing unit for performing denoising processing on the initial synthesized sample through the diffusion model network to obtain an optimized synthesized sample; A post-processing unit for performing post-processing on the optimized synthesized sample to obtain a synthesized sound sample.
[0051] In one embodiment, the sound classification module 20 further includes: A convolutional base network determination unit for determining a convolutional base network according to a first pre-trained network, a second pre-trained network, and a third pre-trained network; A classification layer network determination unit for constructing a classification layer network based on a linear layer; An initial classification model construction unit for obtaining an initial classification model according to the convolutional base network and the classification layer network.
[0052] In one embodiment, the sound classification module 20 further includes: A training sample data determination unit for determining training sample data of the initial classification model according to the initial sound sample and the synthesized sound sample; An index training value acquisition unit for performing transfer learning training on the initial classification model through the training sample data to obtain an index training value; An index training value judgment unit for judging whether the index training value reaches a preset index threshold; A preset classification model determination unit, configured to, if the index training value has not reached the preset index threshold, adjust the parameters of the initial classification model, and perform transfer learning training on the optimized classification model obtained by parameter adjustment through the training sample data until the index training value of the optimized classification model reaches the preset index threshold, and determine the optimized classification model as the preset classification model.
[0053] In one embodiment, the voice classification module 20 further includes: A new voice sample generation unit, configured to generate a new voice sample according to the real-time environmental voice data and the voice classification result; A model update unit, configured to perform update training on the preset classification model according to the new voice sample to obtain an updated preset classification model.
[0054] For the specific limitations of the environmental voice classification device, reference may be made to the limitations of the environmental voice classification method in the foregoing text, which will not be elaborated herein. Each module in the above environmental voice classification device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0055] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database of the computer device is used to store the data involved in the environmental voice classification method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, an environmental voice classification method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0056] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor. When the processor executes the computer-readable instructions, the following steps are implemented: Obtain real-time environmental voice data, and preprocess the real-time environmental voice data to obtain a voice sample to be analyzed; Performing data analysis and processing on the to-be-analyzed sound sample through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthesized sound sample, the synthesized sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0057] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. Computer-readable instructions are stored on the readable storage media, and when the computer-readable instructions are executed by one or more processors, the following steps are implemented: Obtaining real-time environmental sound data, and performing preprocessing on the real-time environmental sound data to obtain a to-be-analyzed sound sample; Performing data analysis and processing on the to-be-analyzed sound sample through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; the preset classification model is obtained by training an initial classification model with an initial sound sample and a synthesized sound sample, the synthesized sound sample is obtained by performing data synthesis processing on the initial sound sample through a preset synthesis model, and the preset synthesis model includes a generative adversarial network and a diffusion model network.
[0058] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0059] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0060] The non-company software tools or components that appear in the embodiments of this application are only introduced by way of example and do not represent actual use. The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A method for classifying environmental sounds, characterized in that: include: Acquire real-time environmental sound data, and pre-process the real-time environmental sound data to obtain a sound sample to be analyzed; Performing data analysis and processing on the sound sample to be analyzed by using a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; The preset classification model is obtained by training the initial classification model with initial sound samples and synthetic sound samples. The synthetic sound samples are obtained by performing data synthesis processing on the initial sound samples by the preset synthesis model. The preset synthesis model includes a generative adversarial network and a diffusion model network.
2. The environmental sound classification method according to claim 1, characterized in that: The preprocessing of the real-time environmental sound data to obtain a sound sample to be analyzed includes: Performing Fourier transform processing on the real-time environmental sound data to obtain a linear spectrum diagram; Converting the linear spectrogram into a Mel spectrogram; The Mel-spectrogram is standardized to obtain a sound sample to be analyzed.
3. The environmental sound classification method according to claim 1, characterized in that: Before performing data analysis on the sound sample to be analyzed by using a preset classification model, the method includes: Acquiring historical environmental sound data, and preprocessing the historical environmental sound data to obtain an initial sound sample; Performing data synthesis processing on the initial sound sample through a preset synthesis model to obtain a synthesized sound sample; An initial classification model is obtained, and the initial classification model is trained based on the initial sound sample and the synthesized sound sample to obtain a preset classification model.
4. The environmental sound classification method according to claim 3, characterized in that: The step of performing data synthesis processing on the initial sound sample by using a preset synthesis model to obtain a synthesized sound sample includes: Performing spectrum synthesis processing on the initial sound sample through the generative adversarial network to obtain an initial synthesized sample; Performing denoising processing on the initial synthetic sample through the diffusion model network to obtain an optimized synthetic sample; The optimized synthesized sample is post-processed to obtain a synthesized sound sample.
5. The environmental sound classification method according to claim 3, characterized in that: Before obtaining the initial classification model, the method includes: Determine a convolutional base network according to the first pre-trained network, the second pre-trained network, and the third pre-trained network; Build a classification layer network based on the linear layer; An initial classification model is obtained according to the convolutional base network and the classification layer network.
6. The environmental sound classification method according to claim 3, characterized in that: The training of the initial classification model based on the initial sound sample and the synthesized sound sample to obtain a preset classification model includes: Determine the training sample data of the initial classification model according to the initial sound sample and the synthesized sound sample; Performing transfer learning training on the initial classification model through the training sample data to obtain an indicator training value; Determine whether the indicator training value reaches a preset indicator threshold; If the indicator training value has not reached the preset indicator threshold, the parameters of the initial classification model are adjusted, and the optimized classification model obtained by the parameter adjustment is subjected to transfer learning training through the training sample data until the indicator training value of the optimized classification model reaches the preset indicator threshold, and the optimized classification model is determined as the preset classification model.
7. The environmental sound classification method according to claim 1, characterized in that: After performing data analysis and processing on the sound sample to be analyzed by using a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data, the method includes: Generate a new sound sample according to the real-time environmental sound data and the sound classification result; The preset classification model is updated and trained according to the newly added sound samples to obtain an updated preset classification model.
8. An environmental sound classification device, characterized in that: include: A sound preprocessing module is used to obtain real-time environmental sound data, preprocess the real-time environmental sound data, and obtain a sound sample to be analyzed; A sound classification module is used to perform data analysis and processing on the sound sample to be analyzed through a preset classification model to obtain a sound classification result corresponding to the real-time environmental sound data; The preset classification model is obtained by training the initial classification model with initial sound samples and synthetic sound samples. The synthetic sound samples are obtained by performing data synthesis processing on the initial sound samples by the preset synthesis model. The preset synthesis model includes a generative adversarial network and a diffusion model network.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that: When the processor executes the computer-readable instructions, the environmental sound classification method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by one or more processors, the one or more processors perform the environmental sound classification method according to any one of claims 1 to 7.
Citation Information
Cited By
Ambient sound feature-based authenticity analysis method, apparatus and device, and medium
CN120612960A