Environmental sound classification method, model, medium and equipment based on neural evolution
By adopting a neural evolution-based environmental sound classification method in intelligent wireless acoustic sensor network, the feature hyperparameters, convolutional neural network hyperparameters and cross-layer connection structure are optimized, and the problem of sound event classification in resource-constrained environments is solved, achieving efficient and low-complexity classification effect.
Patent Information
- Application Number
- CN202510346352.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing intelligent wireless acoustic sensor network (WASN) is difficult to achieve efficient sound event classification in resource-constrained environments, and the network structure design and hyperparameter adjustment of existing deep learning methods lack guiding principles, resulting in cumbersome and time-consuming manual testing and evaluation.
A neural evolution-based environmental sound classification method is adopted to search for feature hyperparameters, convolutional neural network hyperparameters and cross-layer connection structures within the preset feasible domain, and optimize these parameters using neural evolution algorithms to build a neural network model with low time complexity and low spatial complexity.
Efficient sound classification on resource-constrained devices is achieved. The optimized network has significantly improved the classification effect, and the network can be minimized at the same time without any degradation in performance.
Smart Images

Figure CN119864054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental sound classification, and in particular to a method for environmental sound classification based on neuroevolution, an environmental sound classification model based on neuroevolution, a computer-readable storage medium, and an electronic device. Background Art
[0002] Driven by the rapid increase in computing power and data quality, the application of machine learning and artificial intelligence in IoT systems has increased significantly, forming the research topic of AI+IoT systems (Artificial Intelligence of Things, AIoT). The integration of artificial intelligence and IoT devices in various fields of daily life has profoundly affected fields such as smart wearable technology, intelligent environmental monitoring, autonomous vehicles and advanced manufacturing systems.
[0003] Smart Wireless Acoustic Sensor Networks (WASN) is a typical AIOT system that realizes the perception of environmental sound information by processing the sound signals received by the sensors. In most smart wireless acoustic sensor network systems, sound information perception aims to solve the problem of sound event classification, that is, to classify the type of active sound in a given audio signal. Sensor nodes in a wireless network are connected to each other through wireless communication links to form a multi-hop network topology. Each sensor node has a certain computing and storage capacity, which can process and analyze the received sound signals, and then send the sound signal characteristics or classification results to the cloud to provide information for the decision-making of other devices in the smart Internet of Things. For example, WASN nodes scattered over a wide area transmit the results of Environment Sound Classification (ESC) to the data collection node; the data collection node acts as an information hub and uploads the fused information to the cloud server; in the cloud, the application generates a series of operation instructions based on the received data, combined with advanced algorithms and decision logic; these instructions will eventually be issued to the actuator device to trigger the corresponding physical action or response.
[0004] In practical applications, most WASNs face the challenge of having to operate in the wild or in environments with a lack of power supply. At the same time, in order to achieve cost-effective and widespread deployment, the manufacturing cost of each device must be strictly controlled. Given these limitations, WASN devices are usually severely constrained in terms of computing resources, energy supply, and budget, making those sound event classification algorithms with higher complexity unsuitable for application in such systems. Therefore, in order to meet practical needs, the sound event classification algorithms deployed on WASN nodes must have the characteristics of low time complexity, low space complexity, and high accuracy. This means that the algorithm needs to optimize its performance and efficiency as much as possible without sacrificing accuracy, so that it can be successfully executed on resource-constrained hardware. The performance of deep learning-based methods depends on their architecture and the choice of hyperparameters. At present, there is no guiding principle for network structure design and hyperparameter adjustment. Designing a suitable network structure and hyperparameter adjustment is regarded as a black box optimization process, which mainly relies on manual testing and evaluation. However, manual testing and evaluation is a tedious and time-consuming process that requires experimenters to have rich experience and a lot of expertise. Summary of the invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a method for classifying environmental sounds based on neuroevolution, which has low temporal complexity and low spatial complexity and can achieve high classification accuracy.
[0006] The present invention also proposes an environmental sound classification model based on neural evolution.
[0007] The present invention also provides a computer-readable storage medium.
[0008] The invention also provides an electronic device.
[0009] According to the first aspect of the present invention, the method for classifying environmental sounds based on neural evolution comprises the following steps:
[0010] S000, search for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures within the preset feasible domain;
[0011] S100, optimizing the feature hyperparameters, the convolutional neural network hyperparameters, and the cross-layer connection structure;
[0012] S200, constructing a network model according to the optimization result, and evaluating the fitness of the optimization result by training the network model;
[0013] The step S100 includes the following steps:
[0014] S110, algorithm initialization;
[0015] S120, encoding the cross-layer connection structure and sending it together with the feature hyperparameters and the convolutional neural network hyperparameters to the network model;
[0016] S130, calculating the fitness of the optimization result;
[0017] S140, using a neuroevolution algorithm to obtain the optimization result;
[0018] S150, determine whether the iteration condition is met, if so, execute step S120, otherwise update the optimization result,
[0019] The step 200 comprises the following steps:
[0020] S210, obtaining a training set and a test set;
[0021] S220, extracting features from the original ambient sound signal according to the feature hyperparameters of step S120;
[0022] S230, processing the extracted feature graph using a convolutional neural network according to the convolutional neural network hyperparameters and the cross-layer connection structure of step S120;
[0023] S240, using the output layer to output the classification result score to the step S130,
[0024] The convolutional neural network includes multiple convolutional network modules and multiple multi-layer perceptrons, the convolutional network modules sequentially include convolutional layers, maximum pooling layers, batch normalization layers and ReLU activation functions, the multi-layer perceptron includes a fully connected layer and the ReLU activation function, the output layer includes the fully connected layer and the Sigmoid activation function, and encoding the cross-layer connection structure includes encoding the connection between the convolutional network modules into a gene sequence, so that the neural evolution algorithm autonomously decides whether to establish cross-layer connections between the convolutional network modules.
[0025] According to the environmental sound classification method based on neuroevolution according to the embodiment of the present invention, the hyperparameters and network architecture of the model are adjusted in multiple rounds of iterations to explore a neural network structure that has both low time complexity and low space complexity and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but the network can also be minimized at the same time with almost no performance degradation. This provides a new solution for efficient sound classification on resource-constrained devices.
[0026] In addition, the environmental sound classification method based on neuroevolution according to the embodiment of the present invention also has the following additional technical features:
[0027] According to some embodiments of the present invention, the output feature maps of different convolutional network modules are adjusted to a uniform size through an adaptive average pooling operation, and then matched with the minimum size in the original feature map, and the resized output feature maps are fused using a channel-level splicing method.
[0028] According to some embodiments of the present invention, a (N+1)×(N+1) truth table matrix is used to represent the connection relationship between the N convolutional network modules, the N convolutional network modules are numbered 0, 1…N-1, respectively, the horizontal axis of the truth table matrix represents the information outflow layer, the vertical axis represents the information inflow layer, and it is specified that i The convolutional network module does not input the 0th to ( i -1) of the convolutional network modules; the first column of the truth table matrix represents the connection relationship between the input module and each of the convolutional network modules, the (N+1)th row represents the connection relationship between each of the convolutional network modules and the output module, and the i Line j The column elements represent the j -1) the convolutional network module and the i The connection relationship between the convolutional network modules.
[0029] According to some embodiments of the present invention, the feature extraction includes signal pre-emphasis, short-time Fourier transform, Mel filter bank processing and logarithmic operation, and the signal pre-emphasis is calculated according to the following formula:
[0030] (1)
[0031] In the formula, x ( t ) is the original ambient sound signal; t is the time variable; is the signal after pre-emphasis; α is a constant coefficient, and the short-time Fourier transform is performed according to the following formula:
[0032] (2)
[0033] In the formula, It's time t frame and frequency k A two-dimensional function of t frame As the starting position of the time frame, it indicates the position of the analyzed time segment in the signal, k The unit is Hz; is the pre-emphasized signal; n is the time sampling point index; w( t ) is the window function; N is the length of the window function, and the Mel filter bank processing is performed according to the following formula:
[0034] (3)
[0035] In the formula, Indicates m Mel filters; f ( m ) is defined as:
[0036] (4)
[0037] In the formula, M is the total number of the Mel filters, m [1, M ]; k l is the lowest frequency in the filter frequency range; k h is the highest frequency within the filter frequency range; k s is the sampling frequency; is in Mel scale and is calculated as follows:
[0038] (5)
[0039] In the formula, express k l or k h ; for is the inverse function of and is calculated according to the following formula:
[0040] (6)
[0041] The Mel-frequency cepstrum is calculated according to the following formula:
[0042] (7)
[0043] In the formula, is the Mel spectrum.
[0044] According to some embodiments of the present invention, each of the convolutional neural networks is taken as an individual, the convolutional neural network hyperparameters are taken as genes, and all individuals in each iteration are taken as a population, and the neuroevolution algorithm comprises the following steps:
[0045] S141, initializing parameters of the population;
[0046] S142, calculating the fitness of the individual;
[0047] S143, taking the individual with the best fitness value as the teacher individual, and the remaining individuals as student individuals, and calculating the average parameters of all the student individuals;
[0048] S144, entering the teaching stage, the individual student learns from the individual teacher;
[0049] S145, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters;
[0050] S146, entering the learning phase, the students learn from each other;
[0051] S147, determining whether the genes of all the individuals exceed the boundary, if so, modifying the genes, otherwise retaining the genes;
[0052] S148, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters;
[0053] S149, determine whether the algorithm stopping condition is met, if so, output the optimal individual parameters, otherwise repeat steps S142 to S149.
[0054] In some embodiments of the present invention, the fitness of the individual is calculated according to the following formula:
[0055] (8)
[0056] In the formula, x is the convolutional neural network hyperparameter; Acc ( x ) is adopted x The prediction accuracy of the convolutional neural network on the test set; T ( x ) is the time complexity of the convolutional neural network; C ( x ) is the spatial complexity of the convolutional neural network; a , b , c is the weight coefficient.
[0057] In some specific embodiments of the present invention, during the teaching phase, the individual students improve their fitness according to the following formula:
[0058] (9)
[0059] In the formula, A new gene for the student individual; is the current gene of the student individual; rand is a random number in the interval [0,1]; β = round (1+ rand (0,1)) is the teaching factor, round Indicates rounding; represents the average of all the individual genes of the teachers; represents the average of all the individual genes of the students;
[0060] During the learning stage, students With individual students Mutual learning, the student individuals with low fitness learn from the student individuals with high fitness, and the learning process is expressed as:
[0061] (10)
[0062] According to the second aspect of the present invention, the environmental sound classification model based on neural evolution includes: a problem modeling layer, the problem modeling layer is used to search for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures within a preset feasible domain; an algorithm optimization layer, the algorithm optimization layer is used to optimize the feature hyperparameters, the convolutional neural network hyperparameters, and the cross-layer connection structure; a model application layer, the model application layer is used to build a network model according to the optimization results, and evaluate the fitness of the optimization results by training the network model; the algorithm optimization layer includes an algorithm initialization module, an encoding module, a fitness calculation module, a neural evolution module and an algorithm output module, the encoding module is used to encode the cross-layer connection structure and send it to the network model together with the feature hyperparameters and the convolutional neural network hyperparameters, the fitness calculation module is used to calculate the fitness of the optimization result, the neural evolution module adopts a neural evolution algorithm to obtain the optimization result, and the algorithm output module The block updates the optimization result when the iteration condition is not met; the model application layer includes an acquisition module, a feature extraction module, a convolutional neural network and an output layer, the acquisition module is used to acquire a training set and a test set, the feature extraction module performs feature extraction on the original environmental sound signal according to the feature hyperparameters, the convolutional neural network is used to process the extracted feature map, the convolutional neural network includes a plurality of convolutional network modules and a plurality of multi-layer perceptrons, the convolutional network module includes a convolutional layer, a maximum pooling layer, a batch normalization layer and a ReLU activation function in sequence, the multi-layer perceptron includes a fully connected layer and the ReLU activation function, the output layer is used to output the classification result score to the fitness calculation module, the output layer includes the fully connected layer and the Sigmoid activation function, the encoding module encodes the connection between the convolutional network modules into a gene sequence, so that the neural evolution algorithm autonomously decides whether to establish a cross-layer connection between the convolutional network modules.
[0063] The environmental sound classification model based on neuroevolution according to the embodiment of the present invention has low time complexity and low space complexity, and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also can be minimized at the same time, with almost no performance degradation. This provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0064] According to the computer-readable storage medium of the third aspect embodiment of the present invention, a computer program is stored thereon, and when the computer program is executed by a processor, the method for classifying environmental sounds based on neural evolution as described in the first aspect embodiment of the present invention is implemented.
[0065] According to the computer-readable storage medium of the embodiment of the present invention, the method implemented has both low time complexity and low space complexity, and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also the network can be minimized at the same time, and the performance is almost not reduced. It provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0066] According to an electronic device of an embodiment of the fourth aspect of the present invention, the electronic device includes a processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and when the computer program is executed by the processor, the environmental sound classification method based on neural evolution as described in the embodiment of the first aspect of the present invention is implemented.
[0067] According to the electronic device of the embodiment of the present invention, the method implemented has both low time complexity and low space complexity and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also the network can be minimized at the same time with almost no performance degradation. This provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0068] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a schematic diagram of an environmental sound classification method based on neural evolution according to an embodiment of the present invention;
[0070] Figure 2 is a schematic diagram of an environmental sound classification model based on neural evolution according to an embodiment of the present invention;
[0071] Figure 3 is a schematic diagram of a Mel spectrogram of an environmental sound classification method based on neural evolution according to an embodiment of the present invention;
[0072] Figure 4 is a schematic diagram of an adaptive average pooling operation of a method for classifying environmental sounds based on neural evolution according to an embodiment of the present invention;
[0073] Figure 5 is a schematic diagram of a truth table matrix of an environmental sound classification method based on neural evolution according to an embodiment of the present invention;
[0074] Figure 6 is a schematic diagram of a truth table matrix application according to an embodiment of the present invention;
[0075] Figure 7 Schematic diagram of the structure of an experimental platform according to an embodiment of the present invention. DETAILED DESCRIPTION
[0076] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0077] The following describes the environmental sound classification method based on neural evolution according to the first embodiment of the present invention with reference to the accompanying drawings.
[0078] like Figure 1-Figure 2 As shown, the environmental sound classification method based on neural evolution according to an embodiment of the present invention includes the following steps:
[0079] S000, looking for low-complexity, high-accuracy network models within the preset feasible domain, including searching for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures;
[0080] S100, optimizes feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures;
[0081] S200, constructing a network model according to the optimization result, and evaluating the fitness of the optimization result by training the network model;
[0082] Step S100 includes the following steps:
[0083] S110, algorithm initialization;
[0084] S120, encoding the cross-layer connection structure and sending it to the network model together with the feature hyperparameters and the convolutional neural network hyperparameters;
[0085] S130, calculating the fitness of the optimization result;
[0086] S140, using a neuroevolutionary algorithm to obtain optimization results;
[0087] S150, determine whether the iteration condition is met, if so, execute step S120, otherwise update the optimization result,
[0088] Step 200 includes the following steps:
[0089] S210, obtaining a training set and a test set;
[0090] S220, extracting features from the original ambient sound signal according to the feature hyperparameters of step S120;
[0091] S230, processing the extracted feature graph using a convolutional neural network according to the convolutional neural network hyperparameters and cross-layer connection structure of step S120;
[0092] S240, using the output layer to output the classification result score to step S130,
[0093] like Figure 2 As shown in the figure, the convolutional neural network includes multiple convolutional network modules and multiple multi-layer perceptrons. The convolutional network module includes convolutional layers, maximum pooling layers, batch normalization layers, and ReLU activation functions in sequence. Such a structure helps to gradually extract and refine features. The multi-layer perceptron includes a fully connected layer and a ReLU activation function to further process the information from the convolution part. The output layer includes a fully connected layer and a Sigmoid activation function to generate the final prediction result.
[0094] Cross-layer connection generally refers to the direct information flow path between different levels in the network. The cross-layer connection structure includes residual connection in residual network, which skips one or more layers by directly adding the input to the output; and dense connection in densely connected network, which ensures that each layer is directly connected to all previous layers. Cross-layer connection significantly improves the parameter utilization and training efficiency of neural network by optimizing the propagation of information and gradient, while enhancing the feature extraction ability and generalization performance, and plays a powerful regularization role. The present invention adopts an encoded cross-layer connection method to encode the connection between convolutional network modules into a gene sequence, so that the neuroevolution algorithm can autonomously decide whether to establish cross-layer connection between each convolutional network module. Through this connection mode, the neuroevolution algorithm can dynamically adjust its cross-layer connection configuration by adjusting the encoding, which brings greater flexibility and automation to the design of neural network architecture and improves the performance and generalization ability of the model.
[0095] According to the environmental sound classification method based on neuroevolution according to the embodiment of the present invention, the hyperparameters and network architecture of the model are adjusted in multiple rounds of iterations to explore a neural network structure that has both low time complexity and low space complexity and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but the network can also be minimized at the same time with almost no performance degradation. This provides a new solution for efficient sound classification on resource-constrained devices.
[0096] According to some embodiments of the present invention, feature extraction includes signal pre-emphasis, short-time Fourier transform, Mel filter bank processing and logarithmic operation. In order to compensate for the high-frequency part suppressed during the sound propagation process, a high-pass filter is used to enhance the high-frequency part of the sound signal. The signal pre-emphasis is calculated according to the following formula:
[0097] (1)
[0098] In the formula, x ( t ) is the original environmental sound signal; t is the time variable; is the signal after pre-emphasis; α is a constant coefficient.
[0099] The short-time Fourier transform analyzes the time-frequency characteristics of the signal by introducing a sliding window mechanism. As the window gradually slides in the time domain, the signal in each window is Fourier transformed. By splicing the results of each Fourier transform, a time-frequency signal whose frequency changes with time is obtained. The short-time Fourier transform is performed according to the following formula:
[0100] (2)
[0101] In the formula, It's time t frame and frequency k A two-dimensional function of t frame As the starting position of the time frame, it indicates the position of the analyzed time segment in the signal, k The unit is Hz; is the signal after pre-emphasis; n is the time sampling point index; w ( t ) is the window function; N is the length of the window function, that is, the number of sampling points per frame.
[0102] The Mel filter bank is a series of triangular filters with a response value of 1 at the center frequency and attenuation to 0 at the center points of the filters on both sides. The Mel filter bank is processed according to the following formula:
[0103] (3)
[0104] In the formula, Indicates m Mel filters; f ( m ) is defined as:
[0105] (4)
[0106] In the formula, M is the total number of Mel filters, m [1, M ]; k l is the lowest frequency in the filter frequency range; k his the highest frequency within the filter frequency range; k s is the sampling frequency; Mel scale is a nonlinear scale used to describe the different sensitivity of human ears to sounds in different frequency domains. It is calculated according to the following formula:
[0107] (5)
[0108] In the formula, express k l or k h ; for is the inverse function of and is calculated according to the following formula:
[0109] (6)
[0110] The Mel-frequency cepstrum is calculated according to the following formula:
[0111] (7)
[0112] In the formula, is the Mel spectrum. Taking the logarithm of the Mel spectrum output by the Mel filter bank, we can get the Mel cepstrum.
[0113] The above process effectively refers to the difference in human sensitivity to different frequencies. By using the Mel frequency scale to perform nonlinear mapping of the signal's time-frequency spectrum, the resolution of the low-frequency part of the time-frequency spectrum is enhanced, and size compression is effectively achieved. Figure 3 In the figure, a visual comparison of 4 different signals in time domain, time-frequency spectrum and Mel spectrum is shown. In the figure, the Mel spectrum has a narrower frequency interval in the high-frequency region compared to the time-frequency spectrum. Since most of the effective information of the sound signal is usually concentrated in the low-frequency region, this feature extraction method is conducive to improving the classification accuracy of the subsequent processing module.
[0114] In the basic framework of neural networks, the parameters involved in neural evolution cover multiple parts: Mel spectrogram extraction module, convolutional network module and multi-layer perceptron module. Among them, the parameters of the Mel spectrogram extraction module include the window length and stride between frames in the short-time Fourier transform, as well as the number of filters in the Mel filter bank. These parameters directly affect the efficiency and accuracy of the conversion from the original audio signal to the Mel spectrogram. The parameters of the convolutional network module involve the depth of the module, that is, the number of convolutional layers included, and the number of output channels of the feature map generated by each convolutional layer. These parameters determine the network's ability to extract features, as well as the model complexity and computational cost. The parameters of the multi-layer perceptron module include the number of multi-layer perceptron modules and the number of neurons in the fully connected layer in each multi-layer perceptron module. These parameters define the network's ability to process high-level abstract information, as well as the depth and width of the network structure.
[0115] To solve the problem that the output feature maps of different convolutional network modules have different sizes, the present invention uses an adaptive splicing technology to integrate output feature maps of different sizes. Figure 4 As shown in the figure, the output feature maps of different convolutional network modules are adjusted to a uniform size through adaptive average pooling operation, and then matched with the minimum size in the original feature map; in order to effectively fuse these resized output feature maps, channel-level splicing is used to fuse the resized output feature maps, such as the U-Net network. In this way, the consistency of the feature maps in the channel dimension can be ensured, so that information from different sources can be effectively integrated at the channel level.
[0116] According to some embodiments of the present invention, Figure 5 As shown in Figure 1, a (N+1)×(N+1) truth table matrix is used to represent the connection relationship between N convolutional network modules. The N convolutional network modules are numbered 0, 1…N-1 respectively. The horizontal axis of the truth table matrix represents the information outflow layer, and the vertical axis represents the information inflow layer. In order to prevent the input and output from forming a loop, it is stipulated that i The convolutional network module does not input the 0th to ( i -1) convolutional network modules, that is, the truth table matrix is a lower triangular matrix. The first column of the truth table matrix represents the connection relationship between the input module and each convolutional network module. The input module inputs the Mel-frequency cepstrum feature map after feature extraction to the convolutional network module; the (N+1)th row represents the connection relationship between each convolutional network module and the output module. The output module is a multi-layer perceptron; the i Line j The column elements represent the j -1) Convolutional network module and the iThe connection relationship between the convolutional network modules. When the element value in the truth table matrix is 1, it means it is connected; when the element value in the truth table matrix is 0, it means it is not connected. Figure 6 Taking the 4×4 truth table matrix on the left as an example, the truth table matrix represents the connection relationship between the three convolutional network modules, and the connection method is shown on the right side of the figure.
[0117] According to some embodiments of the present invention, each convolutional neural network is regarded as an individual, the convolutional neural network hyperparameters are regarded as genes, and all individuals in each iteration are regarded as a population, and the neuroevolution algorithm includes the following steps:
[0118] S141, initialize the population parameters;
[0119] S142, calculate the fitness of individuals;
[0120] S143, taking the individual with the best fitness value as the teacher individual, and the remaining individuals as the student individuals, and calculating the average parameters of all the student individuals;
[0121] S144, entering the teaching phase, individual students learn from individual teachers;
[0122] S145, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters;
[0123] S146, entering the learning stage, individual students learn from each other;
[0124] S147, determine whether the genes of all individuals exceed the boundary, if so, modify the genes, otherwise keep the genes;
[0125] S148, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters;
[0126] S149, determine whether the algorithm stopping condition is met, if so, output the optimal individual parameters, otherwise repeat steps S142 to S149.
[0127] In some embodiments of the present invention, the performance indicators of the neural network include the prediction accuracy of the network for the test set, the time complexity of the neural network and the parameter space complexity of the neural network. The fitness of the individual is calculated according to the following formula:
[0128] (8)
[0129] In the formula, x is the convolutional neural network hyperparameter; Acc ( x ) is adopted xThe prediction accuracy of the convolutional neural network on the test set; T ( x ) is the time complexity of the convolutional neural network; C ( x ) is the spatial complexity of the convolutional neural network; a , b , c is the weight coefficient. The weight coefficient can be obtained by randomly generating multiple hyperparameters and recording the prediction accuracy, time complexity, and space complexity of networks with different hyperparameters. The weight coefficients of different performance indicators are taken as the reciprocal of the mean of their recorded data. The computational complexity of neural networks is measured by the number of floating-point operations (FLOPS), where one floating-point operation is defined as one multiplication and one addition. The space complexity is defined as the sum of the memory space occupied by the total number of parameters and the memory space occupied by the output feature maps of each layer.
[0130] In some specific embodiments of the present invention, during the teaching phase, individual students improve their fitness according to the following formula:
[0131] (9)
[0132] In the formula, New genes for individual students; The current gene for the individual student; rand is a random number in the interval [0,1]; β = round (1+ rand (0,1)) is the teaching factor, round Indicates rounding, the teaching factor can be β Randomly determined to be 1 or 2; represents the average of all teachers’ individual genes; represents the average of all students' individual genes;
[0133] During the learning stage, students With individual students Mutual learning, students with low fitness learn from students with high fitness, and the learning process is expressed as:
[0134] (10)
[0135] Using the teaching and learning optimization algorithm, the hyperparameters in the feature extraction and convolutional network modules, as well as the connection modes between the convolutional network modules, are carefully adjusted in multiple rounds of iterations.
[0136] The following describes the ablation experiment and effect evaluation experiment of the present invention on the Unbransound 8k dataset. The Unbransound 8k dataset is an audio dataset specifically used for automatic urban environmental sound classification research, covering 10 flods, each of which includes 1) air conditioner; 2) car horn; 3) children playing; 4) dog bark; 5) drilling; 6) engine idling; 7) gun shot; 8) jackhammer; 9) siren; 10) street music. First, some data are selected as the evaluation criteria for neural evolution to achieve the optimization of feature hyperparameters, convolutional neural network hyperparameters and cross-layer connection structures. Next, the optimized neural network model is used to perform experimental verification on a computer, the purpose of which is to test the performance of the model at the theoretical level and whether it can achieve the expected effect. Finally, the optimized algorithm is deployed on an artificial intelligence Internet of Things platform based on STM32, and a physical experiment is carried out to evaluate the performance of the network framework in a practical environment.
[0137] Specifically, the audio signals of flod 1 and flod 2 in the Unbransound 8k dataset were used for the experiment. First, the sound signals in the flod 1 and flod 2 files were preprocessed, including reading the original audio data, adjusting the sampling rate to 6250Hz to meet the requirements of the AI IOT platform, and cutting the audio signals according to a 1-second time window. The sampling frequency of 6250Hz was determined after considering the oversampling multiple of the ADC chip and the delay-free requirement for data exchange between the MUC and the ADC. After preprocessing, 80% of the processed audio data was randomly selected as a training set for training the network model in the model application layer. The remaining 20% was used as a test set to evaluate the performance of the network model in subsequent steps. These performance evaluation results will be converted into classification result scores and input into the algorithm optimization layer to calculate the fitness of the network model. In the algorithm optimization layer, the teaching and learning optimization algorithm was selected as the main neuroevolution algorithm. In addition, in order to explore the role of the cross-layer connection module in the model application layer, this module was removed from the framework of the neuroevolution algorithm and an ablation experiment was performed. By comparing the performance of network models with and without the cross-layer connection module, the impact of this module on the performance of the network model is evaluated.
[0138] In the experiment, the number of iterations was determined to be 60, and the population size was set to 40 individuals. Training will continue until the performance of the network model does not improve further in 20 consecutive epochs, or until the total number of epochs trained reaches 500. After training, the performance of each network model is evaluated on the validation set, including calculating the classification accuracy, analyzing the time complexity of the neural network, and the parameter space complexity. These evaluation indicators are then used to calculate the fitness value of each network model through the fitness function.
[0139] After 60 iterations, the individuals with the best fitness, i.e., the optimal network model, were selected according to the results of the fitness function. The teacher-student interaction mechanism not only promotes effective knowledge transfer, but also helps the algorithm maintain a balance between exploration (exploring the unknown solution space) and utilization (utilizing known information) during the search process, thereby increasing the possibility of finding a better solution. In addition, in the ablation experiment, that is, when the cross-layer connection module was removed, the optimal iteration value of the neuroevolution framework only reached 0.8421. This result shows that cross-layer connections play a vital role in the neural network architecture. The cross-layer connection module expands the search space of the neuroevolution framework, thereby providing the possibility of finding a better solution. The presence of cross-layer connections allows information and gradients to flow more freely between different layers of the network, which helps to speed up the learning process and improve the overall performance of the network model.
[0140] In the physical hardware experiment, an artificial intelligence IoT platform based on the STM32 microcontroller was used. Figure 7 The system block diagram of the experimental platform is shown, including sensor module, processor module, positioning module, wireless communication module and power module. The central processing unit (MCU) uses the STM32H743 chip produced by ST, and the self-organizing network part is realized through Zigbee communication technology. The sensor module includes a 16-bit precision analog-to-digital converter (ADC) chip, a signal amplifier with a gain of 60 decibels (dB) and a microphone with a sensitivity of -26 dB.
[0141] According to the optimal network structure obtained in the neuroevolution algorithm experiment and the network parameters trained in the computer simulation experiment, the C language is used for manual reconstruction to generate the neural network code suitable for the STM32 chip. The algorithm is deployed on the AI-IOT experimental platform and takes a total of 1.8342 seconds to run, of which the environmental sound collection takes one second and the network running time is 0.8342 seconds. In order to test the actual application effect of this neural network, the sound signal is played through a Bluetooth wireless speaker at a distance of 15 meters from the AI-IOT platform. These sound signals will be received and recognized by the AI-IOT platform. The recognition results will be sent to a laptop connected to the receiving device through Zigbee wireless communication technology. The sound signals used for testing are selected from the untrained sound signals of the Ubransound 8k data set, and two samples of each sound type are randomly selected. The results show that the algorithm has good generalization ability.
[0142] According to the neural evolution-based environmental sound classification model of the second aspect of the present invention, Figure 1-Figure 2 As shown, it includes: problem modeling layer, algorithm optimization layer and model application layer.
[0143] Specifically, the problem modeling layer is used to search for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures within the preset feasible domain. The algorithm optimization layer is used to optimize feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures; the model application layer is used to build a network model based on the optimization results, and to evaluate the fitness of the optimization results by training the network model. The algorithm optimization layer includes an algorithm initialization module, an encoding module, a fitness calculation module, a neural evolution module, and an algorithm output module. The encoding module is used to encode the cross-layer connection structure and send it to the network model together with the feature hyperparameters and convolutional neural network hyperparameters. The fitness calculation module is used to calculate the fitness of the optimization result. The neural evolution module uses a neural evolution algorithm to obtain the optimization result. The algorithm output module updates the optimization result when the iteration conditions are not met. The model application layer includes an acquisition module, a feature extraction module, a convolutional neural network and an output layer. The acquisition module is used to obtain training sets and test sets. The feature extraction module extracts features from the original environmental sound signal according to feature hyperparameters. The convolutional neural network is used to process the extracted feature map. The convolutional neural network includes multiple convolutional network modules and multiple multi-layer perceptrons. The convolutional network module includes a convolutional layer, a maximum pooling layer, a batch normalization layer and a ReLU activation function in sequence. The multi-layer perceptron includes a fully connected layer and a ReLU activation function. The output layer is used to output the classification result score to the fitness calculation module. The output layer includes a fully connected layer and a Sigmoid activation function. The encoding module encodes the connection between the convolutional network modules into a gene sequence so that the neuroevolution algorithm can autonomously decide whether to establish cross-layer connections between the convolutional network modules.
[0144] The environmental sound classification model based on neuroevolution according to the embodiment of the present invention has low time complexity and low space complexity, and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also can be minimized at the same time, with almost no performance degradation. This provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0145] According to the computer-readable storage medium of the third aspect embodiment of the present invention, a computer program is stored thereon, and when the computer program is executed by a processor, the method for classifying environmental sounds based on neural evolution as described in the first aspect embodiment of the present invention is implemented.
[0146] According to the computer-readable storage medium of the embodiment of the present invention, the method implemented has both low time complexity and low space complexity, and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also the network can be minimized at the same time, and the performance is almost not reduced. It provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0147] According to the electronic device of the fourth aspect embodiment of the present invention, the electronic device includes a processor and a memory, the processor and the memory are connected, the memory is used to store a computer program, and when the computer program is executed by the processor, the environmental sound classification method based on neural evolution as described in the first aspect embodiment of the present invention is implemented.
[0148] According to the electronic device of the embodiment of the present invention, the method implemented has both low time complexity and low space complexity and can achieve high classification accuracy. The optimized network not only has a significant improvement in classification effect, but also the network can be minimized at the same time with almost no performance degradation. This provides a new solution for achieving efficient sound classification on resource-constrained devices.
[0149] Other structures and operations of the electronic device according to the embodiment of the present invention are known to those skilled in the art and will not be described in detail here.
[0150] In the description of the present invention, it is to be understood that the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0151] In the description of the present invention, "a first feature" or "a second feature" may include one or more of the features. A first feature "above" or "below" a second feature may include direct contact between the first and second features, or may include contact between the first and second features through another feature between them instead of direct contact. A first feature "above", "above" or "above" a second feature may include the first feature being directly above or obliquely above the second feature, or may simply mean that the first feature is higher in level than the second feature.
[0152] It should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0153] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "specific embodiments", "examples" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0154] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for classifying environmental sounds based on neural evolution, characterized in that: The following steps are involved: S000, search for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures within the preset feasible domain; S100, optimizing the feature hyperparameters, the convolutional neural network hyperparameters, and the cross-layer connection structure; S200, constructing a network model according to the optimization result, and evaluating the fitness of the optimization result by training the network model; The step S100 includes the following steps: S110, algorithm initialization; S120, encoding the cross-layer connection structure and sending it together with the feature hyperparameters and the convolutional neural network hyperparameters to the network model; S130, calculating the fitness of the optimization result; S140, using a neuroevolution algorithm to obtain the optimization result; S150, determine whether the iteration condition is met, if so, execute step S120, otherwise update the optimization result, The step 200 comprises the following steps: S210, obtaining a training set and a test set; S220, extracting features from the original ambient sound signal according to the feature hyperparameters of step S120; S230, processing the extracted feature graph using a convolutional neural network according to the convolutional neural network hyperparameters and the cross-layer connection structure of step S120; S240, using the output layer to output the classification result score to the step S130, The convolutional neural network includes multiple convolutional network modules and multiple multi-layer perceptrons, the convolutional network modules sequentially include convolutional layers, maximum pooling layers, batch normalization layers and ReLU activation functions, the multi-layer perceptron includes a fully connected layer and the ReLU activation function, the output layer includes the fully connected layer and the Sigmoid activation function, and encoding the cross-layer connection structure includes encoding the connection between the convolutional network modules into a gene sequence, so that the neural evolution algorithm autonomously decides whether to establish cross-layer connections between the convolutional network modules.
2. The method for environmental sound classification based on neural evolution according to claim 1, characterized in that: The output feature maps of different convolutional network modules are adjusted to a uniform size through an adaptive average pooling operation, and then matched with the minimum size in the original feature map, and the output feature maps after the resize are fused by channel-level splicing.
3. The method for classifying environmental sounds based on neural evolution according to claim 1, characterized in that: A (N+1)×(N+1) truth table matrix is used to represent the connection relationship between the N convolutional network modules, the N convolutional network modules are numbered 0, 1…N-1, respectively, the horizontal axis of the truth table matrix represents the information outflow layer, the vertical axis represents the information inflow layer, and the first i The convolutional network module does not input the 0th to ( i -1) of the convolutional network modules; the first column of the truth table matrix represents the connection relationship between the input module and each of the convolutional network modules, the (N+1)th row represents the connection relationship between each of the convolutional network modules and the output module, and the i Line j The column elements represent the j -1) the convolutional network module and the i The connection relationship between the convolutional network modules.
4. The method for environmental sound classification based on neural evolution according to claim 1, characterized in that: The feature extraction includes signal pre-emphasis, short-time Fourier transform, Mel filter bank processing and logarithmic operation, and the signal pre-emphasis is calculated according to the following formula: (1) In the formula, x ( t ) is the original ambient sound signal; t is the time variable; is the signal after pre-emphasis; α is a constant coefficient, and the short-time Fourier transform is performed according to the following formula: (2) In the formula, It's time t frame and frequency k A two-dimensional function of t frame As the starting position of the time frame, it indicates the position of the analyzed time segment in the signal, k The unit is Hz; is the pre-emphasized signal; n is the time sampling point index; w ( t ) is the window function; N is the length of the window function, and the Mel filter bank processing is performed according to the following formula: (3) In the formula, Indicates m Mel filters; f ( m ) is defined as: (4) In the formula, M is the total number of the Mel filters, m [1, M ]; k l is the lowest frequency in the filter frequency range; k h is the highest frequency within the filter frequency range; k s is the sampling frequency; is in Mel scale and is calculated as follows: (5) In the formula, express k l or k h ; for is the inverse function of and is calculated according to the following formula: (6) The Mel-frequency cepstrum is calculated according to the following formula: (7) In the formula, is the Mel spectrum.
5. The method for classifying environmental sounds based on neural evolution according to any one of claims 1 to 4, characterized in that: Taking each of the convolutional neural networks as an individual, the convolutional neural network hyperparameters as genes, and all individuals in each iteration as a population, the neuroevolution algorithm includes the following steps: S141, initializing parameters of the population; S142, calculating the fitness of the individual; S143, taking the individual with the best fitness value as the teacher individual, and the remaining individuals as student individuals, and calculating the average parameters of all the student individuals; S144, entering the teaching stage, the individual student learns from the individual teacher; S145, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters; S146, entering the learning phase, the students learn from each other; S147, determining whether the genes of all the individuals exceed the boundary, if so, modifying the genes, otherwise retaining the genes; S148, judging whether the individual after learning is better than the previous situation, if so, accepting the parameters of the individual after learning, otherwise retaining the previous parameters; S149, determine whether the algorithm stopping condition is met, if so, output the optimal individual parameters, otherwise repeat steps S142 to S149.
6. The method for environmental sound classification based on neural evolution according to claim 5, characterized in that: The fitness of the individual is calculated according to the following formula: (8) In the formula, x is the convolutional neural network hyperparameter; Acc ( x ) is adopted x The prediction accuracy of the convolutional neural network on the test set; T ( x ) is the time complexity of the convolutional neural network; C ( x ) is the spatial complexity of the convolutional neural network; a , b , c is the weight coefficient.
7. The method for environmental sound classification based on neuroevolution according to claim 6, characterized in that: During the teaching phase, the individual students improve their fitness according to the following formula: (9) In the formula, A new gene for the student individual; is the current gene of the student individual; rand is a random number in the interval [0,1]; β = round (1+ rand (0,1)) is the teaching factor, round Indicates rounding; represents the average of all the individual genes of the teachers; represents the average of all the individual genes of the students; During the learning stage, students With individual students Mutual learning, the student individuals with low fitness learn from the student individuals with high fitness, and the learning process is expressed as: (10)。 8. A neural evolution-based environmental sound classification model, characterized in that: include: A problem modeling layer, wherein the problem modeling layer is used to search for feature hyperparameters, convolutional neural network hyperparameters, and cross-layer connection structures within a preset feasible domain; An algorithm optimization layer, the algorithm optimization layer is used to optimize the feature hyperparameters, the convolutional neural network hyperparameters, and the cross-layer connection structure; A model application layer, the model application layer is used to build a network model according to the optimization result, and evaluate the fitness of the optimization result by training the network model; The algorithm optimization layer includes an algorithm initialization module, an encoding module, a fitness calculation module, a neural evolution module and an algorithm output module, wherein the encoding module is used to encode the cross-layer connection structure and send it to the network model together with the feature hyperparameters and the convolutional neural network hyperparameters, the fitness calculation module is used to calculate the fitness of the optimization result, the neural evolution module uses the neural evolution algorithm to obtain the optimization result, and the algorithm output module updates the optimization result when the iteration condition is not met; The model application layer includes an acquisition module, a feature extraction module, a convolutional neural network and an output layer. The acquisition module is used to acquire a training set and a test set. The feature extraction module performs feature extraction on the original ambient sound signal according to the feature hyperparameters. The convolutional neural network is used to process the extracted feature map. The convolutional neural network includes multiple convolutional network modules and multiple multi-layer perceptrons. The convolutional network module includes a convolutional layer, a maximum pooling layer, a batch normalization layer and a ReLU activation function in sequence. The multi-layer perceptron includes a fully connected layer and the ReLU activation function. The output layer is used to output the classification result score to the fitness calculation module. The output layer includes the fully connected layer and the Sigmoid activation function. The encoding module encodes the connection between the convolutional network modules into a gene sequence so that the neural evolution algorithm can autonomously decide whether to establish a cross-layer connection between the convolutional network modules.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the neuroevolution-based environmental sound classification method as described in any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the processor is connected to the memory, and the memory is used to store a computer program. When the computer program is executed by the processor, the neural evolution-based environmental sound classification method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Automated neural network generation using fitness estimation
CA3148847A1
Emotion recognition method and system based on voice text cross-modal fusion
CN117765981A