Audio signal identification method and device, kitchen range system and computer equipment
By using the audio signal recognition network of pulse-encoded neurons to encode the spectrum coefficients, the problem of high computing resources consumption in the prior art is solved, and efficient and accurate audio signal recognition is achieved, especially the recognition of ignition tones in smoke stove systems.
Patent Information
- Application Number
- CN202311565260.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing audio signal recognition technology, DNN calculation consumes too much space and time, and does not consider the historical information of the initial audio signal. RNN and LSTM calculations rely on a large amount of tag data, resulting in low computing efficiency.
The audio signal recognition network of pulse-encoded neurons is adopted. By encoding the spectrum coefficients, pulse-encoded neurons are used for activation processing, and combined with time memory, real-time recognition of audio signals is achieved.
It reduces computing resource consumption, improves recognition accuracy and efficiency, and is suitable for the recognition of simple monotonic audio, especially the recognition of ignition tones in smoke stove systems.
Smart Images

Figure CN120279895A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio recognition, and in particular, to an audio signal recognition method, apparatus, range hood system, and computer device. Background Art
[0002] With the development of artificial intelligence technology, audio recognition technology has emerged, which can automatically recognize whether a target audio is contained in an acquired audio signal. In the prior art, it is usually based on the DNN (Deep Neural Networks) architecture, taking the initial audio signal as the input, and using the results obtained by multiple DNN calculations for weighted calculation to obtain the recognition result; or replacing the DNN with an RNN (Recurrent Neural Network) or LSTM (Long Short Term Memory) with time memory, and then taking the time series of an initial audio as the input to obtain the recognition result.
[0003] However, there are some technical problems in the above prior art. For example, the weighted calculation of the DNN calculation results multiple times consumes a large amount of chip space and time. It often takes a lot of time to generate a calculation result, and the historical information generated during the calculation of the initial audio signal is not considered in the DNN calculation. When using RNN and LSTM to replace DNN, although the historical information of the data is integrated during the calculation, DNN, RNN, and LSTM are essentially ANN (Artificial Neural Network). The data during ANN calculation is usually floating-point numbers, and ANN usually also requires a large amount of labeled data to drive fitting.
[0004] Currently, no effective solution has been proposed for the problems of excessive calculation amount and low calculation efficiency during target audio recognition. Summary of the Invention
[0005] Based on this, it is necessary to provide an audio signal recognition method, apparatus, range hood system, and computer device for the above technical problems.
[0006] In a first aspect, this application provides an audio signal recognition method. The method includes:
[0007] Obtain a current audio signal;
[0008] Perform feature extraction processing on the current audio signal to obtain at least two current spectral coefficients, and perform encoding processing on the current spectral coefficients to obtain current encoded sequence data corresponding to the current spectral coefficients;
[0009] Input all the encoded sequence data into the trained audio signal recognition network, activate the pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain the recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
[0010] Thus, in this application, by encoding the spectral coefficients into encoded sequence data, it is possible to use the audio signal recognition network to recognize all the encoded sequence data, facilitating subsequent recognition processing using the audio signal recognition network. Further, when the target sound is a simple and monotonous audio, since the audio signal recognition network includes the above-mentioned pulse-coded neurons, it consumes less computing power and is closer to the biological model of synapses in the human brain. On the basis of ensuring the recognition accuracy, it can also make the transmission of information more efficient.
[0011] In one embodiment, after obtaining the recognition result corresponding to the current audio signal, the method further includes:
[0012] Obtain the next audio signal; perform feature extraction processing on the next audio signal to obtain the next spectral coefficients, and perform encoding processing on the next spectral coefficients to obtain the initial encoded sequence data;
[0013] Based on the length of the next encoded sequence data, determine and delete the sequence data to be deleted in the current encoded sequence data to obtain the remaining encoded sequence data; splice the remaining encoded sequence data with the initial encoded sequence data to obtain the next encoded sequence data;
[0014] Input the next encoded sequence data into the audio signal recognition network to obtain the next audio recognition result, and repeat the above steps until all audio signals are recognized; among them, based on all the audio recognition results, complete the real-time recognition processing for all audio signals.
[0015] It can be seen that after the current encoded sequence data is recognized, the initial encoded sequence data of the next audio obtained correspondingly is used to delete the current encoded sequence data, and the remaining encoded sequence data and the initial encoded sequence data are spliced together for the next round of recognition. This method can achieve streaming speech processing and can perform real-time recognition processing on multiple obtained audio signals in practical applications. Further, compared with independently recognizing each of the multiple obtained speeches, splicing the multiple speeches correspondingly and then recognizing in this application can make the recognition accuracy of the audio signal recognition network higher and is more suitable for the parameter configuration set in the network.
[0016] In one embodiment, obtaining the audio signal recognition network includes:
[0017] Obtain a preset training set of coding sequences, where the training set of coding sequences carries audio recognition labels;
[0018] Input the training set of coding sequences into the initial audio signal recognition network for training to obtain the predicted result of the training audio signal. Calculate the result of the loss function based on the predicted result of the training audio signal and the audio recognition label, and backpropagate the gradient of the result of the loss function to the initial audio signal recognition network for iterative training to generate a trained complete audio signal recognition network.
[0019] In one embodiment, input the current audio signal into the trained complete audio signal recognition network to obtain the recognition result corresponding to the current audio signal, including:
[0020] Through the current pulse coding neuron, obtain at least one current pulse result output by the historical pulse coding neuron; wherein, the current pulse result is obtained based on the current coding data in the current coding sequence data;
[0021] Based on the current pulse coding neuron, based on the current pulse result, the saved historical pulse result, and the preset pulse weight for the historical pulse coding neuron, obtain the next pulse result for the current coding data, and send the next pulse result to the next pulse coding neuron;
[0022] Repeat the above steps until all pulse coding neurons are traversed to obtain the recognition result for the current audio signal.
[0023] It can be seen from this that the pulse result transmitted to the next pulse coding neuron is jointly composed of the obtained current pulse result and the historical pulse result saved in the pulse coding neuron, so that the target audio recognition with time memory can be realized. Compared with the traditional neural network, it has time memory, improves the robustness of the neural network. Further, the above pulse weight can be configured by relevant technical personnel according to needs, so that the neural network can be applied to more complex actual environments to obtain more accurate recognition results.
[0024] In one embodiment, sending the next pulse result to the next pulse coding neuron includes:
[0025] When it is detected that the current pulse coding neuron is in the integration mode and the current pulse coding neuron detects that the next pulse result is greater than or equal to the preset pulse threshold, switch the integration mode to the pulse mode, and in the case that the current pulse coding neuron is in the pulse mode, send the next pulse result to the next pulse coding neuron;
[0026] When it is detected that the current pulse-coded neuron is in the integration mode and the next pulse result is less than the pulse threshold, accumulation is performed according to the current pulse result and the next pulse result in the integration mode to obtain a new pulse result, and when switching to the pulse mode, the new pulse result is sent to the next pulse-coded neuron.
[0027] It can be seen from this that different modes of the pulse-coded neuron determine whether to transmit the pulse output to the next neuron, greatly enhancing the sparsity of the network. When facing relatively simple and monotonous target sounds for recognition, a large amount of computing power is saved while ensuring the recognition accuracy.
[0028] In one embodiment, obtaining an identification result corresponding to the current audio signal based on the activated pulse-coded neuron includes:
[0029] Obtaining a preset class label; wherein, the class label corresponds to an output unit in the audio signal recognition network;
[0030] Based on the activated pulse-coded neuron, obtaining a pulse activation result, and obtaining an output result generated by the output unit based on the number of the pulse activation results;
[0031] If the output result is greater than or equal to a preset activation threshold, an identification result is obtained based on the class label corresponding to the output unit.
[0032] In one embodiment, encoding the current spectral coefficient to obtain current encoded sequence data corresponding to the current spectral coefficient includes:
[0033] Obtaining a current spectral pixel value corresponding to the current spectral coefficient;
[0034] Encoding the current spectral pixel value based on a preset encoding rule to obtain the current encoded sequence data.
[0035] It can be seen from this that in this application, encoding processing is performed on the spectral coefficients that can reflect the characteristics of the audio signal, and the two-dimensional characteristics of the spectrogram result composed of multiple spectral coefficients are used to encode the spectral coefficients into encoded sequence data, and the spectral coefficients are expressed by 0 and 1, making the processing of multiple spectral coefficients more efficient.
[0036] In a second aspect, this application also provides an audio signal recognition device. The device includes:
[0037] An acquisition module, configured to acquire a current audio signal;
[0038] A calculation module, configured to perform feature extraction processing on the current audio signal to obtain at least two current spectral coefficients, and perform encoding processing on the current spectral coefficients to obtain current encoded sequence data corresponding to the current spectral coefficients;
[0039] A generation module, configured to input all encoded sequence data into a trained audio signal recognition network, perform activation processing on pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain a recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
[0040] In a third aspect, the present application further provides a range hood and stove system, which includes a range hood and a stove.
[0041] The range hood is connected to the stove and is configured to obtain the current audio signal and perform any of the above-mentioned audio signal recognition methods based on the current audio signal.
[0042] It can be seen that the target sounds to be recognized in the range hood and stove system are mostly the ignition sounds generated by the stove. The ignition sounds have characteristics such as simplicity and monotony, and there are obvious differences between the ignition sounds and the ambient noise. Encoding the acquired audio signal in the form of 0 and 1 can more intuitively highlight whether the acquired audio signal contains the target audio. Further, the calculation based on the pulse-coded neurons saves a large amount of calculation costs and retains the time memory, making the recognition result more accurate.
[0043] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0044] Obtain the current audio signal;
[0045] Perform feature extraction processing on the current audio signal to obtain at least two current spectral coefficients, and perform encoding processing on the current spectral coefficients to obtain current encoded sequence data corresponding to the current spectral coefficients;
[0046] Input all the encoded sequence data into a trained audio signal recognition network, perform activation processing on pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain a recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
[0047] For the above-mentioned audio signal recognition method, device, range hood and stove system, the spectral coefficients that can reflect the characteristics of the current audio signal are encoded, which simplifies the data input into the audio recognition network. Further, each input of the encoded sequence data will cause a certain change in the potential of the pulse-coded neurons in the audio recognition network, so that the potential of the neurons contains the memory of the information of past excitations, improving the reliability of the judgment result. Description of the Drawings
[0048] Figure 1 It is an application environment diagram of the audio signal recognition method in an embodiment;
[0049] Figure 2 It is a schematic flowchart of the audio signal recognition method in an embodiment;
[0050] Figure 3 It is a schematic structural diagram of the audio signal recognition network in an embodiment;
[0051] Figure 4 It is a schematic flowchart of a neuron calculating the input encoded sequence data in an embodiment;
[0052] Figure 5 It is a schematic flowchart of the audio signal recognition method in a preferred embodiment;
[0053] Figure 6 It is a structural block diagram of the audio signal recognition device in an embodiment;
[0054] Figure 7 It is a structural block diagram of the smoke and stove system in an embodiment;
[0055] Figure 8 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] The audio signal recognition method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. First, a current audio signal is obtained, and feature extraction processing is performed on the signal to obtain at least two current spectral coefficients, and encoding is performed on them to obtain current encoded sequence data; and the current encoded sequence data is input into the audio signal recognition network, so as to obtain the recognition result output by the audio signal recognition network. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0058] In one embodiment, as Figure 2 shown, an audio signal recognition method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0059] Step S202, obtain the current audio signal.
[0060] Step S204, perform feature extraction processing on the current audio signal to obtain at least two current spectral coefficients, and perform encoding processing on the current spectral coefficients to obtain the current encoded sequence data corresponding to the current spectral coefficients.
[0061] Specifically, in practical applications, the current audio signal is usually converted into an MFCC spectrum based on feature processing. Multiple current spectral coefficients form the above MFCC spectrum. Each of the multiple current spectral coefficients is encoded to obtain the current encoded sequence data. Preferably, the current encoded sequence data is mostly a Poisson encoded sequence composed of 0 and 1. The above encoding processing can be phase encoding, delay encoding, etc.
[0062] Step S206, input all the encoded sequence data into a trained audio signal recognition network, activate the pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain the recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
[0063] Specifically, in practical applications, all the encoded sequence data is usually combined into a string of time-series encoded data, so that all the encoded sequence data is jointly used as the input of the audio signal recognition network. The above pulse-coded neurons can be an LIF model (leaky integrate-and-fire), a QIF model (Quadratic Integrate-and-Fire), etc. The above recognition result can include target sounds, environmental noises, etc. set by relevant technicians according to the actual application environment. The audio signal recognition network realizes the judgment of the input time-series encoded data, that is, the current audio signal, according to the neuron pulse signals activated by the encoded sequence data.
[0064] Through steps S202 to S206, the acquired current audio signal is converted into spectral coefficients with two-dimensional image features and encoded, effectively highlighting the difference between the target sound and ambient noise in the current audio signal; further, based on an audio signal recognition network including pulse-coded neurons, the encoded sequence data is calculated, and the corresponding recognition result is obtained by activating the neurons according to the encoded sequence data. This not only saves a large amount of computing costs, but also the potential information of the neurons contains the memory of the information of previous excitations, improving the reliability of the judgment result.
[0065] In one embodiment, after obtaining the recognition result corresponding to the current audio signal, the method further includes:
[0066] Obtain the next audio signal; perform feature extraction processing on the next audio signal to obtain the next spectral coefficients, and perform encoding processing on the next spectral coefficients to obtain initial encoded sequence data;
[0067] Based on the length of the next encoded sequence data, determine and delete the sequence data to be deleted in the current encoded sequence data to obtain the remaining encoded sequence data; splice the remaining encoded sequence data with the initial encoded sequence data to obtain the next encoded sequence data;
[0068] Input the next encoded sequence data into the audio signal recognition network to obtain the next audio recognition result, and repeat the above steps until all audio signals are recognized; wherein, based on all audio recognition results, the real-time recognition processing for all audio signals is completed.
[0069] Specifically, after the recognition of the current audio signal is completed, the recognition result for the current audio signal is output, and the recognition of the next audio signal is started. The recognition method in this application is different from the prior art in which two audio signals are recognized independently. Instead, the sequence data to be deleted in the current encoded sequence data is deleted. The length of the sequence data to be deleted is determined according to the initial encoded sequence data corresponding to the next audio signal, and the length of the sequence data to be deleted is equal to the length of the initial encoded sequence data. Specifically, taking the first data of the current encoded sequence data as a reference, a part at the beginning of the current encoded sequence data can be deleted, or taking the last data of the current encoded sequence data as a reference, a part at the end of the current encoded sequence data can be deleted, or relevant technicians can determine which data is the sequence data to be deleted. After deletion, the remaining encoded sequence data and the initial encoded sequence data are spliced together to obtain the next encoded sequence data. With the above method, on the one hand, real-time processing of multiple acquired audio signals is achieved, improving the efficiency of audio recognition; on the other hand, the lengths of the encoded sequence data input to the audio signal recognition network are equal, so as to better adapt to the parameter configuration set in the recognition network, thereby efficiently obtaining more accurate recognition results.
[0070] In one embodiment, obtaining an audio signal recognition network includes:
[0071] Obtain a preset encoded sequence training set, and the encoded sequence training set carries audio recognition labels;
[0072] Input the encoded sequence training set into the initial audio signal recognition network for training to obtain a training audio signal prediction result. Calculate the loss function result according to the training audio signal prediction result and the audio recognition label, and backpropagate the gradient of the loss function result to the initial audio signal recognition network for iterative training to generate a trained complete audio signal recognition network.
[0073] Specifically, Figure 3 For the structural schematic diagram of the above audio signal recognition network, multiple encoded sequence data are obtained from the input layer, where X1, X2,..., X n That is, it represents multiple encoded sequence data, where the hidden layer includes the above-mentioned pulse-coded neurons. The pulse-coded neurons determine whether the state of the neuron is an activated state or a non-activated state based on the multiple input encoded sequence data. In the figure, the output layer is used to output the final classification result. For example, Y1 can be the probability indicating that the current audio signal is environmental noise, and Y m Can be the probability indicating that the current audio signal is the target sound. Further, in this application, the loss function for the initial audio signal recognition network can be a mean square error loss function, a Smooth L1 loss function, etc.
[0074] By training the preset audio signal recognition network through the above method, a neural network more suitable for the current application environment can be obtained. Further, the above parameter configuration can be adjusted according to the actual situation so that the training method can be applicable to more environments.
[0075] In one embodiment, inputting the current audio signal into the trained audio signal recognition network to obtain an identification result corresponding to the current audio signal, including:
[0076] Through the current pulse-coded neuron, at least one current pulse result output by the historical pulse-coded neuron is obtained; wherein, the current pulse result is obtained based on the current coding data in the current coding sequence data;
[0077] Based on the current pulse-coded neuron, based on the current pulse result, the saved historical pulse result, and the preset pulse weight for the historical pulse-coded neuron, the next pulse result for the current coding data is obtained, and the next pulse result is sent to the next pulse-coded neuron;
[0078] Repeat the above steps until all pulse-coded neurons are traversed to obtain the identification result for the current audio signal.
[0079] Specifically, the above pulse-coded neuron can be an LIF model (leaky integrate-and-fire), a QIF model (Quadratic Integrate-and-Fire), etc. Taking the LIF model as an example, the LIF model parameters include the weight W ij , the threshold potential V th , the membrane integration delay time constant τ, and the potential update formula of its neuron i is as follows:
[0080]
[0081] In the case where the resting potential Vres is 0, the potential update formula can be simplified to:
[0082]
[0083] Where β is:
[0084]
[0085] It represents that the voltage decays with time.
[0086] Where the definition of the function K(t) is:
[0087]
[0088] It represents that the SNN information transmission is binary, being 1 when there is a pulse at time t and 0 when there is no pulse. The above weight W ij is the pulse weight in the above text. This weight value can be set by the user for each neuron. Adjusting different pulse weights can obtain more accurate recognition results. Through the above method, based on the currently input encoded data, the activation of each neuron is realized. According to the activation result, it is determined whether the current audio signal is the target audio signal, realizing that the potential contains the information memory of past excitations and is associated with historical data. Further, when existing neural networks calculate, the data is usually floating-point numbers, while the data in this application is binary pulses, enhancing the sparsity in the network.
[0089] In one embodiment, sending the next pulse result to the next pulse-encoding neuron includes:
[0090] When it is detected that the current pulse-encoding neuron is in the integration mode and the current pulse-encoding neuron detects that the next pulse result is greater than or equal to a preset pulse threshold, switch the integration mode to the pulse mode, and when the current pulse-encoding neuron is in the pulse mode, send the next pulse result to the next pulse-encoding neuron;
[0091] When it is detected that the current pulse-encoding neuron is in the integration mode and the next pulse result is less than the pulse threshold, accumulate according to the current pulse result and the next pulse result in the integration mode to obtain a new pulse result, and when switching to the pulse mode, send the new pulse result to the next pulse-encoding neuron.
[0092] Specifically, Figure 4 is a schematic flowchart of the calculation of the pulse-encoding neuron in this embodiment for the input encoded sequence data. Among them, X1 and X j etc. are the pulse signals output by the previous pulse-encoding neuron, St-1 is the above historical pulse result, St is the above next pulse result, compare the next pulse result based on a preset threshold. If the next pulse result is greater than or equal to the preset pulse threshold, send the next pulse result to the next pulse-encoding neuron. At this time, Yi is the pulse result output after the threshold comparison. When Vi > Vth, neuron i generates a pulse output and transmits it to the next neuron, and at the same time Vi is reset to 0. When Vi < Vth, neuron i is still in the integration mode, waiting to update Vi in the next time step. The above Vth is the pulse threshold in the above text. If the pulse result is less than the pulse threshold, the integration mode remains unchanged until the accumulated pulse result based on the transmitted pulse result is greater than or equal to the pulse threshold, and then switch from the integration mode to the pulse mode.
[0093] In one embodiment, obtaining an identification result corresponding to a current audio signal based on the activated pulse-coded neurons includes:
[0094] Obtaining a preset class label; wherein the class label corresponds to an output unit in the audio signal recognition network;
[0095] Based on the activated pulse-coded neurons, obtaining a pulse activation result, and obtaining an output result generated by the output unit based on the number of pulse activation results;
[0096] If the output result is greater than or equal to a preset activation threshold, obtaining an identification result based on the class label corresponding to the output unit.
[0097] Specifically, the above class label can be defined by the user, such as environmental noise, target sound, etc., and the class label and the output unit are in one-to-one correspondence. There is at least one of the above class units, that is, for the detection and identification of a target sound. During the training of the neural network, the output of the sample of the target sound will be encoded as pulses, and the output of the sample of the non-target sound will be encoded as 0 pulses. The output unit obtains at least one type of output result based on the pulse activation result of the pulse-coded neurons, and this output result represents the number of pulse activation results output by the pulse-coded neurons. If the output result corresponding to the target sound is greater than or equal to the preset activation threshold, it means that the input audio signal is the target sound. When there are at least two output units, it can be judged which output unit has the most corresponding output result, and its corresponding class label is the identification result of the current audio signal. Through the above method, the classification and identification of the input audio signal can be conveniently completed. Since it is a counting calculation of the pulse activation result, the computing resources consumed are less, and the identification result can be quickly obtained in practical applications.
[0098] In one embodiment, encoding the current spectral coefficient to obtain current encoded sequence data corresponding to the current spectral coefficient includes:
[0099] Obtaining a current spectral pixel value corresponding to the current spectral coefficient;
[0100] Encoding the current spectral pixel value based on a preset encoding rule to obtain current encoded sequence data.
[0101] Specifically, in this embodiment, taking Poisson encoding as an example, calculating the spectral pixel value corresponding to the spectral coefficient, and in practical applications, encoding methods such as relative order encoding and first trigger pulse encoding can be selected.
[0102] Calculating the Poisson sequence through the following formula:
[0103]
[0104] Wherein, λ is the intensity or rate in the Poisson coding process; k represents the number of occurrences of an event.
[0105] The probability of generating a pulse at each time step is proportional to the pixel value. Under a certain probability, the number of pulses generated by a certain pixel within a period of time conforms to the Poisson distribution. For example, the probability of generating a pulse spike corresponding to a spectral coefficient with a pixel value of 255 is 1 / 2, and its Poisson coding can be 01011001. For example, the probability of generating a pulse spike corresponding to a spectral coefficient with a pixel value of 128 is 1 / 4, and its Poisson coding can be 00001001, etc. In this application, the obtained MFCC spectrogram is regarded as a two-dimensional image containing pixel value information, and it is encoded for subsequent calculation of the encoded result. Further, the pixel value is converted into a coding sequence composed of 0s and 1s, which greatly reduces the computing resources required for the neural network to recognize the MFCC spectrogram.
[0106] This embodiment also provides a preferred embodiment of an audio signal recognition method, as Figure 5 shown, Figure 5 is a schematic flowchart of an audio signal recognition method in a preferred embodiment.
[0107] First, obtain the current audio signal and convert the current audio signal into a spectrogram reflecting the frequency distribution characteristics of the audio signal. Taking the MFCC coefficient map as an example, based on the pre-emphasis processing of the above current audio signal, where the pre-emphasis processing can be to let the audio signal pass through a high-pass filter; then the pre-emphasized audio signal is segmented into frames, and each frame is windowed, such as a Hamming window, a rectangular window, etc. The energy distribution of each frame in the frequency spectrum is obtained by performing a fast Fourier transform on the windowed data of each frame. Then, the obtained frequency-domain data is band-pass filtered through a Mel filter, and the logarithmic energy is first calculated for the filtered data, and a two-dimensional MFCC coefficient map is obtained through a discrete cosine transform DCT, where the MFCC coefficient map is composed of the above-mentioned multiple spectral coefficients.
[0108] Secondly, after obtaining multiple spectral coefficients, the spectral coefficients are encoded according to the corresponding pixel values to obtain encoded sequence data corresponding to each pixel point one by one. Then, the encoded sequence data is input into a preset and well-trained audio signal recognition network. The output unit of the audio signal recognition network outputs the number of pulses generated by each category. If the number of pulses output for the label of the target sound is greater than the threshold, it is determined as the target sound; otherwise, it is determined as a non-spark sound. The audio signal recognition network can be a network such as DNN or CNN. Preferably, in this embodiment, DNN is selected as the architecture basis, and then the computing units in the DNN are replaced with pulse-coded neurons. Among them, the training of the audio signal recognition network uses multiple encoded sequence data as the training set for supervised training.
[0109] Finally, after the current audio signal is recognized, the next audio signal is obtained. It can be selected to directly delete the current audio signal corresponding to the length of the next audio signal, and splice the remaining part after the current audio signal is deleted with the next audio signal, and then perform spectrum conversion and encoding uniformly. It is also possible to first encode the next audio signal to obtain initial encoded sequence data, then delete the current encoded sequence data corresponding to the current audio signal according to the length of the initial encoded sequence data to obtain remaining encoded sequence data, and then splice the remaining encoded sequence data with the initial encoded sequence data to obtain the next encoded sequence data, and input the next encoded sequence data into the audio signal recognition network to obtain the recognition result for the next audio signal, thereby completing the real-time recognition process of multiple audio signals.
[0110] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0111] Based on the same inventive concept, an embodiment of the present application further provides an audio signal recognition device for implementing the above-mentioned audio signal recognition method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the audio signal recognition device provided below can refer to the limitations on the audio signal recognition method in the above text, and will not be repeated here.
[0112] In one embodiment, as Figure 6 shown, an audio signal recognition device is provided, including: an acquisition module 61, a calculation module 62, and a generation module 63, where:
[0113] The acquisition module 61 is configured to acquire a current audio signal;
[0114] The calculation module 62 is configured to perform feature extraction processing on the current audio signal to obtain at least two current spectrum coefficients, and perform encoding processing on the current spectrum coefficients to obtain current encoded sequence data corresponding to the current spectrum coefficients;
[0115] A generation module 63 is configured to input all the encoded sequence data into a trained audio signal recognition network, activate the pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain a recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
[0116] Specifically, the acquisition module 61 acquires the current audio signal, which can be a segment of audio input by the user or the recording result of the acquisition module 61 for the environment. After the acquisition module 61 acquires the current audio signal, it sends the current audio signal to the calculation module 62. The calculation module 62 performs feature extraction processing on the current audio signal, converts the current audio signal into a spectrogram that can express its frequency distribution characteristics, and the spectrogram is composed of the above-mentioned multiple current spectral coefficients. The calculation module 62 is also configured to encode the multiple spectral coefficients to obtain the current encoded sequence data corresponding to the spectral coefficients one by one, where the encoding process can be Poisson encoding, phase encoding, etc. The calculation module 62 then inputs the current encoded sequence data into the generation module 63. The generation module 63 performs recognition processing on all the current encoded sequence data based on a preset trained audio signal recognition network to obtain a recognition result for the current audio signal. Among them, the audio signal recognition network is composed of pulse-coded neurons, and the pulse-coded neurons update the potential of the neurons based on the pulse information transmitted by all the current encoded sequence data. When the potential of the neuron is greater than the preset threshold potential, the current neuron sends a pulse output to the next neuron to realize the transmission of the information of the input current encoded sequence data. Finally, the generation module 63 obtains the recognition result corresponding to the current audio signal according to the number of pulses output by the output unit of the audio signal recognition network.
[0117] With the above device, on the one hand, the originally acquired current audio signal is converted into a two-dimensional spectrogram composed of current spectral coefficients, and further the pixel points therein are encoded to obtain the current encoded sequence data. Compared with directly using the audio signal for recognition, recognizing the current encoded sequence data can save a large amount of computing resources, and the spectral coefficients highlight the frequency distribution characteristics of the audio signal, making the final calculation result more accurate. On the other hand, the audio signal recognition network in this application is composed of multiple pulse-coded neurons, and the calculation data of the neurons is binary pulses, which greatly enhances the sparsity in the network. Moreover, the neuron model is closer to the biological model, consumes less computing power, and transmits information more efficiently.
[0118] Each module in the above audio signal recognition device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0119] In one embodiment, a range hood and stove system is provided, as Figure 7 shown. The range hood and stove system includes a range hood 71 and a stove 74;
[0120] The range hood 71 is connected to the stove 74 and is configured to obtain a current audio signal and perform any one of the audio signal recognition methods described above based on the current audio signal.
[0121] Specifically, in this embodiment, the target sound included in the current audio signal is the ignition sound generated by the stove 74. In practical applications, a control device and the above recognition device are also integrated on the range hood 71. The control device controls the operation of the range hood 71 in response to the recognition result corresponding to the current audio signal obtained by the recognition device. Among them, the recognition device calculates the recognition result of the audio signal based on the ignition sound generated by the stove 74.
[0122] In this embodiment, the recognition device obtains a spectrogram corresponding to the current audio signal, which contains multiple spectral coefficients, and performs encoding processing on all spectral coefficients to obtain multiple encoded sequence data. All the encoded sequence data are arranged into a string of time series encoded data and input into the trained audio signal recognition network. The pulse coding neurons in the recognition network update the potential based on the pulse information carried in the time series encoded data, thereby also realizing the memory that the potential of the neuron contains the information of past excitations. While ensuring a certain computational efficiency, the reliability of the judgment result is improved. Further, in this embodiment, the recognition device can obtain multiple segments of audio signals in real time and splice the multiple segments of audio signals according to user configuration to complete the streaming voice processing of the audio signals.
[0123] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data for audio signal recognition. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an audio signal recognition method.
[0124] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0125] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0126] The above-described embodiments only represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. An audio signal recognition method, characterized in that, The method includes: Obtaining a current audio signal; Performing feature extraction processing on the current audio signal to obtain at least two current spectrum coefficients, and performing encoding processing on the current spectrum coefficients to obtain current encoded sequence data corresponding to the current spectrum coefficients; Inputting all the encoded sequence data into a trained audio signal recognition network, activating pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtaining a recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
2. The method according to claim 1, wherein After obtaining the recognition result corresponding to the current audio signal, the method further includes: Obtaining the next audio signal; performing feature extraction processing on the next audio signal to obtain the next spectrum coefficients, and performing encoding processing on the next spectrum coefficients to obtain initial encoded sequence data; Determining and deleting sequence data to be deleted in the current encoded sequence data based on the length of the initial encoded sequence data to obtain remaining encoded sequence data; splicing the remaining encoded sequence data and the initial encoded sequence data to obtain the next encoded sequence data; Inputting the next encoded sequence data into the audio signal recognition network to obtain the next audio recognition result, and repeating the above steps until all audio signals are recognized; wherein, based on all the audio recognition results, real-time recognition processing for all the audio signals is completed.
3. The method according to claim 1, wherein Obtaining the audio signal recognition network includes: Obtaining a preset encoded sequence training set, where the encoded sequence training set carries audio recognition labels; Inputting the encoded sequence training set into an initial audio signal recognition network for training to obtain a training audio signal prediction result, calculating a loss function result based on the training audio signal prediction result and the audio recognition labels, and backpropagating the gradient of the loss function result to the initial audio signal recognition network for iterative training to generate the trained audio signal recognition network.
4. The method according to claim 1, characterized in that The obtaining a recognition result corresponding to the current audio signal based on the activated pulse-coded neurons includes: Obtaining at least one current pulse result output by a historical pulse-coded neuron through a current pulse-coded neuron; wherein, the current pulse result is obtained based on current encoded data in the current encoded sequence data; Based on the current pulse-coded neuron, obtaining a next pulse result for the current encoded data based on the current pulse result, saved historical pulse results, and preset pulse weights for the historical pulse-coded neuron, and sending the next pulse result to the next pulse-coded neuron; Repeating the above steps until all pulse-coded neurons are traversed to obtain the recognition result for the current audio signal.
5. The method according to claim 4, characterized in that The sending the next pulse result to the next pulse-coded neuron includes: When it is detected that the current pulse-coded neuron is in the integration mode and the next pulse result detected by the current pulse-coded neuron is greater than or equal to a preset pulse threshold, the integration mode is switched to the pulse mode, and when the current pulse-coded neuron is in the pulse mode, the next pulse result is sent to the next pulse-coded neuron; When it is detected that the current pulse-coded neuron is in the integration mode and the next pulse result is less than the pulse threshold, accumulation is performed according to the current pulse result and the next pulse result in the integration mode to obtain a new pulse result, and when switching to the pulse mode, the new pulse result is sent to the next pulse-coded neuron.
6. The method according to claim 1, wherein The obtaining the recognition result corresponding to the current audio signal based on the activated pulse-coded neurons includes: Obtaining a preset class label; wherein, the class label corresponds to an output unit in the audio signal recognition network; Based on the activated pulse-coded neurons, obtaining a pulse activation result, and obtaining an output result generated by the output unit based on the number of the pulse activation results; If the output result is greater than or equal to a preset activation threshold, the recognition result is obtained based on the class label corresponding to the output unit.
7. The method according to claim 1, wherein The encoding the current spectral coefficient to obtain the current encoded sequence data corresponding to the current spectral coefficient includes: Obtaining the current spectral pixel value corresponding to the current spectral coefficient; Encoding the current spectral pixel value based on a preset encoding rule to obtain the current encoded sequence data.
8. An audio signal recognition device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a current audio signal; A calculating module, configured to perform feature extraction processing on the current audio signal to obtain at least two current spectral coefficients, and encode the current spectral coefficients to obtain the current encoded sequence data corresponding to the current spectral coefficients; A generating module, configured to input all the encoded sequence data into a trained audio signal recognition network, activate the pulse-coded neurons in the audio signal recognition network based on the encoded sequence data, and obtain the recognition result corresponding to the current audio signal based on the activated pulse-coded neurons.
9. A smoke and stove system, characterized in that, The range hood and stove system includes a range hood and a stove; The range hood, connected to the stove, is configured to obtain a current audio signal and execute the audio signal recognition method according to any one of claims 1 to 7 based on the current audio signal.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.