Audio recognition method, audio recognition model training method, device, and electronic device
By using multiple parameter sets and multiple feature extraction subnets to feature extraction and recognition of audio data, the problem of insufficient audio feature extraction in the prior art is solved, and higher audio recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111213690.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The prior art is difficult to effectively extract audio features, resulting in insufficient accuracy of audio recognition.
N parameter sets are used to extract features of the audio data to be recognized separately, and the combined models of M feature extraction subnets and classification subnets are used to train the audio data recognition model.
Through the combination of multi-parameter set and multi-feature extraction subnetwork, audio features can be extracted more comprehensively, improving the accuracy and effect of audio recognition.
Smart Images

Figure CN113851147B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, in particular to the field of audio processing and deep learning technology, and specifically to an audio data recognition method, a training method for an audio data recognition model, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In many scenarios, audio needs to be recognized and classified, such as classifying the audio source, detecting whether the audio is offensive, and comparing whether the audio features match the expected features. In the process of audio classification, the audio features need to be extracted. A method that can effectively extract audio features and then accurately recognize the audio is desired.
[0003] The methods described in this section are not necessarily methods that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any method described in this section is considered to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the invention
[0004] The present disclosure provides an audio data recognition method, an audio data recognition model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to one aspect of the present disclosure, there is provided an audio data recognition method, comprising: acquiring audio data to be recognized; performing feature extraction on the audio data to be recognized using N parameter sets to obtain N feature data of the audio data to be recognized, wherein each parameter set in the N parameter sets is respectively associated with a different frequency range, and N is a positive integer greater than 1; and classifying the audio data to be recognized based on the N feature data.
[0006] According to another aspect of the present disclosure, a training method for an audio data recognition model is provided, the audio data recognition model comprising M feature extraction subnetworks and a classification subnetwork connected to an output end of each of the M feature extraction subnetworks, M being a positive integer greater than 1, the method comprising: obtaining sample audio data and a true label of the sample audio data; inputting the sample audio data into each of the M feature extraction subnetworks to obtain M feature data for the sample audio data; inputting the M feature data into the classification subnetwork to obtain a predicted label for the sample audio data; calculating a loss function based on the true label and the predicted label; and adjusting parameters of the audio data recognition model based on the loss function.
[0007] According to another aspect of the present disclosure, an audio data recognition device is provided, including: an audio data acquisition unit, used to acquire audio data to be recognized; a feature extraction unit, used to use N parameter sets to perform feature extraction on the audio data to be recognized, so as to obtain N feature data of the audio data to be recognized, wherein each parameter set in the N parameter sets is respectively associated with a different frequency range, and N is a positive integer greater than 1; and a classification unit, used to classify the audio data to be recognized based on the N feature data.
[0008] According to another aspect of the present disclosure, a training device for an audio data recognition model is provided, the audio data recognition model comprising M feature extraction subnetworks and a classification subnetwork connected to an output end of each of the M feature extraction subnetworks, M being a positive integer greater than 1, the training device comprising: a sample acquisition unit for acquiring sample audio data and a true label of the sample audio data; a feature extraction unit for inputting the sample audio data into each of the M feature extraction subnetworks to obtain M feature data for the sample audio data; a classification unit for inputting the M feature data into the classification subnetwork to obtain a predicted label for the sample audio data; a loss function calculation unit for calculating a loss function based on the true label and the predicted label; and a parameter adjustment unit for adjusting the parameters of the audio data recognition model based on the loss function.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the audio data recognition method or the audio data recognition model training method according to one or more embodiments of the present disclosure.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the audio data recognition method or the audio data recognition model training method according to one or more embodiments of the present disclosure.
[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements the audio data recognition method or the audio data recognition model training method according to one or more embodiments of the present disclosure.
[0012] According to one or more embodiments of the present disclosure, better audio feature extraction can be achieved, thereby achieving better audio recognition effect.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A schematic diagram showing an exemplary system in which the various methods described herein may be implemented according to an embodiment of the present disclosure;
[0016] Figure 2 A flowchart of an audio data recognition method according to an embodiment of the present disclosure is shown;
[0017] Figure 3 A flowchart of a method for training an audio data recognition model according to an embodiment of the present disclosure is shown;
[0018] Figure 4A A schematic diagram showing an audio data recognition model to which the method according to an embodiment of the present disclosure may be applied;
[0019] Figure 4B Another schematic diagram showing an audio data recognition model to which the method according to an embodiment of the present disclosure may be applied;
[0020] Figure 5 A structural block diagram of an audio data recognition device according to an embodiment of the present disclosure is shown;
[0021] Figure 6 A structural block diagram of a training device for an audio data recognition model according to an embodiment of the present disclosure is shown;
[0022] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0024] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, timing relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.
[0025] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.
[0026] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 FIG. 1 is a schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0028] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable execution of an audio data recognition method or a training method for an audio data recognition model.
[0029] In some embodiments, server 120 may also provide other services or software applications that may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0030] exist Figure 1In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0031] A user may use client devices 101, 102, 103, 104, 105, and / or 106 to recognize audio data, train an audio data recognition model, input audio, interact with audio recognition results, etc. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure may support any number of client devices.
[0032] Client devices 101, 102, 103, 104, 105 and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablet computers, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Game systems may include various handheld game devices, Internet-enabled game devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0033] The network 110 may be any type of network known to those skilled in the art that may support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0034] Server 120 may include one or more general purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0035] The computing units in the server 120 may run one or more operating systems including any of the above operating systems and any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0036] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0037] In some embodiments, the server 120 may be a server of a distributed system, or a server combined with a blockchain. The server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and virtual private servers (VPS) services.
[0038] The system 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In some embodiments, the databases used by the server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.
[0039] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.
[0040] Figure 1 The system 100 may be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.
[0041] Reference below Figure 2 An audio data recognition method 200 according to an exemplary embodiment of the present disclosure is described.
[0042] In step S201, audio data to be recognized is obtained.
[0043] In step S202, N parameter sets are used to perform feature extraction on the audio data to be identified to obtain N feature data of the audio data to be identified, wherein each parameter set in the N parameter sets is associated with a different frequency range, and N is a positive integer greater than 1.
[0044] In step S203, the audio data to be recognized is classified based on the N feature data.
[0045] According to the method described in the embodiment of the present disclosure, audio data can be recognized more accurately. Specifically, by using at least two different parameter sets to perform feature extraction on audio data respectively, features associated with different frequency ranges in the audio data can be better covered, better audio feature extraction can be achieved, and thus better audio recognition effect can be achieved.
[0046] The audio data to be recognized can be the original time domain waveform of the audio or a fragment of the time domain waveform. Extracting audio features from the time domain waveform can reduce information loss and improve audio recognition effect.
[0047] According to some embodiments, obtaining the audio data to be recognized may include: obtaining original audio data, and in response to determining that the time length of the original audio data is greater than a length threshold, intercepting the original audio data using the length threshold to obtain the audio data to be recognized, wherein the time length of the audio data to be recognized is equal to the length threshold. Thus, the overlong audio can be intercepted to obtain input data with a regular time length.
[0048] As an example, the length threshold may be 0.5s, 1s, 4s, 10s, etc., but those skilled in the art will appreciate that the present disclosure is not limited thereto, and other time lengths adapted to audio features may also be applicable to the method described in the present disclosure.
[0049] As an example, the length of the cut audio can be determined according to the task requirements and the "balance between the computing speed and the amount of content included". For example, if it is too long, the computing speed will be very slow, and if it is too short, the content contained in the audio may not be sufficient as a benchmark for classification. For example, in the case where the length of the audio to be recognized is often short, the real-time required for the recognition audio is very strong, and / or the data contained in the audio data is often obvious, a shorter length threshold such as 0.2s, 0.02s, 0.002s, etc. can be selected. For another example, in the case where the audio to be recognized generally has a longer length, requires lower real-time requirements, or the data contained in the audio data is often not very clear, so a longer audio is required for analysis, a longer length threshold such as 20s, 60s, 120s, etc. can be selected. As a specific non-limiting example, in the case where the audio recognition task is to prevent recording attacks, audio shorter than 2s may not be enough to contain human voices or attack features, so a time length of 3 seconds to 5 seconds (e.g., 4 seconds) can be selected, and it can be understood that the present disclosure is not limited to this.
[0050] According to some embodiments, obtaining the audio data to be recognized may include: obtaining the original audio data, and in response to determining that the time length of the original audio data is less than a length threshold, copying the original audio data until the time length of the copied original audio data is not less than the length threshold; and intercepting the copied original audio data using the length threshold to obtain the audio data to be recognized, wherein the time length of the audio data to be recognized is equal to the length threshold. Thus, audio that is too short can also be copied and intercepted to obtain input data with a regular time length.
[0051] According to some optional embodiments, feature extraction and classification of audio data can be achieved by an audio data recognition model. According to other embodiments, feature extraction and classification of audio data can be achieved by other feature extraction means and classification means known to those skilled in the art. It is to be understood that the present disclosure is not limited thereto.
[0052] Reference below Figure 3 A training method 300 for an audio data recognition model according to another embodiment of the present disclosure is described. The audio data recognition model may include M feature extraction subnetworks and a classification subnetwork connected to an output end of each of the M feature extraction subnetworks, where M is a positive integer.
[0053] In step 301, sample audio data and true labels of the sample audio data are obtained.
[0054] At step 302, the sample audio data is input into each feature extraction sub-network of the M feature extraction sub-networks to obtain M feature data for the sample audio data.
[0055] At step 303, M feature data are input into the classification sub-network to obtain a predicted label for the sample audio data.
[0056] At step 304, a loss function is calculated based on the true label and the predicted label.
[0057] At step 305 , parameters of the audio data recognition model are adjusted based on the loss function.
[0058] According to the method described in the embodiment of the present disclosure, audio data can be recognized more accurately. Specifically, by using at least two different classification subnetworks to perform feature extraction on audio data respectively, better audio feature extraction can be achieved, so that the model trained in this way can achieve better audio recognition effect.
[0059] According to some embodiments, each of the M feature extraction subnetworks is initialized based on a corresponding filter parameter set in the M filter parameter sets, each of the M filter parameter sets including an upper cutoff frequency and a lower cutoff frequency. In such an embodiment, a faster model learning process and a better convergence effect can be achieved by initializing the feature extraction subnetwork using the filter parameters.
[0060] Refer to the following Figure 4A 4 is a schematic diagram of a model 400 to which the method according to an embodiment of the present disclosure may be applied. Figure 4A It is shown in the figure that the model 400 may include M feature extraction subnetworks 410-1, 410-2, ... 410-M and a classification subnetwork 420. The model 400 may also include an optional feature integration unit for integrating the M feature data obtained from the M feature extraction subnetworks, but it can be understood that this is only an example, and the M feature data can be directly input into the input end of the classification subnetwork without an additional feature integration unit, or the classification subnetwork itself may have a unit or one or more layers (such as one or more residual blocks, etc.) for integrating features.
[0061] It is understandable that although Figure 4A It is shown that the model 400 includes at least 3 feature extraction subnetworks, but such a model may also include more or fewer feature extraction subnetworks. For example, the model 400 may include only one feature extraction subnetwork (M=1). In the following, for convenience, the feature extraction part described as M feature extraction subnetworks 410-1, 410-2...410-M may include only one or two subnetworks, or may include more (for example, dozens) of subnetworks. The selection of the M value will be described in more detail below in conjunction with specific embodiments, and the present disclosure is not limited thereto.
[0062] In signal and system theory, it is necessary to extract signal features by performing convolution operations in the time domain through filters. Here, the convolution layer of the neural network can be used to simulate the convolution of the filter, and the traditional filter parameters can be used to initialize the neural network parameters. In this way, the initial state of the neural network is equivalent to a better state that draws on the empirical formula of the filter, which can greatly reduce the learning cost. Specifically, the filter parameters are calculated using the preset upper and lower limit frequencies of the filter, for example, by substituting them into the filter formula for calculation, and such parameters are used as the initial parameters of the neural network, which can greatly utilize the advantages of the existing filter parameters in the traditional empirical formula or traditional theory, and obtain relatively good initial parameters for the model, so the training process is fast and a better convergence value can be obtained.
[0063] In the field of signal, system and audio processing, multiplication in the frequency domain corresponds to convolution in the time domain. Assuming that the function of the filter in the frequency domain is g[n,θ] in the time domain, for the input signal, the signal after x[n] is filtered by the filter is:
[0064] y[n]=x[n]*g[n,θ]
[0065] Where n represents the time of the signal, which is a normalized unitless time sequence number in the case of a discrete signal, * represents a convolution operation, and θ is a set of parameters of the generalized filter, which may include all possible parameters in the filter formula except n, and may include one or more or zero parameters.
[0066] The traditional features can be obtained by filtering using the filter g[n,θ], and then the extracted traditional features are sent to the back-end classifier. Furthermore, in order to utilize the powerful feature extraction capability of the convolutional neural network CNN and minimize information loss (for example, spectrum leakage), CNN can be used as the feature extraction subnetwork part, and feature extraction can be achieved by learning features with the help of CNN's powerful ability. In this process, g[n,θ] can be used as the initialization parameter of the CNN network. Continue to refer to Figure 4A , wherein the M feature extraction subnetworks 410-1, 410-2, ... 410-M may be M CNNs. The sample audio data may be time series data, and specifically, may be a time domain waveform of an audio or a discrete time domain signal, or other original or processed audio signals or audio data that can be understood by those skilled in the art.
[0067] The back-end classifier may use RNN or CNN or other networks. As an example, CNN may be used to extract features, and then a residual block module may be used to integrate the audio features, making the audio features more distinguishable, and finally the features integrated by the residual block module may be integrated into the RNN network, and the ability of RNN to model time series may be used to obtain audio level features. As a more specific non-limiting example, the back-end classifier may use a gated recurrent unit (GRU), but those skilled in the art will appreciate that the present disclosure is not limited thereto.
[0068] Refer to the following Figure 4B The working process of the audio recognition model is described in conjunction with a more specific non-limiting example model 4200. For audio with a sampling frequency of 16k, a maximum frequency of 8000Hz, and a threshold length set to 4s, there are a total of 16000*4=64000 sample points. In this case, the bandwidth of each corresponding filter network 4211, 4212...421M is 8000 / M(Hz). As an example, each filter network 4211, 4212...421M can be a CNN network with a dimension of (129,0,128).
[0069] A pooling layer 4220 may be arranged after the feature extraction subnetwork, such as a maximum pooling layer of batch normalization (Batch Norm) and leaky (Leaky) ReLU function to extract the largest feature point on each feature through pooling. For example, if the pooling coefficient is 3, the data dimension can be calculated as (64000-128) / 3=21900.
[0070] Residual blocks can also be set in the neural network to further reduce the data dimension. As an example, a first residual network 4230 can be set in the model, with a channel number of 128 and a stride of 1. Thus, the data dimension can be calculated as 21290 / 3=7096. Afterwards, the data can be subjected to feature map scaling. One possible implementation is to obtain the corresponding coefficient c for each feature s through a softmax function, and then calculate the updated feature s' through s'=c*s+c. Similarly, two first residual networks 4230 (not shown) can be set, and after passing through the second second residual network, the dimension can be further reduced to 7096 / 3=2365.
[0071] More residual blocks can be set to further reduce the dimension. As an example, Figure 4B As shown, four identical second residual networks 4240 are set, each with 512 channels and a step size of 1, and the dimension can be further reduced to 2365 / 3 / 4=29. The reduced-dimensional data can be input to a gated recurrent unit (GRU) 4250 to synthesize the information of the entire audio. Finally, it passes through a fully connected layer 4260 for linear transformation.
[0072] It is understandable that the above combination Figure 4B The number of modules, module names and division methods, and parameters of audio data described are all examples, and the present disclosure is not limited thereto.
[0073] According to some embodiments, the M filter parameter sets may be set by the following steps: obtaining a predetermined frequency range; dividing the predetermined frequency range into M continuous sub-bands; and setting the lower limit frequency and the upper limit frequency of each sub-band in the M continuous sub-bands to the lower limit cutoff frequency (denoted as f1) and the upper limit cutoff frequency (denoted as f2) in the corresponding filter parameter set. Thus, the network is initialized using filter parameters covering the entire frequency range, so that the trained neural network can better cover various features of the neural network.
[0074] For example, in the total frequency band range including the frequency range of 0-8k Hz, and a filtering scheme of M=8 filters is set, that is, there are 8 feature extraction subnetworks (e.g., 8 CNNs). In this case, as an exemplary frequency averaging scheme, the start and end points of each frequency band can be 0-1k, 1k-2k... and so on. In this case, a CNN can be set as a filter in each frequency band to extract features within the frequency band.
[0075] The number of filters or feature extraction subnetworks can be set based on experience, accuracy requirements, system computing power, etc. For example, a smaller M value (or M=1, that is, no frequency band division) will result in a larger frequency band interval and a simpler system architecture and computing speed; a larger M value (for example, a dozen, dozens or even hundreds) can result in a finer interval granularity and a more accurate effect; but in the case where the M value is too large, it may also result in a slower calculation speed, or because the information in each interval is too little to characterize the characteristics of the audio, resulting in the accuracy not continuing to increase with the increase of the M value. As an example, 10, 20 or 40 frequency bands can be set for an 8k sampling rate, and those skilled in the art will understand that the present disclosure is not limited to this.
[0076] It is understandable that the frequency band can be divided equally or unevenly, which may depend on different filter types or filter formulas. For example, for rectangular filter parameters, a frequency band equal division method can be adopted, but for other filter formulas (such as Mel filter), an uneven division scheme can also be adopted. It is understandable that the present disclosure is not limited to this.
[0077] The following uses a rectangular filter as an example to describe the method according to an optional embodiment of the present disclosure. It is understandable that such filter parameters are only examples, and other filtering formulas in the field of signal and system processing may also be applicable to the initialization and training of the audio recognition model according to the embodiment of the present disclosure. According to some embodiments, each filter parameter set in the M filter parameter sets may correspond to the parameters of a rectangular filter in the frequency domain, and wherein dividing the predetermined frequency range into M continuous sub-bands may include averaging the predetermined frequency range to obtain M sub-bands of equal width. For the rectangular filter, according to the empirical formula in the signal field, a frequency averaging scheme is applicable.
[0078] The divided frequency range is denoted as [f 1 ,f 2 ] Set the initialization parameters of the neural network for each frequency band. Adding a rectangular filter in the frequency domain corresponds to the Sinc function in the time domain. Here is an example of the filter parameters corresponding to the Sinc function in the time domain:
[0079] g[n,f 1 ,f 2 ]=2f 2 sinc(2πf 2 n)-2f 1 sinc(2πf 1 n)
[0080] Therefore, we can use the g[n,f 1 ,f 2] Initialize the weights of the feature extraction subnetwork. As an example, the first layer of the neural network can be initialized using the parameters in the above function, while the subsequent layers can be randomly initialized. Afterwards, such a model can be trained so that the convolution kernel corresponding to the filter can be continuously updated to learn the convolution kernel parameters suitable for the required task.
[0081] According to some embodiments, the real label of the sample audio data may include a label that the sample data is a real human voice or machine-generated audio. The model trained in this way can realize voice liveness recognition and prevent recording attacks. The recording anti-attack system needs to ensure that the audio received by the voiceprint system is real human audio to ensure the security of the voiceprint system. It can be understood that the application here is only an example, and such training methods and training models can be applied to other purposes such as speech recognition and classification.
[0082] Further, in order to reduce or prevent spectrum leakage to obtain a better feature extraction effect, according to some embodiments, each of the N filter parameter sets may correspond to a parameter set of a filter that has been windowed. Continuing with the example of the rectangular filter above, the filter parameter formula for initialization after windowing may be as follows:
[0083] g w [n,f 1 ,f 2 ]=g[n,f 1 ,f 2 ]·w[n]
[0084] Where w[n] is the window function. As an example, the Hamming window function shown below can be used:
[0085]
[0086] Among them, n represents time and L represents the length of the convolution kernel.
[0087] In such an embodiment, the g calculated thereby can be used w [n,f 1 ,f 2 ] Initialize the weights of the feature extraction subnetwork, for example, initialize the first layer of the neural network, etc. It can be understood that the above filter types, window function types and filter formulas are all examples, and those skilled in the art will understand that other parameter sets related to filtering and feature extraction can also be used to initialize the feature extraction subnetwork according to the embodiment of the present disclosure and obtain better results than random initialization.
[0088] According to some embodiments, obtaining sample audio data may include: obtaining original audio data; and in response to determining that the time length of the original audio data is greater than the sample length threshold, intercepting the original audio data using the sample length threshold to obtain at least one sample audio data, wherein the time length of each sample audio data in the at least one sample audio data is equal to the sample length threshold. According to some embodiments, obtaining sample audio data may include: obtaining original audio data; in response to determining that the time length of the original audio data is less than the sample length threshold, copying the original audio data until the time length of the copied original audio data is not less than the sample length threshold; and intercepting the copied original audio data using the sample length threshold to obtain sample audio data, wherein the time length of the sample audio data is equal to the sample length threshold. Thus, the sample can be processed to obtain a regular sample length. The input of a neural network is often fixed-length data, so in actual training and testing, data longer than the length needs to be intercepted, and data shorter than the length needs to be self-copied and intercepted. As described above, the sample length threshold can be determined based on the task requirements to be used by the model and the "balance between operation speed and content content". As an example, the length threshold can be 0.5s, 1s, 4s, 10s, etc.
[0089] As a specific non-limiting example, when the sample length threshold is 4 seconds, the method may include first segmenting the audio. If the audio length is greater than 4 seconds, then all the audio is cut into segments of 4 seconds each; if the audio is less than 4s, then copy the audio from the beginning, and after the copying is completed, extract 4s of audio from the copied audio.
[0090] According to one or more embodiments of the present disclosure, the model can be a binary classification model, trained using manually labeled true and false samples (for example, in the case of recording anti-attack or liveness recognition, samples labeled as attack audio and real human voices, respectively), and the loss function can be designed as a cross-entropy function.
[0091] According to one or more embodiments of the present disclosure, audio data can be first obtained, and the time-series audio data can be segmented and padded to a predetermined length; and the number of filters N can be set as needed, and the upper and lower cutoff frequencies f1 and f2 of each frequency band can be calculated according to the number of filters N and the sampling rate, and then the parameter set of each filter, that is, the initialization parameters of each feature extraction subnetwork, is calculated according to the upper and lower cutoffs f1 and f2. After that, each feature extraction subnetwork (such as CNN) is initialized using the parameter. After the model is built and initialized, the sample is used for training, and then the model can be tested. During the test, each audio can also be regularized to a predetermined length (for example, 4s length). After the training and testing are completed, the trained model can be used for judgment. For example, in the case of using the real label of the sample audio data, the real human voice or the label of the machine-generated audio, the trained model can judge the high probability as the real human voice and the low probability as the attack audio.
[0092] Reference now Figure 5 An audio data recognition device 500 according to an embodiment of the present disclosure is described. The audio data recognition device 500 may include an audio data acquisition unit 501, a feature extraction unit 502, and a classification unit 503. The audio data acquisition unit 501 is used to acquire audio data to be recognized. The feature extraction unit 502 is used to use N parameter sets to perform feature extraction on the audio data to be recognized, so as to obtain N feature data of the audio data to be recognized, wherein each parameter set in the N parameter sets is associated with a different frequency range, and N is a positive integer greater than 1. The classification unit 503 is used to classify the audio data to be recognized based on the N feature data.
[0093] According to the apparatus of the embodiment of the present disclosure, audio data can be recognized more accurately.
[0094] Reference now Figure 6A training device 600 for an audio data recognition model according to an embodiment of the present disclosure is described. The audio data recognition model may include M feature extraction subnetworks and a classification subnetwork connected to the output end of each feature extraction subnetwork of the M feature extraction subnetworks, where M is a positive integer greater than 1. The training device 600 for the audio data recognition model may include a sample acquisition unit 601, a feature extraction unit 602, a classification unit 603, a loss function calculation unit 604, and a parameter adjustment unit 605. The sample acquisition unit 601 is used to obtain sample audio data and the true label of the sample audio data. The feature extraction unit 602 is used to input the sample audio data into each feature extraction subnetwork of the M feature extraction subnetworks to obtain M feature data for the sample audio data. The classification unit 603 is used to input the M feature data into the classification subnetwork to obtain the predicted label of the sample audio data. The loss function calculation unit 604 is used to calculate the loss function based on the true label and the predicted label. The parameter adjustment unit 605 is used to adjust the parameters of the audio data recognition model based on the loss function. In such an embodiment, each of the M feature extraction subnetworks is initialized based on a corresponding filter parameter set in the M filter parameter sets, and each of the M filter parameter sets includes an upper cutoff frequency and a lower cutoff frequency.
[0095] According to the device described in the embodiment of the present disclosure, it is possible to achieve a faster model learning process and a better convergence effect by initializing the feature extraction subnetwork using filtering parameters.
[0096] In the technical solution disclosed herein, the collection, acquisition, storage, use, processing, transmission, provision and public application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0097] According to an embodiment of the present disclosure, an electronic device, a readable storage medium and a computer program product are also provided.
[0098] refer to Figure 7 , a block diagram of an electronic device 700 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0099] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0100] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 may be any type of device capable of inputting information to the device 700, the input unit 706 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device, and may include but is not limited to a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 707 may be any type of device capable of presenting information, and may include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 708 may include but is not limited to a disk, an optical disk. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a 1302.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0101] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as methods 200 and / or 300 and their variants, etc. For example, in some embodiments, methods 200 and / or 300 and their variants, etc. may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods 200 and / or 300 and their variants, etc. described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the methods 200 and / or 300 and their variations, etc., in any other appropriate manner (eg, by means of firmware).
[0102] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0103] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0104] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0106] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0107] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0108] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0109] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but only by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. An audio data recognition method, comprising: Acquire audio data to be identified, wherein the audio data to be identified includes a time domain waveform segment of a predetermined length; By inputting the time domain waveform segment of the predetermined length into corresponding feature extraction subnetworks corresponding to N parameter sets respectively, using the N parameter sets to extract features of the audio data to be identified respectively, so as to obtain N feature data of the audio data to be identified, wherein each parameter set in the N parameter sets is respectively associated with a different frequency range, N is a positive integer greater than 1, wherein each parameter set in the N parameter sets is respectively associated with a traditional filter parameter, wherein the N parameter sets respectively correspond to feature extraction subnetworks, each feature extraction subnetwork is initialized based on the corresponding traditional filter parameters, and each filter parameter set includes an upper cutoff frequency and a lower cutoff frequency; and The audio data to be identified is classified based on the N feature data.
2. The method according to claim 1, wherein: Obtaining audio data to be recognized includes: Get the original audio data; In response to determining that the time length of the original audio data is greater than a length threshold, the original audio data is intercepted based on the length threshold to obtain audio data to be identified, wherein the time length of the audio data to be identified is equal to the length threshold.
3. The method according to claim 1, wherein: Obtaining audio data to be recognized includes: Get the original audio data; In response to determining that the time length of the original audio data is less than a length threshold, copying the original audio data until the time length of the copied original audio data is not less than the length threshold; and The copied original audio data is cut based on the length threshold to obtain the audio data to be recognized, wherein the time length of the audio data to be recognized is equal to the length threshold.
4. A training method for an audio data recognition model, the audio data recognition model comprising M feature extraction subnetworks and a classification subnetwork connected to an output end of each of the M feature extraction subnetworks, M being a positive integer greater than 1, the method comprising: Acquire sample audio data and a true label of the sample audio data, wherein the sample audio data includes a time domain waveform segment of a predetermined length; Inputting the sample audio data as the time domain waveform segment of the predetermined length into each feature extraction subnetwork of the M feature extraction subnetworks to obtain M feature data for the sample audio data; Inputting the M feature data into the classification subnetwork to obtain a predicted label for the sample audio data; Calculating a loss function based on the true label and the predicted label; and Based on the loss function, adjusting the parameters of the audio data recognition model, Among them, each feature extraction subnetwork in the M feature extraction subnetworks is initialized based on a corresponding filter parameter set in the M filter parameter sets, each filter parameter set corresponds to a traditional filter parameter, and each filter parameter set in the M filter parameter sets includes an upper cutoff frequency and a lower cutoff frequency.
5. The method according to claim 4, wherein: The M filter parameter sets are set by the following steps: obtaining a predetermined frequency range; Dividing the predetermined frequency range into M consecutive sub-frequency bands; and The lower limit frequency and the upper limit frequency of each of the M consecutive sub-frequency bands are set to the upper cutoff frequency and the lower cutoff frequency in the corresponding filter parameter set.
6. The method according to claim 4, wherein: Each of the M filter parameter sets corresponds to a parameter set of a rectangular filter in the frequency domain.
7. The method according to claim 5, wherein: The dividing the predetermined frequency range into M consecutive sub-frequency bands comprises: The predetermined frequency range is evenly divided to obtain M sub-frequency bands with the same width.
8. The method according to any one of claims 4 to 7, wherein: Each of the M filter parameter sets corresponds to a parameter set of a filter that has been windowed.
9. The method according to any one of claims 4 to 7, wherein: Getting sample audio data includes: Obtaining raw audio data; and In response to determining that the time length of the original audio data is greater than a sample length threshold, the original audio data is truncated based on the sample length threshold to obtain at least one sample audio data, wherein the time length of each sample audio data in the at least one sample audio data is equal to the sample length threshold.
10. The method according to any one of claims 4 to 7, wherein: Getting sample audio data includes: Get the original audio data; In response to determining that the time length of the original audio data is less than a sample length threshold, copying the original audio data until the time length of the copied original audio data is not less than the sample length threshold; and The copied original audio data is truncated based on the sample length threshold to obtain sample audio data, wherein the time length of the sample audio data is equal to the sample length threshold.
11. The method according to any one of claims 4 to 7, wherein: The true label of the sample audio data includes a label that the sample audio data is a real human voice or machine-generated audio.
12. An audio data recognition device, comprising: An audio data acquisition unit, used to acquire audio data to be recognized, wherein the audio data to be recognized includes a time domain waveform segment of a predetermined length; a feature extraction unit, configured to extract features from the audio data to be identified by inputting the time domain waveform segment of the predetermined length into corresponding feature extraction subnetworks corresponding to N parameter sets, respectively, using the N parameter sets to obtain N feature data of the audio data to be identified, wherein each parameter set in the N parameter sets is associated with a different frequency range, respectively, and N is a positive integer greater than 1, wherein each parameter set in the N parameter sets is associated with a traditional filter parameter, respectively, wherein the N parameter sets correspond to feature extraction subnetworks, each feature extraction subnetwork is initialized based on corresponding traditional filter parameters, and each filter parameter set includes an upper cutoff frequency and a lower cutoff frequency; and A classification unit is used to classify the audio data to be identified based on the N feature data.
13. A training device for an audio data recognition model, the audio data recognition model comprising M feature extraction subnetworks and a classification subnetwork connected to an output end of each of the M feature extraction subnetworks, M being a positive integer greater than 1, the training device comprising: A sample acquisition unit, used to acquire sample audio data and a true label of the sample audio data, wherein the sample audio data includes a time domain waveform segment of a predetermined length; A feature extraction unit, configured to input the sample audio data as a time domain waveform segment of the predetermined length into each feature extraction subnetwork of the M feature extraction subnetworks to obtain M feature data for the sample audio data; A classification unit, used for inputting the M feature data into the classification subnetwork to obtain a predicted label of the sample audio data; A loss function calculation unit, configured to calculate a loss function based on the true label and the predicted label; and a parameter adjustment unit, configured to adjust the parameters of the audio data recognition model based on the loss function, Among them, each feature extraction subnetwork in the M feature extraction subnetworks is initialized based on a corresponding filter parameter set in the M filter parameter sets, each filter parameter set corresponds to a traditional filter parameter, and each filter parameter set in the M filter parameter sets includes an upper cutoff frequency and a lower cutoff frequency.
14. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3 or 4-11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-3 or 4-11.
16. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method of any one of claims 1 to 3 or 4 to 11 is implemented.
Citation Information
Patent Citations
Audio event detection method and apparatus, electronic device and storage medium
CN111899760A