Audio recognition method, device, electronic device and storage medium
By extracting features from audio frames and generating correlation matrices, the problem of low accuracy in existing audio recognition is solved, and higher-precision audio recognition is achieved.
Patent Information
- Application Number
- CN202411620332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-13
AI Technical Summary
In existing audio recognition technology, the method of extracting Mel spectrum features results in low recognition accuracy.
By extracting features from the audio frames to be identified, a first feature matrix is generated, and based on this matrix, second and third feature matrices representing the association relationship between audio frames are determined to enhance the feature representation of the audio frames, and the association relationship features between multiple audio frames are integrated for identification.
It improves the accuracy of audio recognition, enhances the feature expression ability of audio frames, and improves the accuracy of recognition results.
Smart Images

Figure CN119694337B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and specifically to an audio recognition method, device, electronic device, and storage medium. Background Art
[0002] Audio recognition (including speech recognition) refers to determining the recognition result of audio by analyzing audio signals. For example, the recognition result can be the user's emotional information in the audio, and can identify the emotions expressed in the audio signal. It has a wide range of applications in customer service systems, virtual assistants, sentiment analysis, medical diagnosis and other fields.
[0003] Currently, the general method is to extract features from audio signals, such as Mel spectrum feature extraction, to obtain the corresponding audio feature representation, and then perform audio recognition based on the audio feature representation to obtain a recognition result corresponding to the audio signal. However, this method results in low audio recognition accuracy. Summary of the Invention
[0004] The present application provides an audio recognition method, device, electronic device and storage medium, which improve the accuracy of audio recognition.
[0005] In a first aspect, the present application provides an audio recognition method, the method comprising:
[0006] Perform feature extraction on each audio frame in the audio to be recognized to obtain a first feature matrix;
[0007] Determining, based on the first feature matrix, a second feature matrix for characterizing a correlation relationship between any two audio frames in the plurality of audio frames;
[0008] Determining, based on the second feature matrix, a third feature matrix for characterizing an association relationship between each audio frame and the plurality of audio frames;
[0009] Recognition is performed based on the first feature matrix and the third feature matrix to obtain a recognition result corresponding to the audio to be recognized.
[0010] In a second aspect, the present application provides an audio recognition device, the device comprising: an acquisition unit and a processing unit;
[0011] An acquisition unit, configured to acquire audio to be recognized;
[0012] The processing unit is configured to extract features from each audio frame in the audio to be recognized to obtain a first feature matrix; determine, based on the first feature matrix, a second feature matrix for characterizing the association relationship between any two audio frames in the plurality of audio frames; determine, based on the second feature matrix, a third feature matrix for characterizing the association relationship between each audio frame and the plurality of audio frames; and perform recognition based on the first feature matrix and the third feature matrix to obtain a recognition result corresponding to the audio to be recognized.
[0013] In a third aspect, the present application provides an electronic device comprising: a processor and a memory, the processor being connected to the memory, the memory being used to store computer programs, and the processor being used to execute the computer programs stored in the memory, so that the electronic device performs the method of the first aspect.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method of the first aspect is performed.
[0015] In a fifth aspect, the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, performs the method of the first aspect.
[0016] The implementation of this application has the following beneficial effects:
[0017] After extracting features from multiple audio frames in the audio to be recognized to obtain a first feature matrix, recognition is not performed directly based on the first feature matrix. Instead, a second feature matrix representing the association relationship between any two audio frames in the multiple audio frames is first determined based on the first feature matrix. That is, each element in the second feature matrix corresponds to the association relationship between the two audio frames, that is, the association relationship between audio frames or time nodes can be focused on. Then, a third feature matrix for representing the association relationship between each audio frame and the multiple audio frames is determined based on the second feature matrix. That is, each element in the third feature matrix represents the overall association relationship between each audio frame and the multiple audio frames of the audio to be recognized, which can reflect a feature representation of each audio frame in the entire audio to be recognized. That is, each audio frame aggregates the associations with multiple audio frames, so that the feature representation corresponding to each audio frame finally integrates the associations with multiple audio frames, enriching the feature representation of each audio frame and enhancing the feature expression capability of each audio frame. Then, recognition is performed based on the feature representation of each audio frame in the first feature matrix and the feature representation of the overall association relationship between each audio frame in the third feature matrix, which can improve the accuracy of audio recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of an audio recognition method provided in an embodiment of the present application;
[0020] Figure 2 A flowchart of a method for training an audio recognition model provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the structure of an audio recognition model provided in an embodiment of the present application;
[0022] Figure 4 A schematic diagram of the structure of another audio recognition model provided in an embodiment of the present application;
[0023] Figure 5 A schematic diagram of the structure of another audio recognition model provided in an embodiment of the present application;
[0024] Figure 6 A schematic diagram of the structure of another audio recognition model provided in an embodiment of the present application;
[0025] Figure 7 A schematic diagram of an audio recognition system provided in an embodiment of the present application;
[0026] Figure 8 A schematic diagram of a system for training an audio recognition model provided in an embodiment of the present application;
[0027] Figure 9 A block diagram of the functional units of an audio recognition device provided in an embodiment of the present application;
[0028] Figure 10 A block diagram of the functional units of a training device for an audio recognition model provided in an embodiment of the present application;
[0029] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] The terms "first," "second," "third," and "fourth," etc., in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a particular order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0032] References herein to "embodiments" mean that a particular feature, result, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0033] See Figure 1 , Figure 1 This is a flow chart of an audio recognition method provided in an embodiment of the present application. The method is applied to an audio recognition device and includes but is not limited to steps S101-S104:
[0034] S101 : Extract features from multiple audio frames in the audio to be recognized to obtain a first feature matrix.
[0035] In an embodiment of the present application, Mel-Frequency Cepstral Coefficient (MFCC) can be used to extract features of the audio to be identified, and the extracted Mel spectrum is determined as a first feature matrix. At this time, the first feature matrix includes a feature vector of each audio frame in multiple audio frames of the audio to be identified. Of course, other feature extraction methods can also be used, which is not limited in this application.
[0036] In an optional embodiment, feature extraction can also be performed on each audio frame in the audio to be identified (for example, using Mel spectrum feature extraction) to obtain an initial feature matrix, that is, the initial feature matrix in this case includes the feature vector corresponding to each audio frame of the audio to be identified; then, based on the initial feature matrix, multiple fourth feature matrices corresponding to multiple channels are determined, for example: a channel dimension is added to the feature map corresponding to the initial feature matrix to obtain a first feature map, in which case the first feature map is a single-channel feature map, and then the channel dimension in the first feature map is channel-amplified, for example, a two-dimensional convolution is performed on the first feature map to obtain a second feature map, in which case the second feature map is a multi-channel feature map, and then the second feature map is separated according to the channel to obtain multiple third feature maps corresponding to multiple channels, and then multiple fourth feature matrices corresponding to multiple channels are determined based on the multiple third feature maps. It should be noted that by amplifying the channel feature map through two-dimensional convolution, that is, by generating a multi-channel feature map, high-level time-frequency features can be effectively extracted and combined, which not only captures local features, but also learns more complex and richer high-level feature representations through deep networks, thereby improving the ability to process time series and frequency domain data.
[0037] Then, based on multiple fourth characteristic matrices corresponding to multiple channels, the first characteristic matrix is determined. For example, multiple fourth characteristic matrices can be fused (such as splicing, addition, etc., which is not limited in this application) to obtain the first characteristic matrix. Alternatively, each fourth characteristic matrix is linearly processed (for example, including linear transformation and activation processing) to obtain a fifth characteristic matrix corresponding to each channel, and each fourth characteristic matrix is dimensionality reduced to obtain a sixth characteristic matrix corresponding to each channel. For example, the characteristic dimension of each audio frame in each fourth characteristic matrix can be reduced to 1 by one or multiple dimensionality reductions (each dimensionality reduction can use linear transformation and activation processing, and the linear transformation parameters used in the multiple dimensionality reductions are different), that is, the characteristic dimension of each audio frame in the fourth characteristic matrix is 1. Of course, before performing linear processing and dimensionality reduction processing on each fourth characteristic matrix, each fourth characteristic matrix can also be linearly transformed and activated in sequence, wherein the parameters of the linear layer used for the linear transformation of each fourth characteristic matrix can be different, and then each fourth characteristic matrix obtained by the linear transformation and activation processing is linearly processed and dimensionality reduced to obtain the fifth characteristic matrix and the sixth characteristic matrix corresponding to each channel, respectively. This is not limited in this application.
[0038] Then, based on the sixth feature matrix corresponding to each channel, the first feature weight of each audio frame under each channel is determined. For example, the sixth feature matrix corresponding to each channel is normalized, such as using a softmax function, to obtain the first feature weight of each audio frame under each channel; then, based on the first feature weight of each audio frame under each channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel is processed to obtain the seventh feature matrix corresponding to each channel. For example, the first feature weight of each audio frame under each channel is weighted with the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel, such as multiplying them, to obtain the seventh feature matrix corresponding to each channel; then, the fourth feature matrix corresponding to each channel and the seventh feature matrix corresponding to each channel are fused (such as splicing, adding, etc., which is not limited in this application) to obtain the first fused feature matrix corresponding to each channel. Of course, before fusion, the seventh feature matrix corresponding to each channel can also be linearly processed first, and then the feature matrix after the linear processing of the seventh feature matrix corresponding to each channel is fused with the fourth feature matrix corresponding to each channel to obtain the first fused feature matrix corresponding to each channel, which is not limited in this application.
[0039] Then, based on the first fused feature matrix corresponding to each channel, the first feature matrix is determined. For example, the first fused feature matrix corresponding to each channel is fused (such as splicing, adding, etc., which is not limited in this application) to obtain the first feature matrix. Alternatively, the second feature weight corresponding to each channel can be determined based on the first fused feature matrix corresponding to each channel. For example, a two-dimensional convolution is performed on the first fused feature corresponding to each channel, and then the feature matrix corresponding to each channel obtained after the two-dimensional convolution is pooled (such as global average pooling, global maximum pooling, etc., which is not limited in this application) to obtain the second feature weight corresponding to each channel; of course, the first fused feature matrix corresponding to each channel can also be fused (such as splicing, addition, etc., which is not limited in this application) to obtain a third fused feature matrix, in which case the third fused feature matrix includes features of multiple channels, and then a two-dimensional convolution is performed on the third fused feature matrix to obtain a fourth fused feature matrix, and then a pooling process is performed on the fourth fused feature matrix (such as global average pooling, global maximum pooling, etc., which is not limited in this application) to obtain a fourteenth feature matrix, in which case the fourteenth feature matrix includes the second feature weight corresponding to each channel in the multiple channels, that is, the dimension of the number of audio frames in the fourteenth feature matrix and the feature dimension of each audio frame are both 1. This application does not specifically limit the method for determining the second feature weight corresponding to each channel.
[0040] Then, based on the second feature weight corresponding to each channel, the first fused feature matrix corresponding to each channel is processed to obtain an eighth feature matrix corresponding to each channel. For example, the second feature weight corresponding to each channel is weighted by the feature vector of each audio frame in the first fused feature matrix corresponding to each channel, such as multiplying them together, to obtain the eighth feature matrix corresponding to each channel; then, based on the eighth feature matrix corresponding to each channel, the first feature matrix is determined. For example, the eighth feature matrix corresponding to each channel is averaged according to the channel dimension to obtain the first feature matrix. At this time, the first feature matrix is the feature matrix of a single channel.
[0041] It should be explained that in the above embodiment, in the process of determining the first feature matrix, by amplifying multiple channels and generating multi-channel feature maps, high-level and rich time-frequency features can be obtained, and then the first feature weight from each channel to each audio frame is dynamically determined based on the sixth feature matrix of each channel, which can pay more attention to features related to recognition results such as emotional information, enhance the importance of enhanced features, and the audio recognition results often have long-range dependencies in time and frequency. By using the first feature weight of each audio frame under each channel, the feature vectors corresponding to each audio frame in the fifth feature matrix corresponding to each channel are weighted to obtain the seventh feature matrix corresponding to each channel, so that these dependencies are captured globally, so that the feature matrix of each channel can better represent the recognition results such as emotional information, enhance the expressive ability of the features, and also use the second feature weight corresponding to each channel to weight the first fusion feature matrix corresponding to each channel, so that the high-level features between different channels can be fused and influenced by each other, thereby improving the representation ability of recognition results such as emotional information, and then improving the accuracy of recognition.
[0042] S102: Determine, based on the first characteristic matrix, a second characteristic matrix for characterizing an association relationship between any two audio frames in the plurality of audio frames.
[0043] In an embodiment of the present application, the first feature matrix can be copied to obtain a copied feature matrix, and then the first feature matrix and the copied feature matrix are fused to obtain a second feature matrix for characterizing the association relationship between any two audio frames in a plurality of audio frames. Each audio frame can be regarded as a time node, that is, the association relationship between the two audio frames at the time nodes can be obtained. For example, the first feature matrix and the transpose of the copied feature matrix are matrix multiplied to obtain the second feature matrix. Alternatively, optionally, the first characteristic matrix is first linearly mapped to obtain multiple fifteenth characteristic matrices. For example, the first characteristic matrix is respectively input into multiple linear layers to obtain multiple fifteenth characteristic matrices corresponding to the multiple linear layers, wherein the parameters of the multiple linear layers may be the same or different, and each fifteenth characteristic matrix has the same dimension as the first characteristic matrix; then, based on the remaining fifteenth characteristic matrices in the multiple fifteenth characteristic matrices except the sixteenth characteristic matrix, the second characteristic matrix is determined. For example, the remaining fifteenth characteristic matrices are fused (the fusion method is similar to the fusion method of the above-mentioned first characteristic matrix and the copied characteristic matrix, which will not be repeated here) to obtain the second characteristic matrix, wherein the sixteenth characteristic matrix is any one of the multiple fifteenth characteristic matrices. This application does not limit the method for determining the second characteristic matrix.
[0044] S103: Determine, based on the second characteristic matrix, a third characteristic matrix for characterizing an association relationship between each audio frame and multiple audio frames.
[0045] In an embodiment of the present application, after obtaining the second feature matrix used to characterize the correlation relationship between any two audio frames in a plurality of audio frames, at least one feature aggregation process can be performed on the feature representation of the correlation relationship between each audio frame in the second feature matrix and each audio frame in the plurality of audio frames to obtain a feature matrix corresponding to each feature aggregation process, wherein the at least one feature aggregation process includes one of the following processes: summation process, averaging process, standard deviation process, etc., that is, at least one feature aggregation process such as summation, averaging, standard deviation, etc. is performed on the feature representation of the correlation relationship between each audio frame in the second feature matrix and each audio frame in the plurality of audio frames to obtain a feature matrix corresponding to each feature aggregation process; then the feature matrices corresponding to each feature aggregation process are fused to obtain a third feature matrix, for example, the feature vectors of each audio frame in the feature matrix corresponding to each feature aggregation process are spliced, added, etc. to obtain the third feature matrix, which is not limited in the present application.
[0046] S104 : Perform recognition based on the first feature matrix and the third feature matrix to obtain a recognition result of the audio to be recognized.
[0047] In an embodiment of the present application, first, based on the third feature matrix, the third feature weight corresponding to each audio frame is determined. For example, the third feature matrix is subjected to linear layer operation and normalization processing (such as using a softmax function) to obtain the third feature weight of each audio frame; then, based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix is processed, such as weighted, to obtain a ninth feature matrix, so as to focus on the importance of different time nodes and further enhance the feature representation of the time series feature vector. It should be noted that if, when determining the second feature matrix, the first feature matrix is linearly mapped to obtain multiple fifteenth feature matrices, At this time, the first feature matrix can be the above-mentioned sixteenth feature matrix, that is, based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the sixteenth feature matrix is weighted to obtain a ninth feature matrix, which is not limited in this application; then the first feature matrix and the ninth feature matrix are fused (such as splicing, addition, etc., which is not limited in this application) to obtain a second fused feature matrix. Of course, optionally, before fusion, the ninth feature matrix can also be linearly processed to smooth it and precipitate node features, and then the feature matrix after the linear processing of the ninth feature matrix is fused with the first feature matrix to obtain a second fused feature matrix, which is not limited in this application.
[0048] Recognition is then performed based on the second fused feature matrix to obtain a recognition result corresponding to the audio to be recognized. For example, the second fused features are directly input into a classifier to output a recognition result corresponding to the audio to be recognized. Alternatively, before recognition is performed based on the second fused feature matrix, the second fused features may be downsampled, and then recognition is performed based on the feature matrix after the downsampling of the second fused features to obtain a recognition result corresponding to the audio to be recognized. Or optionally, the similarity corresponding to each audio frame can be determined based on the feature vector corresponding to each audio frame in the second fusion feature matrix and the label vector corresponding to each audio frame (the type of similarity is not limited in this application), and then based on the similarity corresponding to each audio frame, the feature vector corresponding to each audio frame in the second fusion feature matrix is sorted in order of similarity from large to small, and based on a preset screening rule (for example, the similarity is less than a preset threshold), the multiple feature vectors corresponding to the sorted multiple audio frames are screened to obtain multiple first feature vectors corresponding to multiple first audio frames, wherein the number of the multiple first audio frames is less than or equal to the number of the multiple audio frames corresponding to the second fusion feature matrix, and then based on the multiple first feature vectors, a second feature vector of a preset dimension is determined, for example, the multiple first feature vectors are averaged according to the dimension of the audio frame to obtain a second feature vector, and then recognition is performed based on the second feature vector to obtain a recognition result corresponding to the audio to be recognized.
[0049] In an optional embodiment, the second fused feature matrix can also be subjected to graph convolution processing to obtain a tenth feature matrix. Specifically: first, an adjacency matrix corresponding to the audio to be identified is obtained. For example, the adjacency matrix can be undirected, and the audio frame nodes that are adjacent in time sequence are connected according to the timing information of the audio frame to obtain the adjacency matrix; then, based on the adjacency matrix, a self-connected adjacency matrix corresponding to the audio to be identified is determined. For example, the adjacency matrix corresponding to the audio to be identified is added to the unit matrix to obtain a self-connected adjacency matrix corresponding to the audio to be identified; then, the self-connected adjacency matrix is symmetrically normalized to obtain a seventeenth feature matrix. For example, based on the degree matrix, The inverse of the square of the degree corresponding to each audio frame (also known as each node) is calculated, that is, the importance of the audio frame is weighted, and then the self-connected adjacency matrix corresponding to the audio to be identified and the inverse of the square of the degree corresponding to each audio frame are multiplied together, and then multiplied with the inverse of the square of the degree corresponding to each audio frame to obtain the seventeenth characteristic matrix, which realizes symmetric normalization. In this way, the eigenvector corresponding to each audio frame in the obtained seventeenth characteristic matrix has been adjusted in degree, ensuring that the relationship between adjacent nodes / adjacent audio frames can be correctly reflected, preventing some nodes from dominating the feature propagation too much due to more connections, ensuring the effective propagation of features, and enhancing stability.
[0050] Then, for the k-th graph convolution processing, feature aggregation is performed based on the feature vector corresponding to each audio frame obtained by the k-1-th graph convolution processing and the feature vector of each audio frame in the seventeenth feature matrix to obtain the eighteenth feature matrix corresponding to the k-th graph convolution processing, until multiple graph convolution processing is performed to obtain the tenth feature matrix, wherein the k-th graph convolution processing is any one of the multiple graph convolution processings. When k=1, the feature vector of each audio frame obtained by the k-1-th graph convolution processing is the feature vector of each audio frame in the second fusion feature matrix, wherein the k-th graph convolution processing can be obtained by formula (1):
[0051]
[0052] Among them, H k represents the eighteenth characteristic matrix corresponding to the k-th graph convolution process, A represents the adjacency matrix, I represents the identity matrix, represents the degree matrix, represents the seventeenth characteristic matrix, H k-1 represents the feature vector of each audio frame obtained by the k-1th graph convolution process, W k represents the weight matrix corresponding to the i-th graph convolution process, and σ() represents the activation function, such as the RelU function. After the graph convolution operation, each audio frame or each node can capture the information of its adjacent audio frames or adjacent nodes, which can enhance the representation ability of its feature vector.
[0053] Of course, it is optional that before performing graph convolution processing on the second fused feature matrix, the second fused feature matrix can also be instance normalized to remove internal variability in the audio signal, such as noise, which is conducive to capturing relevant information of the recognition result and discarding other irrelevant information. Then, the feature matrix after instance normalization is subjected to graph convolution processing to obtain the sixth feature matrix. The principle refers to the explanation of the above graph convolution processing and will not be repeated here. It should be explained that by normalizing the first fused feature instance, removing audio noise, and symmetric normalizing the adjacency matrix, the seventeenth feature matrix is obtained. Then, feature aggregation is performed based on the feature vector of each audio frame obtained by the k-1th graph convolution processing and the feature vector of each audio frame in the seventeenth feature matrix. This can avoid the problems of gradient explosion and gradient disappearance in the activation process. Multiple graph convolution processing is then used to realize multiple aggregation and activation operations of the features, so that the features of each audio frame in the tenth feature matrix simultaneously capture the information of itself and the adjacent audio frames, thereby enhancing the expressive power of the audio frame features.
[0054] Then, an average pooling operation is performed on the feature vector of each audio frame in the tenth feature matrix to obtain an eleventh feature matrix. Then, a linear mapping is performed on the feature vector of each audio frame in the eleventh feature matrix to obtain a twelfth feature matrix, where the dimension of the twelfth feature matrix is the same as the dimension of the tenth feature matrix. For example, assuming that the dimension of the tenth feature matrix is (N, F), the dimension of the eleventh feature matrix obtained after the average pooling operation is (N, 1). Then, the eleventh feature matrix is linearly mapped to obtain a twelfth feature matrix with a dimension of (N, F). Then, based on the twelfth feature matrix, the fourth feature weight corresponding to each audio frame is determined. For example, the twelfth feature matrix is activated based on the sigmoid function to obtain the fourth feature weight of each audio frame, so as to obtain the importance of different audio frames, or the importance of different time nodes. Then, based on the fourth feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the tenth feature matrix is processed to determine a thirteenth feature matrix. That is, the fourth feature weight corresponding to each audio frame is weighted with the feature vector corresponding to each audio frame in the tenth feature matrix to obtain the thirteenth feature matrix, so as to enhance the position representation of each audio frame. It should be explained that the average pooling operation is a compression operation, which can further enhance the ability to express the features of different nodes. Compression is highly abstract, which may filter out information irrelevant to the task, such as retaining only information related to the recognition results, and may also enhance the representation of other important information. Then, based on the tenth feature matrix, linear mapping and reactivation processing can be performed to obtain the importance of different nodes, and then the enhanced feature representation of each node can be obtained.
[0055] Then, recognition is performed based on the thirteenth characteristic matrix to obtain a recognition result corresponding to the audio to be recognized, such as inputting the thirteenth characteristic matrix into the classifier, and determining the result with the maximum classification probability as the recognition result corresponding to the audio to be recognized. Of course, optionally, before recognition, the thirteenth characteristic matrix can also be subjected to a first linear transformation or a first affine transformation to obtain a nineteenth characteristic matrix, and then the nineteenth characteristic matrix and the second fused characteristic matrix are fused (such as addition, splicing, etc., which are not limited in this application), or the nineteenth characteristic matrix and the characteristic matrix after instance normalization of the second fused characteristic matrix are fused (such as residual addition, splicing, etc., which are not limited in this application) to obtain a twentieth characteristic matrix; then the twentieth characteristic matrix is subjected to a second affine transformation or a second linear transformation to obtain a twenty-first characteristic matrix, which can enhance the feature expression ability of different nodes; finally, recognition is performed based on the twenty-first characteristic matrix to obtain a recognition result corresponding to the audio to be recognized, such as inputting the twenty-first characteristic matrix into the classifier, and determining the result with the maximum classification probability as the recognition result corresponding to the audio to be recognized. Alternatively, optionally, the principle of performing recognition based on the twenty-first feature matrix to obtain a recognition result corresponding to the audio to be recognized is similar to the principle of performing recognition based on the second fused feature matrix mentioned above, that is, the similarity corresponding to each audio frame can also be determined based on the feature vector of each audio frame in the twenty-first feature matrix and the label vector corresponding to each audio frame (the type of similarity is not limited in this application), and then based on the similarity corresponding to each audio frame, the feature vector of each audio frame in the twenty-first feature matrix is sorted in descending order according to the similarity, and then based on a preset screening rule (for example, the similarity is less than a preset threshold), the multiple feature vectors corresponding to the sorted multiple audio frames are screened to obtain multiple third feature vectors corresponding to multiple second audio frames, wherein the number of the multiple second audio frames is less than or equal to the number of the multiple audio frames corresponding to the twenty-first feature matrix, and then based on the multiple third feature vectors, a fourth feature vector of a preset dimension is determined, for example, the multiple third feature vectors are averaged according to the dimension of the audio frame to obtain a fourth feature vector, and then recognition is performed based on the fourth feature vector to obtain a recognition result corresponding to the audio to be recognized.
[0056] It should be explained that, since the dimensions of the high-level feature information finally obtained for audio signals or voice information of different durations are different, and for the classifier, it requires that the input high-level feature information have the same dimension, and the existing classifiers either cannot use the high-level feature information as the input of the classifier, or directly aggregate the high-level feature information into a fixed dimension, which makes it difficult to capture global information. Therefore, in an embodiment of the present application, the second fused feature matrix is subjected to multiple graph convolution processes to obtain a tenth feature matrix, and the feature vector of each audio frame in the tenth feature matrix is average pooled to obtain the fourth feature weight corresponding to each audio frame. Based on the fourth feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the tenth feature matrix is weighted to obtain a thirteenth feature matrix, and then the thirteenth feature matrix is dimensionally aggregated to obtain a feature vector of a preset dimension. This not only can capture global information while enhancing feature representation, but also can aggregate the features of audio of any indefinite duration into a feature vector of a preset dimension that is more capable of distinguishing classification categories, that is, it satisfies the input dimension of the classifier and effectively solves the problem of audio recognition for any duration.
[0057] In an optional embodiment, the first feature matrix and the third feature matrix can also be fused, for example, the feature vectors of each audio frame in the first feature matrix and the third feature matrix are spliced, added, weighted averaged, etc. to obtain a fifth fused feature matrix, and then recognition is performed based on the fifth fused feature matrix to obtain a recognition result corresponding to the audio to be recognized. The principle of performing recognition based on the fifth fused feature matrix to obtain a recognition result corresponding to the audio to be recognized can be similar to the principle of performing recognition based on the second fused feature matrix or the thirteenth feature matrix to obtain a recognition result corresponding to the audio to be recognized, and the same technical effect can be achieved, which will not be repeated here.
[0058] In an optional embodiment, the above audio recognition method can also be performed by an audio recognition model, and the training method of the audio recognition model is applied to the training device of the audio recognition model, see Figure 2 , Figure 2 A flowchart of a method for training an audio recognition model provided in an embodiment of the present application is provided. The training method includes steps S201-S206:
[0059] S201: Obtain audio to be trained.
[0060] S202 : Perform feature extraction on multiple audio frames in the training audio to obtain a twenty-second feature matrix.
[0061] The principle of determining the twenty-second characteristic matrix is similar to the principle of obtaining the first characteristic matrix, and will not be repeated here.
[0062] S203 : Based on the twenty-second characteristic matrix, determine a twenty-third characteristic matrix for characterizing the association relationship between any two audio frames in the plurality of audio frames to be trained.
[0063] The principle of determining the twenty-third characteristic matrix is similar to the principle of determining the second characteristic matrix corresponding to the audio to be recognized, and will not be repeated here.
[0064] S204 : Based on the twenty-third characteristic matrix, determine a twenty-fourth characteristic matrix for characterizing the association relationship between each audio frame of the to-be-trained audio and the multiple audio frames of the to-be-trained audio.
[0065] The principle of determining the twenty-fourth characteristic matrix is similar to the principle of determining the third characteristic matrix corresponding to the audio to be recognized, and will not be repeated here.
[0066] S205 : Perform recognition based on the twenty-second characteristic matrix and the twenty-fourth characteristic matrix to obtain a predicted recognition result of the audio to be trained.
[0067] The principle of determining the predicted recognition result corresponding to the audio to be trained is similar to the principle of obtaining the recognition result corresponding to the audio to be recognized, and will not be repeated here.
[0068] S206: Train the audio recognition model based on the predicted recognition result and label of the audio to be trained.
[0069] For example, the loss is determined based on the predicted recognition result and the label. The type of loss is, for example, cross entropy loss, etc., which is not limited in this application. Then, the audio recognition model is trained based on the loss until the loss meets the preset conditions to obtain a trained audio recognition model.
[0070] In an alternative embodiment, see Figure 3 , Figure 3 This is a structural diagram of an audio recognition model provided in an embodiment of the present application. Figure 3As shown, the audio recognition model includes a feature extraction module, a feature processing module, a pooling module and a classification module. The feature extraction module can be a mel-spectrogram extractor. Based on the explanation of the above embodiment, the audio to be recognized is input into the feature extraction module to obtain a first feature matrix; the first feature matrix is then input into the feature processing module, and the feature processing module performs feature extraction on the first feature matrix to obtain a second feature matrix for characterizing the association relationship between any two audio frames in a plurality of audio frames. Refer to the above principle explanation, which will not be repeated here, and the high-level feature extraction model performs feature extraction on the second feature matrix to obtain a third feature matrix for characterizing the association relationship between each audio frame and a plurality of audio frames. Refer to the above principle explanation, which will not be repeated here; then the feature processing module obtains a second fused feature matrix based on the first feature matrix and the third feature matrix, which will be explained with reference to the above principle explanation, which will not be repeated here; then the second fused feature matrix is input into the pooling module to output a thirteenth feature matrix, which will be explained with reference to the above principle explanation, which will not be repeated here; then the thirteenth feature matrix is input into the classification module to perform recognition based on the thirteenth feature matrix to obtain a recognition result corresponding to the audio to be recognized.
[0071] In an optional embodiment, based on Figure 3 For example, see Figure 4 , Figure 4 This is a structural diagram of another audio recognition model provided in the embodiment of the present application. Figure 4 The feature processing module in the audio recognition model shown includes a first feature processing module, a second feature processing module, and a downsampling module. Figure 4 Other modules shown refer to Figure 3; then the audio to be recognized is input into the feature extraction module, and an initial feature matrix is output; the initial feature matrix is then input into the first feature processing module, and a first feature matrix is output. The principle is explained with reference to the above, and no further details are given here; the first feature matrix is then input into the second feature processing module, and the second feature processing module performs feature extraction on the first feature matrix to obtain a second feature matrix, and the explanation is explained with reference to the above principle, and no further details are given here; and the second feature processing module performs feature extraction on the second feature matrix to obtain a third feature matrix, and the explanation is explained with reference to the above principle, and no further details are given here; then the second feature processing module performs feature extraction on the third feature matrix to obtain a ninth feature matrix, and the explanation is explained with reference to the above principle, and no further details are given here; then the feature processing module performs feature extraction on the third feature matrix based on the first feature matrix Array, the ninth feature matrix, obtain the second fused feature matrix, refer to the above principle explanation, no further details are given here; then the second fused feature matrix is input into the downsampling module, until it is executed N times, and the output of the feature processing module after N times, that is, the output of the downsampling module, is input into the pooling module to obtain the thirteenth feature matrix, refer to the above principle explanation of determining the thirteenth feature matrix based on the second fused feature, no further details are given here. It should be noted that the audio recognition method in the above embodiment of the present application is an example explanation based on the execution of the feature processing module once. If it is executed N times, the principle of each execution is similar, and it is not elaborated here; then the thirteenth feature matrix is input into the classification module to perform recognition based on the thirteenth feature matrix to obtain a recognition result corresponding to the audio to be recognized.
[0072] In an optional embodiment, based on Figure 4 For example, see Figure 5 , Figure 5 This is a structural diagram of another audio recognition model provided in the embodiment of the present application. Figure 5 The pooling module in the audio recognition model shown includes an instance normalization module, a graph convolution module, an average pooling module, a linear mapping module, an activation module, a first linear transformation module (or a first affine transformation module, Figure 5 Not shown), a second linear transformation module (or a second affine transformation module, Figure 5 (not shown), feature dimension processing module, Figure 5 Other modules shown refer to Figure 4; after obtaining the second fused feature matrix based on the above embodiment, the second fused feature matrix is input into the instance normalization module to perform instance normalization on the second fused feature matrix; the feature matrix after instance normalization is then input into the graph convolution module for graph convolution processing to obtain a tenth feature matrix. The principle refers to the explanation of the above graph convolution processing and will not be repeated here; then the feature vector of each audio frame in the tenth feature matrix is averaged by the pooling module to obtain an eleventh feature matrix; then the feature vector of each audio frame in the eleventh feature matrix is linearly mapped by the linear mapping module to obtain a twelfth feature matrix; then the twelfth feature matrix is activated by the activation module to obtain the fourth feature weight of each audio frame; then the pooling module performs an average pooling operation on the feature vector of each audio frame in the tenth feature matrix based on the fourth feature weight corresponding to each audio frame. The feature vector is processed, such as weighted, to obtain a thirteenth feature matrix; the thirteenth feature matrix is then subjected to a first linear transformation through a first linear transformation module to obtain a nineteenth feature matrix; the pooling module then fuses the nineteenth feature matrix with the second fused feature matrix, or fuses the feature matrix after instance normalization of the nineteenth feature matrix and the second fused feature matrix to obtain a twentieth feature matrix; the second linear mapping module then performs a second linear transformation on the twentieth feature matrix to obtain a twenty-first feature matrix; the feature dimension processing module then performs feature processing on the twenty-first feature matrix to obtain a second feature vector. With reference to the above principles, no further details will be given here. The second feature vector is then input into the classification module for recognition based on the second feature vector to obtain a recognition result corresponding to the audio to be recognized. With reference to the above principles, no further details will be given here.
[0073] For ease of understanding, the following examples illustrate the network structure of each module in the audio recognition model. Figure 3-Figure 5 For example, see Figure 6 , Figure 6 A structural diagram of another audio recognition model provided in an embodiment of the present application.
[0074] Before that, Figure 6 The relevant contents involved are explained:
[0075] Linear layer+Relu(): It includes linear layer and activation layer, which correspond to linear transformation processing and activation processing respectively. The numbers in () are only for Figure 6 To facilitate distinction, the parameters of the linear layers in Linear layer+Relu() with different numerical sequences can be the same or different, and are mainly adjusted according to actual needs;
[0076] Linear layer(): represents the linear layer, corresponding to linear transformation processing or affine transformation processing, the numerical sequence in () is also for Figure 6 To facilitate distinction, the parameters of the linear layers of different numerical sequences can be the same or different, and are mainly adjusted according to actual needs;
[0077] BN+Relu(), BN+Sigmoid(): both represent normalization layer and activation layer, corresponding to normalization processing and activation processing respectively. The numerical sequence in () is also for Figure 6 To facilitate distinction, the parameters of the normalization layers in BN+Relu() and BN+Sigmoid() with different numerical sequences can be the same or different, and are mainly adjusted according to actual needs.
[0078] like Figure 6 As shown, the embodiment of the present application takes a sample number as an example to illustrate, the audio to be recognized is input as input, input into the feature extraction module, and an initial feature matrix is obtained. Assume that the dimension is (T, D) where T is the number of audio frames and D is the feature dimension of each audio frame; then the first feature processing module in the feature processing module performs feature extraction processing on the initial feature matrix, specifically including: first, adding the channel dimension to the initial feature matrix, that is, the corresponding Figure 6 Addchannel dimention in the , get the first feature map (1, T, D); then the first feature map is channel amplified, for example, the first feature map is subjected to a deep two-dimensional convolution, which corresponds to Figure 6 Depth-Conv2D in the image is used to obtain the second feature map (3, T, D). It should be noted that the embodiment of the present application does not limit the number of channels. Here, 3 channels are used as an example for explanation. Then, the second feature map is separated according to the channel, that is, the corresponding Figure 6 Channel split in , obtains multiple third feature maps corresponding to multiple channels, and then determines multiple fourth feature matrices corresponding to multiple channels based on the multiple third feature maps, that is, three fourth feature matrices (T, D); then, one of the three fourth feature matrices (T, D) is used as an example to illustrate that the fourth feature matrix is linearly processed, that is, the corresponding Figure 6 The Linear layer+Relu(4) in the channel is used to obtain the fifth feature matrix corresponding to the channel, namely (T, D), and the fourth feature matrix is subjected to multiple dimensionality reduction processes, namely, first based on linear transformation and activation processing, namely, the corresponding Figure 6 The Linear layer+Relu(7) in the image is reduced to (T,D / 2), and then the corresponding linear transformation and activation processing are performed. Figure 6The Linear layer+Relu(8) in the equation is obtained to get (T,D / 4), and then the corresponding linear transformation and activation processing is obtained. Figure 6 The Lin ear layer+Relu (9) in the channel obtains the sixth feature matrix corresponding to the channel, namely (T, 1); then the sixth feature matrix corresponding to the channel is normalized, namely the Softmax function is used to obtain the first feature weight of each audio frame under the channel, namely (T, 1); then based on the first feature weight of each audio frame under the channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to the channel is weighted to obtain the seventh feature matrix corresponding to the channel, namely (T, D); then the seventh feature matrix corresponding to the channel is linearly processed, namely the corresponding Figure 6 Linear layer+Relu(16) in the above equation, and then linear processing is performed on it, which corresponds to Figure 6 The features obtained by Linear layer+Relu(16) in the channel are fused with the fourth feature matrix corresponding to the channel, and then linear processing is performed based on Linear layer+Relu(19) to obtain the first fused feature matrix corresponding to the channel, namely (T, D). Similarly, the first fused feature matrix (T, D) corresponding to each channel can be obtained, which will not be repeated here; then the first fused feature matrix corresponding to each channel is fused, namely the corresponding Figure 6 The concat in the third fusion feature matrix is obtained, namely (3, T, D); then the third fusion feature matrix is subjected to two-dimensional convolution, namely the corresponding Figure 6 conv2D in the fourth fusion feature matrix, namely (3, T, D); then the fourth fusion feature matrix is subjected to global mean pooling processing, namely the corresponding Figure 6 The global average in the 14th feature matrix is obtained, that is, (3, T, 1), that is, the second feature weight corresponding to each channel is obtained; then the first fusion feature matrix corresponding to each channel is weighted based on the second feature weight corresponding to each channel to obtain the eighth feature matrix corresponding to each channel, that is, (3, T, D); then the eighth feature matrix corresponding to each channel is averaged according to the channel dimension, that is, the corresponding Figure 6 The mean in , we get the first feature matrix (T, D).
[0079] Then, the second feature processing module in the feature processing module performs feature extraction processing on the first feature matrix, specifically including: first, linear mapping is performed on the first feature matrix through multiple linear layers, namely Linear layer (1), Linear layer (2), Linear layer (3), to obtain multiple fifteenth feature matrices, namely K = (T, D), M = (T, D), N = (T, D). It should be noted that only three fifteenth feature matrices are used as examples here; then, matrix multiplication is performed based on the fifteenth feature matrix K and the fifteenth feature matrix M, namely the corresponding Figure 6 Dot(K,M T ), and obtain the second feature matrix (T, T); then perform at least one feature aggregation process on the second feature matrix, namely, the summation process, namely, the corresponding Figure 6 The Sum and mean processing in the corresponding Figure 6 Mean and standard deviation in the corresponding Figure 6 The standard deviation in , respectively, obtains the feature matrix (T,1) corresponding to each feature aggregation process; then the feature matrix corresponding to each feature aggregation process is fused, namely contact, to obtain the third feature matrix (T,3); then the third feature matrix is linearly processed, namely the corresponding Figure 6 The Linear layer (4) and the normalization process, i.e., the softmax function, are used to obtain the third feature weight of each audio frame, i.e., Q = (T, 1); then, based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix is weighted, i.e., the corresponding Figure 6 Q*N in the equation, we get the ninth characteristic matrix (T, D); then we perform linear processing on the ninth characteristic matrix, which corresponds to Figure 6 The Linear layer (5) and normalization process in the corresponding Figure 6 The Layernorm feature matrix (T, D) in is fused with the first feature matrix to obtain the second fused feature matrix (T, D).
[0080] Then the downsampling module in the feature processing module performs downsampling processing on the second fused feature matrix, specifically including: performing two-dimensional convolution processing on the second fused feature matrix, that is, corresponding to Figure 6 Conv2d_1 in the above code is then normalized and activated to correspond to Figure 6 BN+Relu(1) in the function gets the feature matrix (T, D), and then performs two-dimensional convolution on the feature matrix, which corresponds to Figure 6 Conv2d_2 in the above code is then normalized and activated to correspond to Figure 6 BN+Relu(2) in the equation, we get the feature matrix (T / 2N ,D / 2 N ).
[0081] Then the pooling module performs the feature matrix (T / 2 N ,D / 2 N) Perform pooling processing, specifically including: feature matrix (T / 2 N ,D / 2 N ) to normalize the instance, that is, Figure 6 InstanceNorm in, and then perform graph convolution processing, which corresponds to Figure 6 The GCN in the tenth feature matrix is obtained; then the feature vector of each audio frame in the tenth feature matrix is average pooled, which corresponds to Figure 6 The Mean_pool in the eleventh feature matrix is obtained; then the eigenvector of each audio frame in the eleventh feature matrix is linearly mapped to the corresponding Figure 6 The Linear layer (6) in the twelfth feature matrix is obtained; the twelfth feature matrix is then activated based on the sigmoid function to obtain the fourth feature weight of each audio frame; then, based on the fourth feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the tenth feature matrix is weighted to determine the thirteenth feature matrix; then the thirteenth feature matrix is subjected to the first linear transformation, that is, the corresponding Figure 6 Linearlayer (7) in the fusion matrix, the nineteenth feature matrix is obtained; then the nineteenth feature matrix is fused with the feature matrix after the second fusion feature matrix is instance-normalized to obtain the twentieth feature matrix; then the second linear transformation is performed on the twentieth feature matrix, which corresponds to Figure 6 Li near layer (8) in the 21st characteristic matrix is obtained; then the characteristic dimension of the 21st characteristic matrix is processed to correspond to Figure 6 Feature dimension in , and get the fourth eigenvector.
[0082] Then the classification module performs recognition based on the fourth eigenvector, specifically including: performing linear processing on the fourth eigenvector, i.e., corresponding Figure 6 The Linear layer (9) in the , and then normalized and activated corresponding to Figure 6 The fifth eigenvector is obtained by applying BN+Sigmoid(1) in the equation, and the dimension of the fifth eigenvector is the same as that of the fourth eigenvector. Then the fifth eigenvector is linearly processed, which corresponds to Figure 6 The Linear layer (10) in the , and then normalized and activated corresponding to Figure 6The BN+Sigmoid(2) in the algorithm is used to obtain the classification probabilities under the preset classification categories. Then, the recognition result corresponding to the audio to be recognized can be determined based on the classification probabilities under the preset classification categories. For example, the category with the largest classification probability or the category with a classification probability greater than a preset threshold is used as the recognition result corresponding to the audio to be recognized.
[0083] It can be seen that by amplifying multiple channels through the first feature processing module in the feature processing module and generating multiple fourth feature matrices corresponding to multiple channels, high-level and rich time-frequency features can be obtained. Then, the first feature weight from each channel to each audio frame is dynamically determined based on the fourth feature matrix of each channel, which can pay more attention to features related to recognition results such as emotional information, enhance the importance of enhanced features, and the audio recognition results often have long-range dependencies in time and frequency. By adding the first feature weight of each audio frame under each channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel is weighted to obtain the seventh feature matrix corresponding to each channel, so that these dependencies are captured globally, so that the feature matrix of each channel can better represent the recognition results such as emotional information, enhance the expressive ability of the features, and also by adding the second feature weight corresponding to each channel, the first fusion feature matrix corresponding to each channel is weighted, so that the high-level features between different channels can be integrated and influenced by each other, thereby improving the representation ability of recognition results such as emotional information, and then improving the accuracy of audio recognition. Accuracy; in addition, the first feature matrix is processed by the second feature processing module to obtain a second feature matrix for characterizing the correlation relationship between any two audio frames in the multiple audio frames in turn, so as to obtain the relationship between time nodes or audio frames, and then a third feature matrix for characterizing the correlation relationship between each audio frame and the multiple audio frames is determined based on the second feature matrix, and the third feature weight of each audio frame is calculated based on the third feature matrix and weighted fused with the first feature matrix to obtain a second fused feature matrix, which enhances the feature representation of the audio frame and pays more attention to the relationship between time nodes or audio frames; and the second fused feature matrix is also subjected to graph convolution processing to capture the information of its adjacent audio frames or adjacent nodes, further enhancing the features, and the feature vectors of each audio frame are adjusted to ensure that the relationship between adjacent nodes / adjacent audio frames can be correctly reflected, which can prevent some nodes from dominating the feature propagation due to more connections, ensure the effective propagation of features, enhance stability, and thus improve the accuracy of audio recognition.
[0084] See Figure 7 , Figure 7 A schematic diagram of an audio recognition system provided in an embodiment of the present application.
[0085] Figure 7The system shown includes an audio recognition device and a client; the audio recognition device can be a server, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as basic cloud computing services such as big data and artificial intelligence platforms, which are not specifically limited in this application; the client can be a smart phone, tablet computer, laptop computer, desktop computer, smart TV, desktop computer, smart watch, smart car and other smart terminals, but is not limited to these.
[0086] The user can upload the audio to be recognized through the client, or obtain the user's audio (such as voice) through the client, and then the client sends the audio to be recognized to the audio recognition device, and then the audio recognition device recognizes the audio to be recognized, specifically as follows: feature extraction is performed on each audio frame in the audio to be recognized to obtain a first feature matrix; based on the first feature matrix, a second feature matrix is determined for characterizing the association relationship between any two audio frames in multiple audio frames; based on the second feature matrix, a third feature matrix is determined for characterizing the association relationship between each audio frame and multiple audio frames; recognition is performed based on the first feature matrix and the third feature matrix to obtain a recognition result corresponding to the audio to be recognized; and then the recognition result corresponding to the audio is sent to the client.
[0087] It should be noted that Figure 7 The audio recognition device in the embodiment may also correspondingly execute other steps of the audio recognition method executed by the audio recognition device in the above embodiment, which will not be described in detail here.
[0088] See Figure 8 , Figure 8 A schematic diagram of a system for training an audio recognition model provided in an embodiment of the present application.
[0089] Figure 8 The system shown includes a training device and a client for an audio recognition model; the training device for the audio recognition model can be a server, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as basic cloud computing services such as big data and artificial intelligence platforms, which are not specifically limited in this application; the client can be a smart phone, tablet computer, laptop computer, desktop computer, smart TV, desktop computer, smart watch, smart car and other smart terminals, but is not limited to these.
[0090] The user can upload the audio to be trained through the client, and then the client sends the audio to be trained to the training device of the audio recognition model, and then the training device of the audio recognition model performs model training based on the audio to be trained, specifically as follows: obtain the audio to be trained; then perform feature extraction on each audio frame in the audio to be trained to obtain a twenty-second feature matrix; then, based on the twenty-second feature matrix, determine a twenty-third feature matrix for characterizing the correlation between audio frames of the audio to be trained; then, based on the twenty-third feature matrix, determine a twenty-fourth feature matrix for characterizing the global features of each audio frame of the audio to be trained; then, perform recognition based on the twenty-second feature matrix and the twenty-fourth feature matrix to obtain a predicted recognition result corresponding to the audio to be trained; then, train the audio recognition model based on the predicted recognition result and label corresponding to the audio to be trained.
[0091] It should be noted that Figure 8 The audio recognition model training device in the embodiment can also correspondingly execute other steps of the audio recognition model training method in the above embodiment, which will not be repeated here.
[0092] It should be noted that the main scenarios applicable to audio recognition in this application may include but are not limited to the following scenarios:
[0093] (1) In the fields of customer service, intelligent outbound calls, etc., when a virtual agent communicates with a user through a client, the call audio with the user can be obtained in real time, and then the client sends the call audio to the audio recognition device. The audio recognition device then recognizes the call audio according to the audio recognition method of the example of this application, obtains the user's current emotion, and can observe the user's emotional fluctuations in real time, and then can match the user with a corresponding call speech strategy based on the user's current emotion.
[0094] (2) In the field of teaching, the audio information of students in class can be obtained and then sent to an audio recognition device. The audio recognition device then recognizes the call audio according to the audio recognition method of the example of this application to obtain the emotional state of the students, thereby understanding the students' participation and emotions in class, and thus helping teachers to effectively adjust teaching strategies; and audio information from other places on campus can also be used to recognize students' emotions, so as to timely understand the students' current status and reduce the probability of dangerous events such as school violence.
[0095] (3) Mental health monitoring: Identifying and analyzing the patient's conversation audio based on the audio recognition method of this application can help psychologists understand the patient's emotional fluctuations and thus provide more targeted treatment plans.
[0096] (4) Social media monitoring: Analyze users’ voice or audio information on social platforms through the audio recognition method of this application to gain insights into public sentiment and public opinion trends.
[0097] It should be noted that the above scenarios are only some examples, and the audio recognition method of the present application can be adaptively adjusted and applied in corresponding fields.
[0098] See Figure 9 , Figure 9 This is a block diagram of the functional units of an audio recognition device provided in an embodiment of the present application. The audio recognition device 900 includes: an acquisition unit 901 and a processing unit 902;
[0099] An acquisition unit 901 is configured to acquire audio to be recognized;
[0100] Processing unit 902 is used to extract features from multiple audio frames in the audio to be recognized to obtain a first feature matrix; based on the first feature matrix, determine a second feature matrix for characterizing the association relationship between any two audio frames in the multiple audio frames; based on the second feature matrix, determine a third feature matrix for characterizing the association relationship between each audio frame and the multiple audio frames; and perform recognition based on the first feature matrix and the third feature matrix to obtain a recognition result of the audio to be recognized.
[0101] In one embodiment of the present application, in determining a third feature matrix for characterizing an association relationship between each audio frame and multiple audio frames based on the second feature matrix, the processing unit 902 is specifically configured to:
[0102] Performing at least one feature aggregation process on the feature vector of each audio frame in the second feature matrix to obtain a feature matrix corresponding to each feature aggregation process;
[0103] The feature matrices corresponding to each feature aggregation process are fused to obtain a third feature matrix;
[0104] The at least one feature aggregation process includes at least one of the following processes: summation process, mean value process, and standard deviation process.
[0105] In one embodiment of the present application, in terms of performing recognition based on the first feature matrix and the third feature matrix to obtain a recognition result of the audio to be recognized, the processing unit 902 is specifically configured to:
[0106] Determining a third feature weight corresponding to each audio frame based on the third feature matrix;
[0107] Based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix is processed to obtain a ninth feature matrix;
[0108] Fusing the first feature matrix and the ninth feature matrix to obtain a second fused feature matrix;
[0109] Recognition is performed based on the second fused feature matrix to obtain a recognition result of the audio to be recognized.
[0110] In one embodiment of the present application, in extracting features from multiple audio frames in the audio to be recognized to obtain a first feature matrix, the processing unit 902 is specifically configured to:
[0111] Extract features from multiple audio frames in the audio to be recognized to obtain an initial feature matrix;
[0112] Based on the initial characteristic matrix, determining a plurality of fourth characteristic matrices corresponding to the plurality of channels;
[0113] A first characteristic matrix is determined based on a plurality of fourth characteristic matrices corresponding to the plurality of channels.
[0114] In one embodiment of the present application, in determining the first characteristic matrix based on a plurality of fourth characteristic matrices corresponding to a plurality of channels, the processing unit 902 is specifically configured to:
[0115] Perform linear processing on each fourth characteristic matrix to obtain a fifth characteristic matrix corresponding to each channel;
[0116] Perform dimensionality reduction on each fourth characteristic matrix to obtain a sixth characteristic matrix corresponding to each channel;
[0117] Determining a first feature weight of each audio frame under each channel based on a sixth feature matrix corresponding to each channel;
[0118] Based on the first feature weight of each audio frame under each channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel is processed to obtain a seventh feature matrix corresponding to each channel;
[0119] The fourth feature matrix corresponding to each channel and the seventh feature matrix corresponding to each channel are fused to obtain a first fused feature matrix corresponding to each channel;
[0120] A first feature matrix is determined based on the first fused feature matrix corresponding to each channel.
[0121] In one embodiment of the present application, in determining the first feature matrix based on the first fused feature matrix corresponding to each channel, the processing unit 902 is specifically configured to:
[0122] Determine a second feature weight corresponding to each channel based on the first fusion feature matrix corresponding to each channel;
[0123] Based on the second feature weight corresponding to each channel, the first fusion feature matrix corresponding to each channel is weighted to obtain an eighth feature matrix corresponding to each channel;
[0124] A first characteristic matrix is determined based on the eighth characteristic matrix corresponding to each channel.
[0125] In one embodiment of the present application, in terms of performing recognition based on the second fused feature matrix to obtain a recognition result of the audio to be recognized, the processing unit 902 is specifically configured to:
[0126] Perform graph convolution on the second fused feature matrix to obtain the tenth feature matrix;
[0127] Performing average pooling on the feature vectors of each audio frame in the tenth feature matrix to obtain an eleventh feature matrix;
[0128] Performing linear mapping on the eigenvector of each audio frame in the eleventh eigenmatrix to obtain a twelfth eigenmatrix;
[0129] Determining a fourth feature weight corresponding to each audio frame based on the twelfth feature matrix;
[0130] Based on the fourth feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the tenth feature matrix is processed to determine a thirteenth feature matrix;
[0131] Recognition is performed based on the thirteenth characteristic matrix to obtain a recognition result corresponding to the audio to be recognized.
[0132] In a specific implementation, the acquisition unit 901 and the processing unit 902 described in the embodiment of the present invention may also execute other implementations described in the embodiment of the audio recognition method provided by the embodiment of the present invention, which will not be described in detail here.
[0133] See Figure 10 , Figure 10 This is a block diagram of the functional units of an audio recognition model training device provided in an embodiment of the present application. The audio recognition model training device 1000 includes: a first acquisition unit 1001 and a first processing unit 1002;
[0134] A first acquiring unit 1001 is configured to acquire audio to be trained;
[0135] The first processing unit 1002 is configured to perform feature extraction on multiple audio frames in the audio to be trained to obtain a twenty-second feature matrix; determine, based on the twenty-second feature matrix, a twenty-third feature matrix for characterizing the correlation between the audio frames of the audio to be trained; determine, based on the twenty-third feature matrix, a twenty-fourth feature matrix for characterizing the global features of each audio frame of the audio to be trained; perform recognition based on the twenty-second and twenty-fourth feature matrices to obtain a predicted recognition result corresponding to the audio to be trained; and train an audio recognition model based on the predicted recognition result and label corresponding to the audio to be trained.
[0136] In a specific implementation, the first acquisition unit 1001 and the first processing unit 1002 described in the embodiment of the present invention may also execute other implementation methods described in the embodiment of the training method of the audio recognition model provided by the embodiment of the present invention, which will not be repeated here.
[0137] See Figure 11 , Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 11 As shown, the electronic device 1100 includes a transceiver 1101, a processor 1102, and a memory 1103. These are connected via a bus 1104. The memory 1103 is used to store computer programs and data, and can transmit the data stored in the memory 1103 to the processor 1102.
[0138] The electronic device 1100 may be the audio recognition device 900 or the audio recognition model training device 1000;
[0139] When the electronic device 1100 is the audio recognition device 900, the processor 1102 is configured to read the computer program in the memory 1103 and perform the following operations:
[0140] Controlling the transceiver 1101 to obtain the audio to be recognized;
[0141] Perform feature extraction on multiple audio frames in the audio to be recognized to obtain a first feature matrix;
[0142] Determining, based on the first feature matrix, a second feature matrix for characterizing a correlation relationship between any two audio frames in the plurality of audio frames;
[0143] Determining, based on the second feature matrix, a third feature matrix for characterizing an association relationship between each audio frame and the plurality of audio frames;
[0144] Recognition is performed based on the first characteristic matrix and the third characteristic matrix to obtain a recognition result of the audio to be recognized.
[0145] In one embodiment of the present application, in determining a third feature matrix for characterizing an association relationship between each audio frame and multiple audio frames based on the second feature matrix, the processor 1102 is specifically configured to perform the following operations:
[0146] Performing at least one feature aggregation process on the feature vector of each audio frame in the second feature matrix to obtain a feature matrix corresponding to each feature aggregation process;
[0147] The feature matrices corresponding to each feature aggregation process are fused to obtain a third feature matrix;
[0148] The at least one feature aggregation process includes at least one of the following processes: summation process, mean value process, and standard deviation process.
[0149] In one embodiment of the present application, in terms of performing recognition based on the first feature matrix and the third feature matrix to obtain a recognition result of the audio to be recognized, the processor 1102 is specifically configured to perform the following operations:
[0150] Determining a third feature weight corresponding to each audio frame based on the third feature matrix;
[0151] Based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix is processed to obtain a ninth feature matrix;
[0152] Fusing the first feature matrix and the ninth feature matrix to obtain a second fused feature matrix;
[0153] Recognition is performed based on the second fused feature matrix to obtain a recognition result of the audio to be recognized.
[0154] In one embodiment of the present application, in extracting features from multiple audio frames in the audio to be recognized to obtain a first feature matrix, the processor 1102 is specifically configured to perform the following operations:
[0155] Extract features from multiple audio frames in the audio to be recognized to obtain an initial feature matrix;
[0156] Based on the initial characteristic matrix, determining a plurality of fourth characteristic matrices corresponding to the plurality of channels;
[0157] A first characteristic matrix is determined based on a plurality of fourth characteristic matrices corresponding to the plurality of channels.
[0158] In one embodiment of the present application, in determining the first characteristic matrix based on a plurality of fourth characteristic matrices corresponding to a plurality of channels, the processor 1102 is specifically configured to perform the following operations:
[0159] Perform linear processing on each fourth characteristic matrix to obtain a fifth characteristic matrix corresponding to each channel;
[0160] Perform dimensionality reduction on each fourth characteristic matrix to obtain a sixth characteristic matrix corresponding to each channel;
[0161] Determining a first feature weight of each audio frame under each channel based on a sixth feature matrix corresponding to each channel;
[0162] Based on the first feature weight of each audio frame under each channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel is processed to obtain a seventh feature matrix corresponding to each channel;
[0163] The fourth feature matrix corresponding to each channel and the seventh feature matrix corresponding to each channel are fused to obtain a first fused feature matrix corresponding to each channel;
[0164] A first feature matrix is determined based on the first fused feature matrix corresponding to each channel.
[0165] In one embodiment of the present application, in determining the first feature matrix based on the first fused feature matrix corresponding to each channel, the processor 1102 is specifically configured to perform the following operations:
[0166] Determine a second feature weight corresponding to each channel based on the first fusion feature matrix corresponding to each channel;
[0167] Based on the second feature weight corresponding to each channel, the first fusion feature matrix corresponding to each channel is weighted to obtain an eighth feature matrix corresponding to each channel;
[0168] A first characteristic matrix is determined based on the eighth characteristic matrix corresponding to each channel.
[0169] In one embodiment of the present application, in terms of performing recognition based on the second fused feature matrix to obtain a recognition result of the audio to be recognized, the processor 1102 is specifically configured to perform the following operations:
[0170] Perform graph convolution on the second fused feature matrix to obtain the tenth feature matrix;
[0171] Performing average pooling on the feature vectors of each audio frame in the tenth feature matrix to obtain an eleventh feature matrix;
[0172] Performing linear mapping on the eigenvector of each audio frame in the eleventh eigenmatrix to obtain a twelfth eigenmatrix;
[0173] Determining a fourth feature weight corresponding to each audio frame based on the twelfth feature matrix;
[0174] Based on the fourth feature weight corresponding to each audio frame, weight the feature vector corresponding to each audio frame in the tenth feature matrix to determine a thirteenth feature matrix;
[0175] Recognition is performed based on the thirteenth characteristic matrix to obtain a recognition result corresponding to the audio to be recognized.
[0176] In a specific implementation, the transceiver 1101 and the processor 1102 described in the embodiment of the present invention may also execute other implementations described in the embodiment of the audio recognition method provided in the embodiment of the present invention, which will not be described in detail here.
[0177] When the electronic device 1100 is a training device 1000 for an audio recognition model, the processor 1102 is configured to read the computer program in the memory 1103 and perform the following operations:
[0178] Controlling the transceiver 1101 to obtain the audio to be trained;
[0179] Perform feature extraction on multiple audio frames in the training audio to obtain a twenty-second feature matrix;
[0180] Determine, based on the twenty-second feature matrix, a twenty-third feature matrix for characterizing association relationships between audio frames of the audio to be trained;
[0181] Determine, based on the twenty-third feature matrix, a twenty-fourth feature matrix for characterizing global features of each audio frame of the audio to be trained;
[0182] Perform recognition based on the twenty-second characteristic matrix and the twenty-fourth characteristic matrix to obtain a predicted recognition result corresponding to the audio to be trained;
[0183] The audio recognition model is trained based on the predicted recognition results and labels corresponding to the audio to be trained.
[0184] In a specific implementation, the transceiver 1101 and the processor 1102 described in the embodiment of the present invention may also execute other implementation methods described in the embodiment of the training method of the audio recognition model provided by the embodiment of the present invention, which will not be repeated here.
[0185] Specifically, the transceiver 1101 may be Figure 9 The acquisition unit 901 of the audio recognition device 900 or Figure 10 The first acquisition unit 1001 of the audio recognition model training device 1000 of the embodiment, the processor 1102 can be Figure 9 The processing unit 902 of the audio recognition device 900 of the embodiment or Figure 10 The first processing unit 1002 of the audio recognition model training apparatus 1000 of the embodiment.
[0186] It should be understood that the electronic device in this application can be an audio recognition device or a training device for an audio recognition model. Both the audio recognition device and the training device for an audio recognition model can be servers, such as independent physical servers, or server clusters or distributed systems composed of multiple physical servers. It can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms. This application does not make specific limitations. The above-mentioned electronic devices are only examples, not exhaustive, and include but are not limited to the above-mentioned electronic devices.
[0187] It should be understood that the embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any audio recognition method or audio recognition model training method recorded in the above method embodiments.
[0188] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any audio recognition method or audio recognition model training method described in the above method embodiments.
[0189] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0190] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0191] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0192] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0193] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software program modules.
[0194] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0195] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0196] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An audio recognition method, characterized in that: The method comprises: Perform feature extraction on multiple audio frames in the audio to be recognized to obtain a first feature matrix; determining, based on the first feature matrix, a second feature matrix for characterizing an association relationship between any two audio frames among the plurality of audio frames; Determining, based on the second feature matrix, a third feature matrix for characterizing an association relationship between each audio frame and the plurality of audio frames; Based on the third feature matrix, a third feature weight corresponding to each audio frame is determined; based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix is processed to obtain a ninth feature matrix; the first feature matrix and the ninth feature matrix are fused to obtain a second fused feature matrix; and recognition is performed based on the second fused feature matrix to obtain a recognition result of the audio to be recognized.
2. The method according to claim 1, characterized in that The determining, based on the second feature matrix, a third feature matrix for characterizing the association relationship between each audio frame and the multiple audio frames includes: performing at least one feature aggregation process on a feature representation of an association relationship between each audio frame in the second feature matrix and each audio frame in the plurality of audio frames, to obtain a feature matrix corresponding to each feature aggregation process; The feature matrices corresponding to each feature aggregation process are fused to obtain the third feature matrix.
3. The method according to claim 1 or 2, characterized in that The step of extracting features from multiple audio frames in the audio to be recognized to obtain a first feature matrix includes: Performing feature extraction on the multiple audio frames in the audio to be recognized to obtain an initial feature matrix; Based on the initial characteristic matrix, determining a plurality of fourth characteristic matrices corresponding to a plurality of channels; The first characteristic matrix is determined based on the plurality of fourth characteristic matrices corresponding to the plurality of channels.
4. The method according to claim 3, characterized in that The determining the first characteristic matrix based on the plurality of fourth characteristic matrices corresponding to the plurality of channels includes: Perform linear processing on each fourth characteristic matrix to obtain a fifth characteristic matrix corresponding to each channel; Perform dimensionality reduction on each fourth characteristic matrix to obtain a sixth characteristic matrix corresponding to each channel; Determining a first feature weight of each audio frame under each channel based on a sixth feature matrix corresponding to each channel; Based on the first feature weight of each audio frame under each channel, the feature vector corresponding to each audio frame in the fifth feature matrix corresponding to each channel is processed to obtain a seventh feature matrix corresponding to each channel; The fourth feature matrix corresponding to each channel and the seventh feature matrix corresponding to each channel are fused to obtain a first fused feature matrix corresponding to each channel; Based on the first fusion feature matrix corresponding to each channel, the first feature matrix is determined.
5. The method according to claim 4, characterized in that The determining the first feature matrix based on the first fusion feature matrix corresponding to each channel includes: Determine a second feature weight corresponding to each channel based on the first fusion feature matrix corresponding to each channel; Based on the second feature weight corresponding to each channel, the first fusion feature matrix corresponding to each channel is processed to obtain an eighth feature matrix corresponding to each channel; The first characteristic matrix is determined based on the eighth characteristic matrix corresponding to each channel.
6. The method according to claim 1, characterized in that The performing recognition based on the second fusion feature matrix to obtain a recognition result of the audio to be recognized includes: Performing graph convolution processing on the second fused feature matrix to obtain a tenth feature matrix; performing average pooling on the feature vector of each audio frame in the tenth feature matrix to obtain an eleventh feature matrix; Performing linear mapping on the eigenvector of each audio frame in the eleventh eigenmatrix to obtain a twelfth eigenmatrix; Determining a fourth feature weight corresponding to each audio frame based on the twelfth feature matrix; Based on the fourth feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the tenth feature matrix is processed to determine a thirteenth feature matrix; Recognition is performed based on the thirteenth characteristic matrix to obtain a recognition result of the audio to be recognized.
7. An audio recognition device, characterized in that: The audio recognition device includes: an acquisition unit and a processing unit; The acquisition unit is used to acquire the audio to be recognized; The processing unit is configured to perform feature extraction on multiple audio frames in the audio to be recognized to obtain a first feature matrix; determine, based on the first feature matrix, a second feature matrix for characterizing the association relationship between any two audio frames in the multiple audio frames; determine, based on the second feature matrix, a third feature matrix for characterizing the association relationship between each audio frame and the multiple audio frames; determine, based on the third feature matrix, a third feature weight corresponding to each audio frame; process, based on the third feature weight corresponding to each audio frame, the feature vector corresponding to each audio frame in the first feature matrix to obtain a ninth feature matrix; fuse the first feature matrix and the ninth feature matrix to obtain a second fused feature matrix; and perform recognition based on the second fused feature matrix to obtain a recognition result of the audio to be recognized.
8. An electronic device, characterized in that: include: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice recognition method and device, computer equipment and readable storage medium
CN113763927A
Voice processing method and device, equipment and storage medium
CN113823313A