Speech audio analysis method and device, electronic device and readable storage medium

By using the tensor weight value of the target character in speech audio analysis to determine the endpoint audio frame, the problem of environmental interference in the prior art is solved, and more accurate and efficient speech endpoint detection is achieved.

CN113971963BActive Publication Date: 2025-05-13MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010709120.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-22
Publication Date
2025-05-13
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

In the prior art, the method of voice endpoint detection through time domain or frequency domain characteristics is greatly disturbed by the environment, resulting in poor detection effect.

Method used

By determining the N weight values ​​corresponding to the tensor of the target character in the speech audio to be analyzed, selecting the M weight values ​​whose weight values ​​are larger than or equal to the preset threshold value, and thus determining the endpoint audio frame corresponding to the target character.

Benefits of technology

This method can avoid environmental interference during the endpoint detection process and improve the accuracy and efficiency of voice endpoint detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971963B_ABST
    Figure CN113971963B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for analyzing speech audio, an electronic device, and a readable storage medium, wherein the method includes: determining the tensor of a target character in the speech audio to be analyzed; obtaining N weight values ​​corresponding to the tensor of the target character, wherein the weight values ​​are used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, wherein N is a positive integer; selecting the top M weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, wherein M is a positive integer, and M is less than N; determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values. Through the present application, the problem that the method of performing endpoint detection through time domain or frequency domain features in the prior art is greatly affected by environmental interference, resulting in poor effect of speech endpoint detection is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a method and device for analyzing speech audio, an electronic device, and a readable storage medium. Background Art

[0002] As a means of human-computer interaction, voice endpoint detection is of great significance in freeing human hands. At the same time, there are various background noises in the working environment, which will seriously reduce the quality of voice and thus affect the effect of voice applications, such as reducing the recognition rate. Uncompressed voice audio will cause large network traffic in network interactive applications, thereby reducing the success rate of voice applications.

[0003] At present, endpoint detection technology is mainly based on some time domain or frequency domain features of speech for distinction. Time domain parameter endpoint detection is based on characteristic parameters in the time domain for distinction, mainly including time domain energy size, time domain average zero crossing rate, short-term correlation analysis, energy change rate, logarithmic energy, subband energy, GMM (Gaussian Mixture Model) hypothesis test and other methods. The noise resistance of frequency domain parameters will be better than that of time domain, but the calculation consumption is also high. The mainstream technology includes spectral entropy, frequency domain subband, adaptive wavelet, fundamental frequency and other methods. The existing method of endpoint detection through time domain or frequency domain features is greatly affected by environmental interference, resulting in poor detection effect. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a speech audio analysis method and device, an electronic device and a readable storage medium, which can solve the problem that the method of performing endpoint detection through time domain or frequency domain features in the prior art is greatly affected by environmental interference, resulting in poor speech endpoint detection effect.

[0005] In order to solve the above technical problems, this application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for analyzing speech audio, comprising: determining a tensor of a target character in the speech audio to be analyzed; obtaining N weight values ​​corresponding to the tensor of the target character, the weight values ​​being used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, wherein N is a positive integer; selecting M weight values ​​with the top M largest weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, wherein M is a positive integer and M is less than N; determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values.

[0007] In a second aspect, an embodiment of the present application provides a speech audio analysis device, comprising: a first determination module, used to determine the tensor of a target character in the speech audio to be analyzed; an acquisition module, used to obtain N weight values ​​corresponding to the tensor of the target character, the weight value being used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, wherein N is a positive integer; a processing module, used to select M weight values ​​with the top M largest weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, wherein M is a positive integer and M is less than N; a second determination module, used to determine the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values.

[0008] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] In an embodiment of the present application, N weight values ​​corresponding to the tensor of the target character in the speech audio to be analyzed are determined, and then from the N weight values, M weight values ​​with the top M largest weight values ​​are selected, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, and the endpoint audio frame corresponding to the target character is determined from the M audio frames corresponding to the M weight values. It can be seen that the endpoint audio frame corresponding to the target character is determined by the weight value corresponding to the tensor of the target character. Since the endpoint audio frame of the target character is determined by the weight value of the audio frame, it will not be affected by environmental interference during the endpoint detection process, thereby solving the problem that the method of performing endpoint detection through time domain or frequency domain features in the prior art is subject to greater environmental interference, resulting in poor effect of speech endpoint detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic diagram of a model of an attention mechanism in an embodiment of the present application;

[0012] Figure 2 is a schematic diagram of the calculation process of Attention in an embodiment of the present application;

[0013] Figure 3 is a flow chart of a method for analyzing speech audio according to an embodiment of the present application;

[0014] Figure 4is a schematic diagram of a Transformer model of an embodiment of the present application;

[0015] Figure 5 It is a schematic diagram of the structure of Scaled dot-product attention in an embodiment of the present application;

[0016] Figure 6 It is a schematic diagram of the structure of Multi-head attention in an embodiment of the present application;

[0017] Figure 7 It is a structural diagram of the speech audio analysis device of an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0019] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the audio used in this way can be interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0020] See also Figure 1 , Figure 1 is a flow chart of the method for analyzing speech audio according to an embodiment of the present application, such as Figure 1 As shown, the steps of the method include:

[0021] Step S102, determining the tensor of the target character in the speech audio to be analyzed;

[0022] It should be noted that the speech audio to be analyzed may include multiple target characters, each of which has a corresponding tensor, such as a tensor corresponding to a target character may be (1, 512). Of course, this is only an example of the specific value of the tensor corresponding to a target character, and the tensors of other target characters may also have other values ​​according to actual conditions.

[0023] Step S104, obtaining N weight values ​​corresponding to the tensor of the target character, where the weight value is used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, where N is a positive integer;

[0024] In a specific application scenario, assuming that the target character corresponds to 90 audio frames, it is necessary to obtain the weight distribution of the tensor of the target character on the corresponding 90 audio frames, that is, the value of N is 90. It should be noted that the value of N is 90 for example only, and may also be 80 or 100, that is, it may be set or determined accordingly according to the actual situation.

[0025] Step S106, selecting M weight values ​​with the largest weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, where M is a positive integer and M is less than N;

[0026] Among them, it should be noted that there are many ways to select M weight values ​​with the largest weight values ​​from N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold. For example: sort the N weight values ​​in descending order, and then select M weight values ​​with the largest weight values, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold; or, sort the N weight values ​​in ascending order, and then select M weight values ​​with the largest weight values ​​from the sorting results, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold; or, first select the largest one from the N weight values, and then select the largest one from the remaining weight values, until M weight values ​​with the largest weight values ​​are selected, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold.

[0027] The following is an example of combining these three methods to illustrate the method of selecting the M weight values ​​with the largest weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold;

[0028] 1) Descending sorting means sorting the N weight values ​​from large to small. Therefore, the M weight values ​​selected in sequence are the selected weight values ​​after sorting from large to small. For example, if the value of N is 10, the corresponding number of audio frames is also 10. For example, the 10 audio frames are: a1, a2, a3, a4, a5, a6, a7, a8, a9, a10, a111, a2, a3, a4, a5, a6, a7, a8, a9, a112, a113, a114, a115, a116, a117, a118, a119, a120, a121, a122, a123, a124, a125, a126, a127, a128, a129, a130, a131 10 For example, the 10 audio frames are sorted in time series, and are sorted in time series by the digital labels of the audio frames, that is, a1 is the audio frame that is ranked first in the time series, and a 10is the audio frame that is ranked last in the time series; further, if the weight distribution corresponding to the above 10 audio frames is: a1=0.01, a2=0.01, a3=0.11, a4=0.21, a5=0.31, a6=0.21, a7=0.11, a8=0.01, a9=0.01, a 10 =0.01. The result of descending sorting according to the weight is: a5≥a4≥a6≥a3≥a7≥a1≥a2≥a8≥a9≥a 10 In the embodiment of the present application, if the preset threshold is set to 0.84, the selected audio frames are a5, a4, a6, a3, that is, a3+a4+a5+a6=0.84≥0.84, that is, the value of M is 4.

[0029] 2) Sort the N weight values ​​in ascending order (ascending order means sorting the N weight values ​​from small to large), and then select the M weight values ​​with the largest weight values ​​from the sorting results, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold. For example, the value of N is 10, that is, the corresponding number of audio frames is also 10. Take 10 audio frames as: a1, a2, a3, a4, a5, a6, a7, a8, a9, a10, a111, a12, a3, a4, a5, a6, a7, a8, a9, a112, a13, a4, a5, a6, a7, a8, a9, a10, a113, a114, a124, a13, a13, a14, a15, a16, a17, a8, a9, a10 10 For example, the 10 audio frames are sorted in time series, and are sorted in time series by the digital labels of the audio frames, that is, a1 is the audio frame that is ranked first in the time series, and a 10 is the audio frame that is ranked last in the time series; further, if the weight distribution corresponding to the above 10 audio frames is: a1=0.01, a2=0.01, a3=0.11, a4=0.21, a5=0.31, a6=0.21, a7=0.11, a8=0.01, a9=0.01, a 10 = 0.01. The result of sorting them in ascending order according to the weight is: a 10 ≤a9≤a8≤a2≤a1≤a7≤a3≤a6≤a4≤a5. In the embodiment of the present application, if the preset threshold is set to 0.84, the selected audio frames are a5, a4, a6, a3, that is, a3+a4+a5+a6=0.84≥0.84, that is, the value of M is 4.

[0030] 3) First select the largest one from the N weight values, and then select the largest one from the remaining weight values, until the M weight values ​​with the largest weight values ​​are selected, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold. For example, if the value of N is 10, the corresponding number of audio frames is also 10. Take 10 audio frames as: a1, a2, a3, a4, a5, a6, a7, a8, a9, a10, a111, a12, a3, a4, a5, a6, a7, a8, a9, a112, a123, a13, a14, a15, a16, a17, a18, a19, a10 10For example, the 10 audio frames are sorted in time series, and are sorted in time series by the digital labels of the audio frames, that is, a1 is the audio frame that is ranked first in the time series, and a 10 is the audio frame that is ranked last in the time series; further, if the weight distribution corresponding to the above 10 audio frames is: a1=0.01, a2=0.01, a3=0.11, a4=0.21, a5=0.31, a6=0.21, a7=0.11, a8=0.01, a9=0.01, a 10 =0.01; and the preset threshold is set to 0.84. Therefore, a5 with the largest weight value is selected first. Since the weight value of a5 is less than the preset threshold, a4 with the largest weight value is selected from the remaining audio frames; it is determined whether a5+a4 is greater than or equal to the preset threshold. Since the sum of a5+a4 is less than 0.84, a6 with the largest weight value is selected from the remaining audio frames; it is determined whether a5+a4+a6 is greater than or equal to the preset threshold. Since the sum of a5+a4+a6 is less than 0.84, a3 with the largest weight value is selected from the remaining audio frames again. It is determined that the sum of a5+a4+a6+a3 is greater than or equal to the preset threshold. Since a5+a4+a6+a3=0.84, the condition of being greater than or equal to the preset threshold is met, and the selection is stopped. That is, in this application scenario, the selected audio frames are a5, a4, a6, a3, and the value of M is 4.

[0031] Step S108, determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values.

[0032] Through the above-mentioned steps S102 to S108 of the embodiment of the present application, by determining the N weight values ​​corresponding to the tensor of the target character in the speech audio to be analyzed, and then selecting the top M weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to the preset threshold, and determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values. It can be seen that the endpoint audio frame corresponding to the target character is determined by the weight value corresponding to the tensor of the target character. Since the endpoint audio frame of the target character is determined by the weight value of the audio frame, it will not be affected by environmental interference during the endpoint detection process, thereby solving the problem that the method of performing endpoint detection through time domain or frequency domain features in the prior art is greatly affected by environmental interference, resulting in poor effect of speech endpoint detection.

[0033] Optionally, the method for determining the tensor of the target character in the speech audio to be analyzed involved in step S102 of the embodiment of the present application may further be:

[0034] Step S102-11, obtaining the target character in the speech audio to be analyzed;

[0035] The step S102-11 may be further implemented in the following manner:

[0036] Step S11, extracting an audio frame feature vector from the speech audio to be analyzed;

[0037] Specifically, the audio frame feature vector may be an FBank feature (Filter bank features) or an MFCC (Mel Frequency Cepstrum Coefficient) feature.

[0038] Step S12, inputting the audio frame feature vector into a first preset model to obtain a tensor of the audio frame;

[0039] Specifically, the audio frame feature vector may be (100, 40, 1), and the tensor of the audio frame obtained after inputting it into the first preset model is (100, 512). In this embodiment, the first preset model may be a residual network, and other models may also be selected according to actual needs.

[0040] Step S13, input the tensor of the audio frame into the second preset model to obtain a character sequence, and determine any character in the character sequence as the target character. Wherein, after the tensor of the audio frame is input into the second preset model, a character sequence can be obtained. If the tensor of the above audio frame is (100, 512), the number of characters in the character sequence obtained after inputting into the second preset model is 100, that is, the sequence of target characters is: b1, b2, b3, ... b 100 . Among them, the numbers used to identify their character sequences are different; of course, in some scenarios, the corresponding sorting relationship can also be represented by the digital label of the character according to actual needs. In this embodiment, the second preset model is a neural network model with an attention mechanism. Optionally, the neural network model with an attention mechanism involved in the embodiments of the present application can be a transformer model, and of course it can also be other neural network models with an attention mechanism, such as: a seq2seq model.

[0041] Among them, the attention mechanism in the embodiment of the present application can be implemented by an attention mechanism model (Attention model), such as Figure 2As shown in the figure, the Attention model is an alignment model that outputs a word in the Target sentence and each word in the input Source sentence, that is, imagine the constituent elements in the Source as a series of<Key,Value> The audio pair is composed of a certain element Query in the Target. By calculating the similarity or correlation between the Query and each Key, the weight coefficient of the Value corresponding to each Key is obtained, and then the Value is weighted and summed to obtain the final Attention value. Therefore, the Attention mechanism is to perform a weighted sum of the Value values ​​of the elements in the Source, and the Query and Key are used to calculate the weight coefficient of the corresponding Value. The Attention mechanism can be further expressed by formula (1):

[0042]

[0043] Among them, Lx=||Source|| represents the length of Source.

[0044] The specific calculation process of Attention includes two processes: 1) Calculate the weight coefficient based on the query and key; 2) Perform weighted summation of the value based on the weight coefficient. Among them, process 1) can be further divided into two stages: the first stage calculates the similarity or correlation between the query and the key. The second stage normalizes the original score of the first stage; in this way, the calculation process of Attention can be abstracted as follows: Figure 3 The three stages are shown.

[0045] In the first stage, different functions and calculation mechanisms can be introduced to calculate the similarity or correlation between the query and a key. The methods for calculating the similarity and correlation between the two include: calculating the dot product of their vectors, calculating the cosine similarity of their vectors, or evaluating by introducing an additional neural network. The similarity calculation formula is as follows: Formulas (2) to (4):

[0046] Dot product: Similarity(Qurey,key i )=Qurey·key i (2)

[0047] Cosine Similarity:

[0048] MLP network: Similarity (Qurey, key i )=MLP(Qurey,key i ) (4)

[0049] The scores generated in the first stage have different numerical ranges depending on the specific generation method. In the second stage, a calculation method similar to SoftMax is introduced to convert the scores of the first stage into numerical values. On the one hand, normalization can be performed to organize the original calculated scores into a probability distribution in which the sum of the weights of all elements is 1; on the other hand, the weights of important elements can be more highlighted through the internal mechanism of SoftMax.

[0050] The calculation result ai in the second stage is the weight coefficient corresponding to value, and then the weighted sum is performed to get the Attention value, which can be obtained by the following formula (5):

[0051]

[0052] It should be noted that the attention mechanism used in the transformer model of the embodiment of the present application mainly includes: Scaled dot-product attention and Multi-head attention, where:

[0053] Scaled dot-product attention can be expressed by the following formula (6):

[0054]

[0055] The above formula (6) means that each query-key will undergo a dot multiplication process and be divided by a constant of dimension to prevent the value from being too large; then, they are normalized using softmax; and finally, they are multiplied by V(values) to be used as the attention vector.

[0056] Step S102-12, performing word embedding processing on the target character to obtain a tensor of the target character;

[0057] Among them, word embedding processing refers to converting a word into a vector representation. In the embodiment of the present application, the target character is converted into a corresponding tensor.

[0058] Optionally, in the embodiment of the present application, the method for obtaining N weight values ​​corresponding to the tensor of the target character involved in step S104 in the embodiment of the present application may further be:

[0059] Step S104-11, obtaining the similarity between the tensor of the target character and the tensor of the audio frame corresponding to the target character;

[0060] For the acquisition of similarity in step S104-11, in the embodiment of the present application, based on the above-mentioned Attention model, it can be calculated by one of the above-mentioned formulas (2) to (4); if the above-mentioned formula (2) is taken as an example, and it is assumed that there are 90 audio frames, the key and value in the Attention model are tensors of audio frames, the key tensor is (90, 512) in size, and the value tensor is (90, 512) in size; the query is a tensor of a decoded character (equivalent to the tensor of the target character), and the size of the query tensor is (1, 512). Therefore, by using formula (2) to calculate the similarity, the similarity of the audio frame is dots = query*key, that is, the size of the similarity dots tensor of the audio frame is (1, 90), which represents the relationship between the current query character and the 90 audio frames.

[0061] Step S104 - 12 , normalizing the similarity to obtain N weight values ​​corresponding to the tensor of the target character.

[0062] In the specific application scenario of the embodiment of the present application, the similarity can be normalized by using softmax to normalize all values ​​to between 0 and 1 to obtain a1, a2, ..., a corresponding to the current query character. 90 The weight of . Among them, SoftMax adopts the following formula (7):

[0063]

[0064] That is to say, for the normalization process involved in the above step S104-12, the similarity value is normalized to obtain the corresponding weight value, and the sum of the weight values ​​is 1. For example, the target character corresponds to 7 audio frames, and the weight values ​​after normalization of the similarity between the tensor of the target character and the tensor of the audio frame corresponding to the target character are: a1=0.07, a2=0.1, a3=0.3, a4=0.4, a5=0.03, a6=0.04, a7=0.06; and after sorting from large to small, they are: a4=0.4, a3=0.3, a2=0.1, a1=0.07, a7=0.06, a6=0.04, a5=0.03. The seven audio frames are sorted in time series by digital labels of the audio frames, that is, a1 is the audio frame that is ranked first in the time series, and a7 is the audio frame that is ranked last in the time series.

[0065] Optionally, in the embodiment of the present application, the method for determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values ​​involved in step S108 may further be:

[0066] Step S108-11, selecting the audio frame that is the frontmost in the time sequence from the determined M audio frames as the start frame of the endpoint audio frame corresponding to the target character;

[0067] Step S108-12, selecting the audio frame with the latest time sequence from the determined M audio frames as the end frame of the endpoint audio frame corresponding to the target character.

[0068] For step S108-11 and step S108-12, take the weight values ​​of the audio frames corresponding to the target characters in step S104-12 above, and sort the results from large to small as an example, that is, the sorting results are: a4=0.4, a3=0.3, a2=0.1, a1=0.07, a7=0.06, a6=0.04, a5=0.03; wherein, the 7 audio frames are sorted in time series, and are sorted in time series by the digital labels of the audio frames, that is, a1 is the audio frame that is sorted the most in the time series, and a7 is the audio frame that is sorted the most in the time series. If the preset weight value is set to 0.9, the audio frames whose sum of weight values ​​after sorting is greater than 0.9 are: a4=0.4, a3=0.3, a2=0.1, a1=0.07, a7=0.06. Therefore, a1 is the audio frame at the front of the time sequence, so a1 is the starting frame in the endpoint audio frame corresponding to the target character; a7 is the audio frame at the back of the time sequence, so a7 is the ending frame in the endpoint audio frame corresponding to the target character.

[0069] Optionally, if the speech audio to be analyzed includes multiple target characters, the method steps in the embodiment of the present application further include:

[0070] Step S110, determining the audio frame at the front of the time sequence from the start frames of the endpoint audio frames corresponding to the multiple target characters as the start frame of the endpoint audio frame corresponding to the speech audio to be analyzed;

[0071] Step S112, determining the audio frame with the latest time sequence from the end frames of the endpoint audio frames corresponding to the multiple target characters as the end frame of the endpoint audio frames corresponding to the speech audio to be analyzed.

[0072] Taking the example that the number of target characters in the speech audio to be analyzed is 2, for example, there are the first target character and the second target character, and there are 10 audio frames of the endpoint audio frames corresponding to the two characters. When identifying the first target character, the endpoint audio frames are the 2nd audio frame and the 4th audio frame, and when identifying the second target character, the endpoint audio frames are the 3rd audio frame and the 8th audio frame; then the starting frame of the final endpoint audio frame of the speech audio to be analyzed is the 2nd audio frame, and the ending frame of the endpoint audio frame is the 8th audio frame.

[0073] The present application is described below with reference to the specific implementation of the embodiment of the present application; the specific implementation of the embodiment of the present application provides a method for speech endpoint detection based on an attention mechanism, and the specific implementation is described by taking the first preset model as a residual network and the second preset model as a neural network model Transformer model with an attention mechanism as an example, and the method includes:

[0074] Step S202, obtaining FBank features of the audio;

[0075] Step S204, extracting FBank features through the residual network Resnet;

[0076] Among them, in a specific application scenario, step S404 can be: after extracting the Fbank features of a piece of audio, for example, the Fbank feature dimension is (100, 40, 1), it is input into the residual network. Optionally, the residual network can use 18 layers, and after the residual network is processed, the result of (100, 512) is obtained.

[0077] Step S206, inputting the extracted features into the Transformer model for speech recognition;

[0078] Step S208, identifying each character (or phoneme) corresponding to the audio;

[0079] Step S210, extracting each attention range according to the recognized characters (or phonemes);

[0080] Step S212, integrating all recognition results to determine the start and end positions of the endpoints.

[0081] For the above steps S206 to S212, in a specific application scenario, the result (100, 512) can be input into the Transformer model. Since the Transformer model uses a network model of the attention mechanism, such as Figure 4 As shown in the figure, in the first part of the encoder and decoder, multi-head attention uses a self-attention mechanism network to learn the internal relationship of speech features and the internal relationship of speech tags after Resnet processing. In the second part of the Transformer model decoder, multi-head attention uses the keys and values ​​from the encoder and the queries from the decoder.

[0082] See also Figure 5 , Figure 5It is a schematic diagram of the structure of the Scaled dot-product attention in the embodiment of the present application. It should be noted that the expression ability of the scaled dot-product attention network is still a little simple. Further, a multi-head attention mechanism is proposed in the embodiment of the present application, such as Figure 6 As shown in the figure, Multi-head attention projects Q, K, and V through h different linear transformations, and finally concatenates the different attention results, as shown in the following formula (8):

[0083]

[0084] from Figure 5 and Figure 6 It can be seen that scaled dot-product attention is included in multi-head attention. If there are 8 heads in multi-head attention, Q, K, and V are projected through 8 different linear transformations, and finally the different attention results are spliced ​​together.

[0085] from Figure 4 In the transformer model structure in Fig. 1, we can see that there is a multi-head attention mechanism in the encoder structure. The multi-head attention mechanism is a multi-head self-attention mechanism that encodes information, and self-attention takes the same Q, K, and V.

[0086] from Figure 4From the transformer model structure in , we can see that there are two multi-head attention mechanisms in the decoder structure. The first multi-head attention mechanism is a multi-head self-attention mechanism for decoding information. Self-attention takes Q, K, and V as the same; the second multi-head attention mechanism uses K (keys), V (values) from the encoder and Q (queries) from the decoder. According to the specific calculation process of the Attention mechanism mentioned above, in the first two stages, the similarity or correlation between the Query and a Key is calculated, and the weight coefficient ai, that is, the weight coefficient corresponding to the value, is calculated. i The original calculated scores are organized into a probability distribution where the sum of the weights of all elements is 1. If there are 8 heads, the weights and probability distributions of the 8 heads are added together to obtain a probability distribution where the sum of the weights is 8.

[0087] The following is an example of one head. First, the weight distribution λ of the current single character (or phoneme) in the current decoder's Q (queries) on the encoder's K (keys) is obtained, and the weight threshold T = 0.95 (if there are N heads, just multiply each term by N). The weight distribution λ is:

[0088] λ={a1,a2,a3…a n}

[0089] Among them, a1, a2, a3…a n It is a decimal whose weights sum to 1 after SoftMax calculation, and n is the number of temporal audio frames on K (keys) of the encoder. It can be understood as the current single character (or phoneme) is obtained on which audio frames the weights are distributed on the encoded information.

[0090] For example: n = 10 (i.e., there are 10 audio frames), a1 = 0.01, a2 = 0.01, a3 = 0.011, a4 = 0.21, a5 = 0.31, a6 = 0.21, a7 = 0.11, a8 = 0.01, a9 = 0.01, a 10 =0.01, among which, a1+a2+a3+a4+a5+a6+a7+a8+a9+a 10 =1.0.

[0091] First, for a1, a2, a3…a 10 Sort from largest to smallest and we get: a5≥a4≥a6≥a3≥a7≥a1≥a2≥a8≥a9≥a 10 , take the first m weights and all audio frames that are greater than or equal to the weight threshold T we set (i.e. a3+a4+a5+a6+a7=0.95≥T), it can be seen that the current single character (or phoneme) is generated by the attention mechanism focusing on the 3rd, 4th, 5th, 6th, and 7th audio frames, a3+a4+a5+a6+a7=0.95≥T, T is 0.95. Therefore, the endpoints of the current single character (or phoneme) are set as follows: the start frame is the 3rd frame, and the end frame is the 7th frame.

[0092] It should be noted that in order to find the audio frame range where the current single character (or phoneme) is produced, the attention mechanism focuses on it. Therefore, if there are non-focused audio frames between the longest and shortest audio frames where a single character (or phoneme) is focused on, the speech group between the longest and shortest audio frames will also be taken as the focused audio frames.

[0093] After the Transformer model outputs all characters (or phonemes), the starting endpoints of all characters (or phonemes) are merged to obtain the starting endpoint of the entire audio frame.

[0094] It can be seen that the specific implementation method of the embodiment of the present application extracts valid audio frames through the attention mechanism and then maps them to the voice endpoint time; since it can be accurate to the audio frame, the voice endpoint can be detected more accurately.

[0095] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0096] See also Figure 7 , Figure 7 is a schematic diagram of the structure of the speech audio analysis device in the embodiment of the present application, such as Figure 7 As shown, the device comprises:

[0097] A first determination module 72, used to determine the tensor of the target character in the speech audio to be analyzed;

[0098] An acquisition module 74 is used to acquire N weight values ​​corresponding to the tensor of the target character, where the weight value is used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, where N is a positive integer;

[0099] The processing module 76 is used to select M weight values ​​with the largest weight values ​​from the N weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, where M is a positive integer and M is less than N;

[0100] The second determination module 78 is used to determine the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values.

[0101] Optionally, the first determination module 72 in the embodiment of the present application may further include: a first acquisition unit, used to obtain target characters in the speech audio to be analyzed; and a first processing unit, used to perform word embedding processing on the target characters to obtain a tensor of the target characters.

[0102] Optionally, the acquisition unit may further include: an extraction subunit, used to extract an audio frame feature vector from the speech audio to be analyzed; an output subunit, used to input the audio frame feature vector into a first preset model to obtain a tensor of the audio frame; an acquisition subunit, used to input the tensor of the audio frame into a second preset model to obtain a character sequence, and determine any character in the character sequence as a target character.

[0103] Optionally, the acquisition module 74 in the embodiment of the present application may further include: a second acquisition unit, used to obtain the similarity between the tensor of the target character and the tensor of the audio frame corresponding to the target character; a second processing unit, used to normalize the similarity to obtain N weight values ​​corresponding to the tensor of the target character.

[0104] Optionally, the second determination module 78 in the embodiment of the present application may further include: a first selection unit, used to select the audio frame with the earliest time sequence from the determined M audio frames as the start frame of the endpoint audio frame corresponding to the target character; a second selection unit, used to select the audio frame with the latest time sequence from the determined M audio frames as the end frame of the endpoint audio frame corresponding to the target character.

[0105] Optionally, if the speech audio to be analyzed includes multiple target characters, the apparatus in the embodiment of the present application may further include:

[0106] A third determination module is used to determine the audio frame that is the frontmost in the time sequence from the start frames of the endpoint audio frames corresponding to the multiple target characters as the start frame of the endpoint audio frame corresponding to the speech audio to be analyzed;

[0107] The fourth determination module is used to determine the audio frame with the latest time sequence from the termination frames of the endpoint audio frames corresponding to the multiple target characters as the termination frame of the endpoint audio frames corresponding to the speech audio to be analyzed.

[0108] In the embodiment of the present application, since the endpoint detection of speech audio is performed based on audio frames, no matter the noise level or the noise ratio, as long as the speech recognition is accurate, a good endpoint detection effect can be obtained; since it can be accurate to the audio frame and adjusted by setting the weight threshold, it can have a higher endpoint detection accuracy.

[0109] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned speech audio analysis method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0110] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0111] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned speech audio analysis method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0112] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0113] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0114] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for analyzing speech audio, characterized in that: include: Determine the tensor of the target character in the speech audio to be analyzed; Obtaining N weight values ​​corresponding to the tensor of the target character, the weight values ​​being used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, the N weight values ​​corresponding to the tensor of the target character being determined based on the similarity between the tensor of the target character and the tensors of the N audio frames corresponding to the target character, wherein N is a positive integer; From the N weight values, select the top M weight values ​​with the largest weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, where M is a positive integer and M is less than N; According to the time sequence of the M audio frames corresponding to the M weight values, the endpoint audio frame corresponding to the target character is determined from the M audio frames corresponding to the M weight values.

2. The method according to claim 1, characterized in that The step of determining the tensor of the target character in the speech audio to be analyzed includes: Obtaining the target character in the speech audio to be analyzed; Perform word embedding processing on the target character to obtain a tensor of the target character.

3. The method according to claim 2, characterized in that The step of obtaining the target character in the speech audio to be analyzed comprises: Extracting an audio frame feature vector from the speech audio to be analyzed; Inputting the audio frame feature vector into a first preset model to obtain a tensor of the audio frame; The tensor of the audio frame is input into a second preset model to obtain a character sequence, and any character in the character sequence is determined as the target character.

4. The method according to claim 3, characterized in that The step of obtaining N weight values ​​corresponding to the tensor of the target character includes: Obtaining the similarity between the tensor of the target character and the tensor of the audio frame corresponding to the target character; The similarity is normalized to obtain N weight values ​​corresponding to the tensor of the target character.

5. The method according to claim 1, characterized in that The step of determining the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values ​​comprises: Selecting the audio frame at the front of the time sequence from the determined M audio frames as the starting frame of the endpoint audio frame corresponding to the target character; The audio frame with the latest time sequence is selected from the determined M audio frames as the end frame of the endpoint audio frame corresponding to the target character.

6. The method according to claim 4, characterized in that If the speech audio to be analyzed includes a plurality of target characters, the method further includes: Determine, from the start frames of the endpoint audio frames corresponding to the multiple target characters, the audio frame that is the frontmost in the time sequence as the start frame of the endpoint audio frame corresponding to the speech audio to be analyzed; The audio frame with the latest time sequence is determined from the termination frames of the endpoint audio frames corresponding to the multiple target characters as the termination frame of the endpoint audio frames corresponding to the speech audio to be analyzed.

7. The method according to claim 3, characterized in that The second preset model is a neural network model with an attention mechanism.

8. A speech audio analysis device, characterized in that: include: A first determination module is used to determine the tensor of the target character in the speech audio to be analyzed; An acquisition module, used for acquiring N weight values ​​corresponding to the tensor of the target character, wherein the weight values ​​are used to represent the weight distribution of the tensor of the target character on the N audio frames corresponding to the target character, and the N weight values ​​corresponding to the tensor of the target character are determined based on the similarity between the tensor of the target character and the tensors of the N audio frames corresponding to the target character, wherein N is a positive integer; A processing module, configured to select, from the N weight values, M weight values ​​with the largest weight values, and the sum of the selected M weight values ​​is greater than or equal to a preset threshold, wherein M is a positive integer and M is less than N; The second determination module is used to determine the endpoint audio frame corresponding to the target character from the M audio frames corresponding to the M weight values ​​according to the time sequence of the M audio frames corresponding to the M weight values.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, it implements the steps of the speech audio analysis method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, it implements the steps of the speech audio analysis method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Real-time speech endpoint detection method and device

    CN109545188A

  • Speech recognition method and device

    CN110895929A