Speech recognition method, device and equipment and readable storage medium
By preprocessing and feature extraction of historical audio data of speech recognition system and optimizing model parameters and algorithms, the problem of insufficient accuracy and robustness of speech recognition system in noisy environments is solved, and the recognition accuracy is significantly improved.
Patent Information
- Application Number
- CN202510244038.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-18
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing speech recognition systems have poor accuracy and robustness when processing noise-carrying audio data, especially in noisy environments, where noise interference can affect feature extraction and model performance.
By obtaining historical audio data, including historical voice data carrying noise and historical voice data without noise, preprocessing and feature extraction are performed, and the similarity of speech features is calculated. When the similarity is greater than the set threshold, the target audio data is input to the preprocessing model for preprocessing, obtain the target voice data, and input it to the speech recognition model for recognition.
It significantly improves the recognition accuracy of the speech recognition model, and provides clearer and more accurate input signals by effectively removing noise interference, ensuring that high-quality speech data enters the recognition process.
Smart Images

Figure CN120071901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a speech recognition method, device, equipment and readable storage medium. Background Art
[0002] In the prior art, noise is a key factor affecting the accuracy and robustness of speech recognition systems. Many speech recognition models perform poorly when processing audio data with noise, resulting in an increase in the recognition error rate. Especially in a noisy environment, the noise interference in the speech signal may mask the speech information, affecting feature extraction and model performance. Therefore, there is an urgent need for an effective method to process speech data with noise to improve the accuracy and robustness of speech recognition systems. Summary of the Invention
[0003] The purpose of the present invention is to provide a speech recognition method, device, equipment and readable storage medium to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0004] In a first aspect, the present application provides a speech recognition method, including:
[0005] Obtaining historical audio data, where the historical audio data includes historical speech data with noise and historical speech data without noise;
[0006] Inputting the historical speech data with noise into a preset processing model for preprocessing to obtain processed speech data;
[0007] Performing feature extraction on the processed speech data and the historical speech data without noise to obtain a first speech feature and a second speech feature respectively;
[0008] Calculating the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, inputting preset target audio data into the preprocessing model for preprocessing to obtain target speech data, where the target audio data is audio data to be recognized;
[0009] Inputting the target speech data into a preset speech recognition model for speech recognition to obtain a speech recognition result.
[0010] In a second aspect, the present application further provides a speech recognition device, including:
[0011] A first acquisition unit, configured to acquire historical audio data, where the historical audio data includes historical speech data with noise and historical speech data without noise;
[0012] A first input unit for inputting historical speech data with noise into a preset processing model for preprocessing to obtain processed speech data;
[0013] An extraction unit for extracting features from the processed speech data and the historical speech data without noise to obtain a first speech feature and a second speech feature respectively;
[0014] A first calculation unit for calculating the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, inputting preset target audio data into the preprocessing model for preprocessing to obtain target speech data, where the target audio data is the audio data to be recognized;
[0015] A second input unit for inputting the target speech data into a preset speech recognition model for speech recognition to obtain a speech recognition result.
[0016] In a third aspect, the present application further provides a speech recognition device, including:
[0017] A memory for storing a computer program;
[0018] A processor for implementing the steps of the speech recognition method when executing the computer program.
[0019] In a fourth aspect, the present application further provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above speech recognition method are implemented.
[0020] The beneficial effects of the present invention are:
[0021] By inputting historical speech data with noise into a model for preprocessing and comparing the features of the processed data with the historical speech data without noise, continuously optimizing the parameters and algorithms of the model, the trained preprocessing model can more effectively identify and remove noise interference, provide a clearer and more accurate input signal for the subsequent speech recognition model, ensure that only high-quality speech data enters the recognition process, and significantly improve the recognition accuracy of the speech recognition model.
[0022] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0024] Figure 1 Schematic flow chart of the speech recognition method described in the embodiments of the present invention;
[0025] Figure 2 Parameter mapping diagram described in the embodiments of the present invention;
[0026] Figure 3 Flow chart of the speech recognition model training described in the embodiments of the present invention;
[0027] Figure 4 Schematic structural diagram of the speech recognition device described in the embodiments of the present invention;
[0028] Figure 5 Schematic structural diagram of the speech recognition device described in the embodiments of the present invention.
[0029] Reference numerals in the figure: 10, first acquisition unit; 20, first input unit; 30, extraction unit; 40, first calculation unit; 50, second input unit; 800, speech recognition device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0031] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0032] Embodiment 1:
[0033] This embodiment provides a speech recognition method.
[0034] See Figure 1 , the figure shows that this method includes step S10, step S20, step S30, step S40 and step S50.
[0035] Step S10. Obtain historical audio data, where the historical audio data includes historical speech data with noise and historical speech data without noise;
[0036] Specifically, the historical speech data with noise is audio data collected by a speech acquisition device in a normal environment, and this historical audio data will include the noise in the environment and irrelevant human voices; the historical speech data without noise is the data obtained by filtering, denoising and manually listening and adjusting the historical speech data with noise, and it is considered that the historical speech data without noise only contains the required human voice data.
[0037] Step S20. Input the historical speech data with noise into a preset processing model for preprocessing to obtain processed speech data;
[0038] Specifically, input the historical speech data with noise into the processing model for preprocessing to obtain processed speech data, and determine whether the processing model can perform better denoising processing on the speech data with noise information by comparing the correlation between the speech features of the processed speech data and the historical speech data without noise.
[0039] Step S30. Extract features from the processed speech data and the historical speech data without noise to obtain a first speech feature and a second speech feature respectively;
[0040] Specifically, step S30 specifically includes step S31, step S32, step S33, step S34, step S35, step S36 and step S37:
[0041] Step S31. Calculate the energy value of the processed speech data in each preset time period to obtain a plurality of energy values;
[0042] Step S32. When the energy value is less than the preset energy threshold, the time period corresponding to the energy value is used as the pause time period;
[0043] Step S33. Based on all the pause time periods, the processed speech data is divided to obtain multiple sub-audio data, and the energy value of each sub-audio data is calculated to obtain multiple target energy values;
[0044] Specifically, due to different speaking habits, for the same text, different people may have different speaking styles and there may be continuous situations. Therefore, the sentences can be divided into multiple sub-audio data through the rise and fall points during speaking, that is, the energy values of the corresponding speech segments, making subsequent processing and analysis more flexible, and more detailed feature extraction can also be performed on different sub-audio data.
[0045] Step S34. Based on the clustering algorithm and multiple target energy values, clustering processing is performed on the multiple sub-audio data to obtain multiple clustering sets, and each clustering set contains multiple sub-audio data with relatively small differences in target energy values;
[0046] Specifically, the clustering algorithm can use distance-based clustering algorithms. In this step, by performing more detailed clustering on the sub-audio data, data sets with similar characteristics can be more accurately identified, providing more accurate input for subsequent feature extraction and correlation analysis.
[0047] Step S35. Based on the Pauta criterion, calculate the energy threshold range of all the sub-audio data contained in each clustering set, and use the clustering set corresponding to the smallest threshold range among all the threshold ranges as the abnormal set;
[0048] Step S36. Delete the abnormal sets in all the clustering sets to obtain multiple target clustering sets;
[0049] Specifically, by using the Pauta criterion and analyzing the threshold range, abnormal data can be automatically detected and excluded. For example, when normal speech is collected in a relatively quiet environment, the sudden background noise will cause the energy value of some sub-audio data to increase abnormally. Such noisy sub-data with large energy values will interfere with subsequent analysis and affect accuracy. Therefore, after detecting these high-energy noisy sub-data, they can be directly deleted to improve the accuracy and reliability of subsequent analysis.
[0050] Step S37. Extract features from the sub-audio data in all the target clustering sets, and aggregate the extracted multiple features to obtain the first speech feature;
[0051] Specifically, considering that a target clustering set contains multiple sub-audio data, the speech features of each target clustering set should be jointly constituted by all the audio data within the set, so as to not only effectively offset the influence of accidental noise in a single sub-audio data, but also ensure that the extracted features are more comprehensive and representative.
[0052] Specifically, step S37 specifically includes step S371, step S372, step S373, step S374, step S375, step S376, and step S377:
[0053] Step S371. Perform Fourier transform on each sub-audio data, and input the spectrum obtained after Fourier transform into a Mel filter for processing to obtain multiple energy logarithms. The energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter;
[0054] Step S372. Perform cosine transform on the energy logarithm to obtain the corresponding multiple Mel-frequency cepstral coefficients;
[0055] Specifically, by performing Fourier transform and Mel filter processing on each sub-audio data, the spectral features of the audio can be extracted. The cosine transform further converts the spectral information into Mel-frequency cepstral coefficients, enhancing the anti-noise performance and expression ability of the features.
[0056] Step S373. Take the first twelve coefficients among the multiple Mel-frequency cepstral coefficients corresponding to each sub-audio data as feature coefficients, and encode all the feature coefficients to obtain a feature coefficient matrix;
[0057] Specifically, by screening the first twelve main feature coefficients for encoding to form a feature coefficient matrix, the system can focus on the most significant frequency features, reduce the influence of lower noise, speaking styles, or emotional changes. The obtained feature coefficient matrix after encoding can better represent the speech features of each target clustering set, helping to improve the accuracy and efficiency of subsequent analysis.
[0058] Step S374. Calculate the correlation degree values of pairwise feature coefficients in the feature coefficient matrix based on the grey relational analysis algorithm;
[0059] Step S375. Transmit the information of each feature coefficient to adjacent feature coefficients, and update the features by combining the information of adjacent feature coefficients to obtain updated feature coefficients;
[0060] Specifically, by calculating the correlation degrees among the feature coefficients in the feature coefficient matrix through grey relational analysis, the internal correlation between features can be revealed, and the information of each feature coefficient is transmitted to adjacent features for feature update. This enables the features to incorporate the context information of adjacent features while maintaining individual differences, thereby improving the overall expression ability and robustness of the features. It not only optimizes the discriminative power of the features but also enhances their noise resistance and robustness, providing a more accurate and stable feature description for subsequent speech recognition.
[0061] The calculation formula for feature update is as follows:
[0062]
[0063] Where x’ s is the updated sth feature coefficient; x s is the sth feature coefficient; h sz is the correlation degree value between the sth feature coefficient and the adjacent zth feature coefficient; x z is the adjacent zth feature coefficient to the sth feature coefficient; n is the number of adjacent feature coefficients to the sth feature coefficient.
[0064] Step S376. Perform a pooling operation on all updated feature coefficients to obtain multiple local feature information, where the local feature information is aggregated from adjacent updated feature coefficients;
[0065] Step S377. Concatenate the multiple local feature information to obtain the first speech feature of the processed speech data;
[0066] Specifically, by performing a pooling operation on all updated feature coefficients, the information of adjacent feature coefficients can be aggregated to obtain multiple local feature information, thereby simplifying the feature representation, reducing the interference of noise, and retaining key speech information at the same time; then, through feature concatenation, the multiple local feature information is integrated into the first speech feature, enabling the system to form a comprehensive and expressive feature vector, which provides a data basis for subsequent speech analysis and recognition.
[0067] Step S40. Calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than the first set threshold, input the preset target audio data into the preprocessing model for preprocessing to obtain the target speech data, where the target audio data is the audio data to be recognized;
[0068] Specifically, by continuously adjusting the model parameters of the processing model, when the similarity between the features of the processed historical audio data and the features of the historical speech data reaches the first set threshold, it is considered that the current processing model has good noise removal performance and can be used for the preprocessing of subsequent speech data to be recognized.
[0069] Specifically, step S40 specifically includes steps S41, S42, S43, S44, S45, S46, S47, S49, S410, and S411:
[0070] Step S41. Perform feature transformation on the first voice feature and the second voice feature to obtain a first spectrogram and a second spectrogram respectively;
[0071] Step S42. Perform block processing on the first spectrogram to obtain a plurality of initial modules;
[0072] Step S43. Calculate the sum of the brightness of each primary color in the three primary colors included in each initial module to obtain a plurality of brightness sums, where the brightness sum is the sum of the brightness of each primary color in the initial module;
[0073] Step S44. Determine any one of the brightness sums as the first brightness sum from the plurality of brightness sums;
[0074] Step S45. Determine a second brightness sum within the neighborhood of the first brightness sum based on a preset neighborhood radius;
[0075] Step S46. When the number of second brightness sums is not less than a preset minimum number, divide the first brightness sum and all second brightness sums into a first target set;
[0076] Step S47. Delete the plurality of brightness sums included in the first target set from the plurality of brightness sums, and determine any one of them, and repeat the above steps until a plurality of first target sets are obtained;
[0077] Step S48. Repeat the above steps to determine a plurality of second target sets corresponding to the second spectrogram;
[0078] Specifically, set the neighborhood radius of the brightness sum, that is, the voice energy of the initial modules corresponding to all brightness sums within this neighborhood is approximately the same, set the minimum number to define the number of brightness sums within the minimum neighborhood of a core point, randomly select a first brightness sum, calculate the difference between this brightness sum and all brightness sums in the target clustering set, mark the points with a difference within the neighborhood radius as second brightness sums, when the number of second brightness sums is greater than the minimum number, then this first brightness sum is the core point, divide this first brightness sum and the adjacent second brightness sums into the same target set, repeat the above steps until all brightness sums are divided into the corresponding target sets, so as to determine all target sets, and determine the feature information of the voice in units of target sets.
[0079] Step S49. Determine a first correlation based on the difference between the number of the first target set and the number of the second target set;
[0080] Specifically, through the above-mentioned division of the neighborhood, the differences between the spectrograms corresponding to the first voice feature and the second voice feature can be effectively compared, that is, the differences in voice features.
[0081] The calculation formula for the first correlation is:
[0082]
[0083] Where S 1 is the first correlation; β is the first control parameter; D 1 is the number of the first target sets; D 2 is the number of the second target sets;
[0084] The control parameter is used to control the decline rate of the correlation.
[0085] Step S410. Determine the second correlation based on the difference between the sum of brightnesses in the first target set and the sum of brightnesses in the second target set;
[0086] Specifically, step S410 specifically includes step S4101, step S4102, step S4103, step S4104, step S4105, step S4106, and step S4107:
[0087] Step S4101. Determine a partial set from multiple first target sets based on a random function as the third target set;
[0088] Specifically, all the first target sets can be encoded, each first target set has a corresponding code, and a partial code is randomly determined from multiple codes by using a random function, and the set corresponding to the code is used as the third target set.
[0089] Step S4102. Determine the fourth target set in the second spectrogram that is closest to the position of the third target set based on the position of the third target set;
[0090] Specifically, determine the position coordinates of the third target set in the first spectrogram and the position coordinates of each second target set in the second spectrogram, calculate the distance difference between the position coordinates of the third target set and the position coordinates of each second target set in the second spectrogram, so as to determine the minimum distance difference, and the second target set corresponding to the minimum distance difference is the fourth target set.
[0091] Step S4103. Calculate the position deviation between each third target set and the corresponding fourth target set to obtain the first deviation, and the position deviation includes the horizontal position deviation and the vertical position deviation;
[0092] Step S4104. Calculate the difference between the sum of brightness in each third target set and the sum of brightness in the corresponding fourth target set as the second deviation.
[0093] Step S4105. Calculate the sum of the first deviation and the second deviation to obtain the target deviation.
[0094] Step S4106. Calculate the product of the target deviation and the control parameter to obtain the target product.
[0095] Step S4107. Calculate the exponential function value with the natural constant as the base and the target product as the exponent to obtain the second correlation.
[0096] Specifically, the calculation formula for the second correlation is:
[0097]
[0098] where S 2 is the second correlation; γ is the second control parameter; D position is the position deviation; D brightness is the sum of brightness deviation.
[0099] The control parameter is used to control the decreasing speed of the correlation.
[0100] Step S411. Determine the similarity based on the first correlation and the second correlation.
[0101] Specifically, the similarity calculation formula is:
[0102] S = S 1 + S 2
[0103] where S is the similarity; S 1 is the first correlation; S 2 is the second correlation.
[0104] Step S50. Input the target voice data into a preset voice recognition model for voice recognition to obtain the voice recognition result.
[0105] Specifically, as Figure 3 shown, the training process of the voice recognition model depends on the sparrow optimization algorithm. First, initialize the sparrow population and conduct fitness evaluation, divide the individuals in the sparrows into different categories, and then optimize different categories of sparrows generation by generation through operations such as simulating natural selection, crossover, and mutation until the optimal solution is found, thereby constructing the corresponding voice recognition model.
[0106] Specifically, step S50 specifically includes steps S51, S52, S53, S54, S55, S56, S57, S58, S59, S510, and S511:
[0107] Step S51. Obtain the historical speech text data corresponding to the historical speech data;
[0108] Specifically, the historical speech text data is the text data corresponding to the historical speech data after automatic recognition and manual adjustment.
[0109] Step S52. Initialize the positions of the sparrow population based on chaotic mapping to obtain multiple initial position parameters;
[0110] Specifically, using the Logistic chaotic mapping to initialize the sparrow population can traverse the states of the sparrow population without repetition within a certain range, enabling the sparrow population to be relatively evenly distributed in the entire search space. This not only increases the diversity of the initial sparrow population but also avoids the situation of falling into local optima during the search process of the sparrow algorithm.
[0111] Specifically, step S52 specifically includes steps S521, S522, S523, S524, and S525:
[0112] Step S521. Obtain the initial value and bifurcation parameter of each sparrow in the sparrow population. The initial value is within the first set range, and the bifurcation parameter is within the second set range;
[0113] Step S522. Calculate the product of the initial value and the bifurcation parameter to obtain the first value;
[0114] Step S523. Calculate the difference between the preset value and the initial value to obtain the second value;
[0115] Step S524. Calculate the product of the first value and the second value to obtain the chaotic value;
[0116] Step S525. Repeat calculating the product of the initial value and the bifurcation parameter, calculating the difference between the preset value and the initial value, and calculating the product of the first value and the second value until the number of chaotic values reaches the set number of times. Take all the chaotic values as the initial position parameters;
[0117] Specifically, the calculation formula for the initial position parameter is:
[0118]
[0119] Among them, is the chaotic value of the i-th sparrow at the (t + 1)-th time; It is the chaotic value of the $t$-th time of Sparrow i, and its value range is $[0, 1]$; $\alpha$ is the bifurcation parameter that determines whether the chaotic mapping is in a chaotic state, and its value range is $[0, 4]$.
[0120] In this embodiment, the chaotic value after mapping each parameter two thousand times is used as the initial position parameter of the sparrow. As Figure 2 shown, it is the parameter mapping diagram after two thousand iterations of chaotic mapping.
[0121] Step S53. Build a model based on the initial position parameter to obtain an initial recognition model;
[0122] Step S54. Determine the initial fitness of all sparrows based on the initial position parameters of all sparrows;
[0123] Specifically, assume that the number of sparrows in the sparrow population is $n$, and each sparrow individual is a solution in a $d$-dimensional solution space. The fitness of the sparrow population is:
[0124]
[0125] where $F(X)$ is the fitness of the sparrow population; $x$ n,d is the position of sparrow $n$ in the population in the $d$-th dimension space; $f([x$ n, 1 $x$ n,2 …$x$ n,d ) is the individual fitness of sparrow $n$.
[0126] Step S55. Divide all sparrows into discoverers, followers, and vigilants based on the initial position parameters;
[0127] Step S56. Update the position parameters of the discoverers, followers, and vigilants respectively based on the preset position update formula to obtain the updated position parameters;
[0128] Specifically, dividing the sparrows into discoverers, followers, and vigilants helps to simulate the process of natural selection, enhance the diversity and adaptability of the group, and through the position update formula to dynamically adjust sparrows with different roles, the search strategy can be optimized, the global search ability can be improved, and thus converge to the optimal solution more effectively.
[0129] Considering that the global search relies too much on the position of the discoverer, referring to the spiral upward search method of iterative optimization in the whale optimization algorithm, an adaptive weight factor is introduced, so that when the discoverer in the sparrow algorithm explores the next area after position update, it searches with a larger step size in the early stage and converges and explores with a smaller step size in the later stage.
[0130] In the embodiment of this application, the position update formula of the discoverer is:
[0131]
[0132] Among them, is the j-th dimension position of sparrow i in the population after the (T + 1)-th iteration; is the j-th dimension position of sparrow i in the population after the T-th iteration; δ is the adaptive weight factor; iter max is the maximum number of iterations of the population; θ is a uniformly distributed random number with a value range of (0, 1]; R 2 is the warning value with a value range of [0, 1]; ST is the warning threshold with a value range of [0.5, 1]; Q is a random number subject to a normal distribution; Z is a 1×d matrix; is the worst position of sparrow i in the population after the T-th iteration; is the best position of sparrow i in the population after the T-th iteration; W(T) is the adaptive weight factor, a function that decreases with the increase of the iteration number T; A is the amplitude factor of the spiral search; B is the frequency factor of the spiral search; is the upper bound constraint of sparrow i; is the lower bound constraint of sparrow i; iter i is the current iteration number of sparrow i.
[0133] When R 2 < ST, it indicates that when the warning value is lower than the safety value, there is no predator in the foraging environment, and at this time, the discoverer can conduct extensive searches; when R 2 ≥ ST, it indicates that some sparrows in the population have discovered the predator and issued warnings to other sparrows, and all sparrows need to quickly fly to a safe area to forage.
[0134] In the embodiments of the present application, the position update of the sparrow follower is as follows;
[0135]
[0136] Among them, is the j-th dimension position of sparrow i in the population after the (T + 1)-th iteration; X w is the position of the worst sparrow in the population after the T-th iteration; is the j-th dimension position of sparrow i in the population after the T-th iteration; Q is a random number subject to a normal distribution; is the position of sparrow p in the population after the (T + 1)-th iteration; K is a uniformly distributed random number with a value range of [-1, 1]; d is the spatial dimension; is the j-th dimension position of sparrow p in the population after the T-th iteration; n is the number of individuals in the sparrow population;
[0137] When i > n / 2, it indicates that the i-th follower has not obtained food and is in a state of hunger. At this time, it needs to fly to other places to forage for food to obtain more energy; when i ≤ n / 2, it indicates that the i-th follower can obtain food and does not need to fly to other places to forage for food to obtain energy.
[0138] To improve the convergence speed and accuracy of the sparrow algorithm, the step size factors and B are corrected. The correction formula for the step size factor is:
[0139]
[0140] where, f g is the fitness of the optimal sparrow in the current population; f w is the fitness of the worst sparrow in the current population; t is the current iteration number; Si is the initialization parameter; rand is a random factor with a uniform distribution; and B are step size factors, and the value range of B is [0, 1].
[0141] The position update formula for the sparrow vigilante is:
[0142]
[0143] where, is the j-th dimension position of sparrow i in the population after the (T + 1)-th iteration; is the j-th dimension position of sparrow i in the population after the T-th iteration; is the position of the optimal sparrow in the population after the T-th iteration; is the position of the worst sparrow in the population after the T-th iteration; f i is the individual fitness of sparrow i; f w is the fitness of the worst sparrow in the current population; f g is the fitness of the optimal sparrow in the current population; ε is the smallest constant to prevent the denominator from being zero; and B are step size factors, and the value range of B is [0, 1];
[0144] When f i ≠ f g it means that the sparrow is at the edge of the population at this time and is extremely vulnerable to attacks by predators; when f i = f g it means that the sparrows in the middle of the population are also in danger. At this time, they need to get closer to other sparrows to reduce the risk of being preyed on.
[0145] Step S57. Determine the current iteration number based on the number of position updates;
[0146] Step S58. Calculate the mutation rate based on a preset mutation rate, a preset minimum mutation rate, a preset maximum number of iterations, and the current number of iterations;
[0147] Step S59. Determine multiple target sparrows from the sparrow population based on the mutation rate, and optimize the updated position parameters of the multiple target sparrows using the Cauchy distribution algorithm to obtain optimized position parameters;
[0148] Specifically, the mutation operation can expand the search space of the sparrow population, but it does not require every sparrow individual to perform this operation in each iteration. Instead, the mutation rate is used to determine the target sparrows for mutation, thereby optimizing the position parameters. This approach can appropriately maintain the population diversity while avoiding excessive random perturbations, thus improving the convergence speed and search efficiency of the algorithm.
[0149] Step S510. When the number of sentinels is less than the second set threshold, determine the target sparrow from the current sparrow population, as well as the optimized position parameters and initial fitness corresponding to the target sparrow. The target sparrow is the sparrow with the maximum fitness in the sparrow population;
[0150] Step S511. Adjust and optimize the preset initial recognition model based on the historical speech data, historical speech text data, the optimized position parameters of the target sparrow, and the initial fitness of the target sparrow to obtain a speech recognition model;
[0151] Specifically, considering that when the number of sentinels is large, it is beneficial for the global search of the algorithm; when the number of sentinels is small, it is more conducive to the algorithm to perform fast local search within a small range. Therefore, in the embodiments of the present application, a higher sentinel ratio is set in the early stage of the algorithm to enhance the global exploration ability of the population, and the sentinel ratio is gradually reduced as the number of iterations increases, thereby improving the convergence speed of the algorithm.
[0152] The update formula for the sentinel ratio is:
[0153]
[0154] where P is the sentinel ratio; P 0 is the initial sentinel ratio; T is the current number of iterations; iter max is the maximum number of iterations of the population; P min is the preset minimum sentinel ratio;
[0155] When P > P min , continue the sparrow position update iteration operation until the condition is met; when P ≤ P minWhen the condition is met, the iteration ends. The optimal position and the best fitness value are obtained globally, and the optimal weights and thresholds of the convolutional neural network are determined. The optimal weights and thresholds are sent back to the convolutional neural network for retraining to obtain a speech recognition model.
[0156] Inputting the target speech data into the speech recognition model can achieve fast recognition and conversion of speech into text. When applied in the judicial scenario, it can help judicial mediators record and process case information more efficiently, thereby improving the quality and efficiency of mediation work.
[0157] Embodiment 2:
[0158] As Figure 4 shown, this embodiment provides a speech recognition device, which includes:
[0159] The first acquisition unit 10 is used to acquire historical audio data, where the historical audio data includes historical speech data with noise and historical speech data without noise;
[0160] The first input unit 20 is used to input the historical speech data with noise into a preset processing model for preprocessing to obtain processed speech data;
[0161] The extraction unit 30 is used to extract features from the processed speech data and the historical speech data without noise, respectively obtaining a first speech feature and a second speech feature;
[0162] The first calculation unit 40 is used to calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, the preset target audio data is input into the preprocessing model for preprocessing to obtain target speech data, where the target audio data is the audio data to be recognized;
[0163] The second input unit 50 is used to input the target speech data into a preset speech recognition model for speech recognition to obtain a speech recognition result.
[0164] In a specific implementation manner disclosed in this application, the extraction unit 30 includes:
[0165] The second calculation unit is used to calculate the energy value of the processed speech data in each preset time period to obtain a plurality of energy values;
[0166] The first determination unit is used to, when the energy value is less than a preset energy threshold, use the time period corresponding to the energy value as a pause time period;
[0167] The first division unit is used to divide the processed speech data based on all the pause time periods to obtain a plurality of sub-audio data, and calculate the energy value of each sub-audio data to obtain a plurality of target energy values;
[0168] A clustering unit, configured to perform clustering processing on multiple sub-audio data based on a clustering algorithm and multiple target energy values, to obtain multiple clustering sets, where each clustering set contains multiple sub-audio data with relatively small differences in target energy values;
[0169] A second determination unit, configured to calculate an energy threshold range of all sub-audio data included in each clustering set based on the Chauvenet's criterion, and use the clustering set corresponding to the smallest threshold range among all threshold ranges as an abnormal set;
[0170] A first deletion unit, configured to delete the abnormal set from all clustering sets to obtain multiple target clustering sets;
[0171] An aggregation unit, configured to extract features from the sub-audio data in all target clustering sets, and perform aggregation processing on the extracted multiple features to obtain a first speech feature.
[0172] In a specific implementation manner disclosed in this application, the aggregation unit includes:
[0173] A first transformation unit, configured to perform Fourier transform on each sub-audio data, and input the spectrum obtained after the Fourier transform into a Mel filter for processing to obtain multiple energy logarithms, where the energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter;
[0174] A second transformation unit, configured to perform cosine transform on the energy logarithm to obtain corresponding multiple Mel-frequency cepstral coefficients;
[0175] A third determination unit, configured to use the top twelve coefficients among the multiple Mel-frequency cepstral coefficients corresponding to each sub-audio data as feature coefficients, and encode all feature coefficients to obtain a feature coefficient matrix;
[0176] A third calculation unit, configured to calculate the correlation degree values between pairwise feature coefficients in the feature coefficient matrix based on a grey relational analysis algorithm;
[0177] A transmission unit, configured to transmit the information of each feature coefficient to adjacent feature coefficients, and update the features by combining the information of adjacent feature coefficients to obtain updated feature coefficients;
[0178] A pooling unit, configured to perform a pooling operation on all updated feature coefficients to obtain multiple local feature information, where the local feature information is aggregated from adjacent updated feature coefficients;
[0179] A splicing unit, configured to splice the multiple local feature information to obtain the first speech feature of the processed speech data.
[0180] In a specific implementation manner disclosed in the present application, the second input unit 50 includes:
[0181] A second acquisition unit, configured to acquire historical speech text data corresponding to historical speech data;
[0182] An initialization unit, configured to initialize the positions of the sparrow population based on a chaotic map to obtain a plurality of initial position parameters;
[0183] A construction unit, configured to construct a model based on the initial position parameters to obtain an initial recognition model;
[0184] A first determination unit, configured to determine the initial fitness of all sparrows based on the initial position parameters of all sparrows;
[0185] A second division unit, configured to divide all sparrows into discoverers, followers, and guards based on the initial position parameters;
[0186] An update unit, configured to update the position parameters of the discoverers, followers, and guards respectively based on a preset position update formula to obtain updated position parameters;
[0187] A second determination unit, configured to determine the current iteration number based on the position update times;
[0188] A fourth calculation unit, configured to calculate a mutation rate based on a preset mutation rate, a preset minimum mutation rate, a preset maximum iteration number, and the current iteration number;
[0189] A third determination unit, configured to determine a plurality of target sparrows from the sparrow population based on the mutation rate, and optimize the updated position parameters of the plurality of target sparrows based on the Cauchy distribution algorithm to obtain optimized position parameters;
[0190] A fourth determination unit, configured to, when the number of guards is less than a second set threshold, determine target sparrows from the current sparrow population, as well as the optimized position parameters and initial fitness corresponding to the target sparrows, where the target sparrows are the sparrows with the maximum fitness in the sparrow population;
[0191] An optimization unit, configured to adjust and optimize a preset initial recognition model based on the historical speech data, the historical speech text data, the optimized position parameters of the target sparrows, and the initial fitness of the target sparrows to obtain a speech recognition model.
[0192] In a specific implementation manner disclosed in the present application, the initialization unit includes:
[0193] A third acquisition unit, configured to acquire the initial values and branch parameters of each sparrow in the sparrow population, where the initial values are within a first set range and the branch parameters are within a second set range;
[0194] A product unit for calculating the product of an initial value and a branch parameter to obtain a first numerical value;
[0195] A fifth calculation unit for calculating the difference between a preset numerical value and the initial value to obtain a second numerical value;
[0196] A sixth calculation unit for calculating the product of the first numerical value and the second numerical value to obtain a chaotic value;
[0197] A first repetition unit for repeatedly calculating the product of the initial value and the branch parameter, calculating the difference between the preset numerical value and the initial value, and calculating the product of the first numerical value and the second numerical value until the number of chaotic values reaches a set number of times, and taking all the chaotic values as the initial position parameters.
[0198] In a specific embodiment disclosed in the present application, the first calculation unit 40 includes:
[0199] A conversion unit for performing feature conversion on the first voice feature and the second voice feature to respectively obtain a first spectrogram and a second spectrogram;
[0200] A processing unit for performing block processing on the first spectrogram to obtain a plurality of initial modules;
[0201] A seventh calculation unit for calculating the sum of the brightnesses of the primary colors in each initial module to obtain a plurality of brightness sums, where the brightness sum is the sum of the brightnesses of the primary colors in the initial module;
[0202] A fifth determination unit for determining any one of the brightness sums as the first brightness sum from the plurality of brightness sums;
[0203] A sixth determination unit for determining a second brightness sum located within the neighborhood of the first brightness sum based on a preset neighborhood radius;
[0204] A third partitioning unit for partitioning the first brightness sum and all the second brightness sums into a first target set when the number of the second brightness sums is not less than a preset minimum number;
[0205] A second deletion unit for deleting the plurality of brightness sums included in the first target set from the plurality of brightness sums and determining any one of them, and repeating the above steps until a plurality of first target sets are obtained;
[0206] A second repetition unit for repeating the above steps to determine a plurality of second target sets corresponding to the second spectrogram;
[0207] A seventh determination unit for determining a first correlation based on the difference between the number of the first target sets and the number of the second target sets;
[0208] An eighth determination unit, configured to determine a second correlation based on the difference between the sum of brightnesses in the first target set and the sum of brightnesses in the second target set;
[0209] A ninth determination unit, configured to determine a similarity based on the first correlation and the second correlation.
[0210] In a specific implementation manner disclosed in this application, the eighth determination unit includes:
[0211] A tenth determination unit, configured to determine a partial set from multiple first target sets as a third target set based on a random function;
[0212] An eleventh determination unit, configured to determine a fourth target set in the second spectrogram that is closest to the location of the third target set based on the location of the third target set;
[0213] An eighth calculation unit, configured to calculate the position deviation between each third target set and the corresponding fourth target set to obtain a first deviation, where the position deviation includes a horizontal position deviation and a vertical position deviation;
[0214] A ninth calculation unit, configured to calculate the difference between the sum of brightnesses in each third target set and the sum of brightnesses in the corresponding fourth target set as a second deviation;
[0215] A tenth calculation unit, configured to calculate the sum of the first deviation and the second deviation to obtain a target deviation;
[0216] An eleventh calculation unit, configured to calculate the product of the target deviation and a control parameter to obtain a target product;
[0217] A twelfth calculation unit, configured to calculate the exponential function value with the natural constant as the base and the target product as the exponent to obtain the second correlation.
[0218] It should be noted that for the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0219] Embodiment 3:
[0220] Corresponding to the above method embodiment, a voice recognition device is further provided in this embodiment. A voice recognition device described below can be mutually referred to with a voice recognition method described above.
[0221] Figure 5 It is a block diagram of a voice recognition device 800 shown according to an exemplary embodiment. As Figure 5As shown, the voice recognition device 800 may include: a processor 801 and a memory 802. The voice recognition device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0222] Among them, the processor 801 is used to control the overall operation of the voice recognition device 800 to complete all or part of the steps in the above voice recognition method. The memory 802 is used to store various types of data to support the operation of the voice recognition device 800. These data may include, for example, instructions for any application or method operating on the voice recognition device 800, as well as application-related data, such as contact data, sent and received messages, pictures, audio, video, and so on. The memory 802 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 802 or sent through the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the voice recognition device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 805 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.
[0223] In an exemplary embodiment, the speech recognition device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned speech recognition method.
[0224] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned speech recognition method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 802 including program instructions, and the above-mentioned program instructions can be executed by the processor 801 of the speech recognition device 800 to complete the above-mentioned speech recognition method.
[0225] Embodiment 4:
[0226] Corresponding to the above method embodiment, a readable storage medium is further provided in this embodiment. A readable storage medium described below can be mutually corresponded and referred to with a speech recognition method described above.
[0227] A readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the steps of the speech recognition method in the above method embodiment are implemented.
[0228] Specifically, the readable storage medium can be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0229] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0230] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A speech recognition method, characterized in that: include: Acquire historical audio data, wherein the historical audio data includes historical speech data with noise and historical speech data without noise; Inputting the historical speech data with noise into a preset processing model for preprocessing to obtain processed speech data; Performing feature extraction on the processed speech data and the historical speech data without noise to obtain a first speech feature and a second speech feature respectively; Calculating the similarity between the first speech feature and the second speech feature, and when the similarity is greater than a first set threshold, inputting the preset target audio data into the preprocessing model for preprocessing to obtain target speech data, where the target audio data is the audio data to be recognized; The target speech data is input into a preset speech recognition model for speech recognition to obtain a speech recognition result.
2. The speech recognition method according to claim 1, characterized in that , feature extraction is performed on the processed speech data and the historical speech data without noise, and the first speech feature and the second speech feature are obtained respectively, including: Calculating the energy value of the processed voice data in each preset time period to obtain multiple energy values; When the energy value is less than a preset energy threshold, the time period corresponding to the energy value is used as a pause time period; Dividing the processed speech data based on all pause time periods to obtain a plurality of sub-audio data, and calculating the energy value of each sub-audio data to obtain a plurality of target energy values; Clustering the multiple sub-audio data based on a clustering algorithm and multiple target energy values to obtain multiple cluster sets, each cluster set containing multiple sub-audio data with small differences in target energy values; Calculate the energy threshold range of all sub-audio data contained in each cluster set based on the Laida criterion, and take the cluster set corresponding to the minimum threshold range among all threshold ranges as the abnormal set; Deleting the abnormal set in all cluster sets to obtain multiple target cluster sets; Feature extraction is performed on the sub-audio data in all target cluster sets, and the extracted multiple features are aggregated to obtain the first speech feature.
3. The speech recognition method according to claim 2, characterized in that , extracting features from the sub-audio data in all target cluster sets, and aggregating the extracted multiple features to obtain the first speech feature, including: Performing Fourier transform on each sub-audio data, and inputting the spectrum obtained after Fourier transform into a Mel filter for processing to obtain a plurality of energy logarithms, where the energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter; Performing a cosine transform on the energy logarithm to obtain a corresponding plurality of Mel-frequency cepstrum coefficients; Taking the first twelve coefficients of the plurality of Mel-frequency cepstrum coefficients corresponding to each sub-audio data as characteristic coefficients, and encoding all the characteristic coefficients to obtain a characteristic coefficient matrix; Calculate the correlation values of each pair of characteristic coefficients in the characteristic coefficient matrix based on the grey correlation analysis algorithm; The information of each feature coefficient is passed to the adjacent feature coefficients, and the feature is updated by combining the information of the adjacent feature coefficients to obtain the updated feature coefficients; Performing a pooling operation on all updated feature coefficients to obtain a plurality of local feature information, wherein the local feature information is obtained by aggregating adjacent updated feature coefficients; The plurality of local feature information are concatenated to obtain a first speech feature of the processed speech data.
4. The speech recognition method according to claim 1, characterized in that , constructing the speech recognition model, including: Acquire historical voice text data corresponding to the historical voice data; Initialize the position of the sparrow population based on chaotic mapping to obtain multiple initial position parameters; Building a model based on the initial position parameters to obtain an initial recognition model; Based on the initial position parameters of all sparrows, the initial fitness of all sparrows is determined; Based on the initial position parameters, all sparrows are divided into finders, followers and guards; Based on a preset position update formula, the position parameters of the discoverer, the follower and the sentinel are updated respectively to obtain updated position parameters; Determining a current number of iterations based on the number of position updates; The mutation rate is calculated based on a preset mutation rate, a preset minimum mutation rate, a preset maximum number of iterations and the current number of iterations; Determine a plurality of target sparrows from the sparrow population based on the mutation rate, and optimize the updated position parameters of the plurality of target sparrows based on the Cauchy distribution algorithm to obtain optimized position parameters; When the number of the alerters is less than a second set threshold, a target sparrow is determined from the current sparrow population, as well as the optimized position parameters and initial fitness corresponding to the target sparrow, wherein the target sparrow is the sparrow with the largest fitness in the sparrow population; Based on the historical voice data, the historical voice text data, the optimized position parameters of the target sparrow and the initial fitness of the target sparrow, the preset initial recognition model is adjusted and optimized to obtain the voice recognition model.
5. The speech recognition method according to claim 4, characterized in that ,Based on the chaotic mapping, the position of the sparrow population is initialized, and multiple initial position parameters are obtained, including: Acquire an initial value and a branch parameter of each sparrow in a sparrow population, wherein the initial value is within a first set range, and the branch parameter is within a second set range; Calculating the product of the initial value and the branch parameter to obtain a first value; Calculate the difference between the preset value and the initial value to obtain a second value; Calculate the product of the first value and the second value to obtain a chaos value; Repeat the calculation of the product of the initial value and the branch parameter, the calculation of the difference between the preset value and the initial value, and the calculation of the product of the first value and the second value until the number of chaotic values reaches the set number of times, and use all chaotic values as initial position parameters.
6. The speech recognition method according to claim 5, characterized in that ,The calculation formula of initial position parameters is: in, is the chaos value of sparrow i at the t+1th time; is the t-th chaotic value of sparrow i, and its value range is [0,1]; α is the branch parameter that determines whether the chaotic map is in a chaotic state, and its value range is [0,4].
7. The speech recognition method according to claim 4, characterized in that ,The discoverer’s position update formula is: in, is the j-th position of sparrow i in the population after the T+1th iteration; is the j-th position of sparrow i in the population after the T-th iteration; δ is the adaptive weight factor; iter max is the maximum number of iterations of the population; θ is a uniform random number, ranging from (0,1]; R2 is the warning value, ranging from [0,1]; ST is the warning threshold, ranging from [0.5,1]; Q is a random number that obeys the normal distribution; Z is a matrix with 1 row and d columns; is the worst position of sparrow i in the population after the Tth iteration; is the optimal position of sparrow i in the population after the Tth iteration; W(T) is the adaptive weight factor, a function that decreases as the number of iterations T increases; A is the amplitude factor of the spiral search; B is the frequency factor of the spiral search; is the upper limit constraint of sparrow i; is the lower limit constraint of sparrow i; iter i is the current iteration number of sparrow i.
8. A speech recognition device, characterized in that: include: A first acquisition unit, configured to acquire historical audio data, wherein the historical audio data includes historical speech data with noise and historical speech data without noise; A first input unit is used to input historical speech data carrying noise into a preset processing model for preprocessing to obtain processed speech data; An extraction unit, used to extract features from the processed speech data and the historical speech data without noise, to obtain a first speech feature and a second speech feature respectively; A first calculation unit is used to calculate the similarity between the first speech feature and the second speech feature, and when the similarity is greater than a first set threshold, input the preset target audio data into the preprocessing model for preprocessing to obtain target speech data, where the target audio data is the audio data to be recognized; The second input unit is used to input the target voice data into a preset voice recognition model for voice recognition to obtain a voice recognition result.
9. The speech recognition device according to claim 8, characterized in that: The extraction unit comprises: A second calculation unit, used to calculate the energy value of the processed voice data in each preset time period to obtain multiple energy values; The first unit is used to use the time period corresponding to the energy value as a pause time period when the energy value is less than a preset energy threshold; A first dividing unit is used to divide the processed speech data based on all pause time periods to obtain a plurality of sub-audio data, and calculate the energy value of each sub-audio data to obtain a plurality of target energy values; A clustering unit, used for clustering the multiple sub-audio data based on a clustering algorithm and multiple target energy values to obtain multiple cluster sets, each cluster set containing multiple sub-audio data with small differences in target energy values; The second unit is used to calculate the energy threshold range of all sub-audio data contained in each cluster set based on the Laida criterion, and take the cluster set corresponding to the minimum threshold range among all the threshold ranges as the abnormal set; A first deleting unit is used to delete the abnormal set in all cluster sets to obtain multiple target cluster sets; The aggregation unit is used to extract features from the sub-audio data in all target cluster sets, and aggregate the extracted multiple features to obtain the first speech feature.
10. The speech recognition device according to claim 9, characterized in that: The polymeric unit comprises: A first transform unit is used to perform Fourier transform on each sub-audio data, and input the spectrum obtained after Fourier transform into a Mel filter for processing to obtain a plurality of energy logarithms, where the energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter; A second transformation unit is used to perform a cosine transformation on the energy logarithm to obtain a corresponding plurality of Mel-frequency cepstral coefficients; The third unit is used to take the first twelve coefficients of the multiple Mel-frequency cepstral coefficients corresponding to each sub-audio data as characteristic coefficients, and encode all the characteristic coefficients to obtain a characteristic coefficient matrix; A third calculation unit is used to calculate the correlation value of each pair of characteristic coefficients in the characteristic coefficient matrix based on a grey correlation analysis algorithm; A transmission unit, used to transmit information of each feature coefficient to an adjacent feature coefficient, and perform feature update in combination with the information of the adjacent feature coefficient to obtain an updated feature coefficient; A pooling unit, used for performing a pooling operation on all updated feature coefficients to obtain a plurality of local feature information, wherein the local feature information is obtained by aggregating adjacent updated feature coefficients; The concatenation unit is used to concatenate multiple local feature information to obtain the first speech feature of the processed speech data.
Citation Information
Patent Citations
Position prompting method and device, storage medium and electronic equipment
CN108922523A
Audio recognition method and device, electronic equipment and storage medium
CN113889146A
Robustness speech enhancement method based on self-learning complex convolutional neural network
CN114566178A
Speech recognition model training method, speech recognition method and electronic equipment
CN114582330A
Voice wake-up word detection method and device, storage medium and electronic equipment
CN116705013A