A speech recognition method, apparatus, device, and readable storage medium

By preprocessing and extracting features from noisy speech data, and training the model using the Sparrow Optimization Algorithm, the noise interference problem of speech recognition systems in noisy environments is solved, improving recognition accuracy and robustness.

CN120071901BActive Publication Date: 2025-12-02SICHUAN XINYUNDIAO TECHNOLOGY SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510244038.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-11-18
Filing Date
2025-03-03
Publication Date
2025-12-02
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Existing speech recognition systems suffer from severe noise interference in noisy environments, which affects feature extraction and model performance, leading to an increased recognition error rate.

Method used

By preprocessing historical speech data carrying noise, clustering algorithms and feature extraction techniques are used to remove noise. The speech recognition model is then trained using the Sparrow Optimization Algorithm, and the model parameters and algorithm are optimized to improve recognition accuracy.

Benefits of technology

It significantly improves the recognition accuracy and robustness of the speech recognition model, ensures that high-quality speech data enters the recognition process, and improves the accuracy and stability of the recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071901B_ABST
    Figure CN120071901B_ABST
Patent Text Reader

Abstract

This invention provides a speech recognition method, apparatus, device, and readable storage medium, relating to the field of speech recognition technology. The method includes acquiring historical audio data; obtaining processed speech data; performing feature extraction to obtain a first speech feature and a second speech feature; when the similarity is greater than a first preset threshold, inputting preset target audio data into a preprocessing model for preprocessing to obtain target speech data; and obtaining a speech recognition result. This invention preprocesses noisy historical speech data into a model and compares the processed data with noisy historical speech data, continuously optimizing the model's parameters and algorithms. This allows the trained preprocessing model to more effectively identify and remove noise interference, providing clearer and more accurate input signals for subsequent speech recognition models. This ensures that only high-quality speech data enters the recognition process, significantly improving the recognition accuracy of the speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and more specifically, to a speech recognition method, apparatus, device, and readable storage medium. Background Technology

[0002] Noise is a key factor affecting the accuracy and robustness of speech recognition systems in existing technologies. Many speech recognition models perform poorly when processing audio data with noise, leading to an increased recognition error rate. Especially in noisy environments, noise interference in speech signals may mask speech information and affect feature extraction and model performance. Therefore, there is an urgent need for an effective method to process noisy speech data in order to improve the accuracy and robustness of speech recognition systems. Summary of the Invention

[0003] The purpose of this invention is to provide a speech recognition method, apparatus, device, and readable storage medium to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:

[0004] In a first aspect, this application provides a speech recognition method, including:

[0005] Acquire historical audio data, which includes historical speech data with noise and historical speech data without noise;

[0006] Noisy historical speech data is input into a preset processing model for preprocessing to obtain processed speech data.

[0007] Feature extraction is performed on the processed speech data and the historical speech data without noise to obtain the first speech feature and the second speech feature, respectively.

[0008] Calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, input the preset target audio data into the preprocessing model for preprocessing to obtain target speech data. The target audio data is the audio data to be identified.

[0009] The target speech data is input into a preset speech recognition model for speech recognition to obtain the speech recognition result.

[0010] Secondly, this application also provides a voice recognition device, comprising:

[0011] The first acquisition unit is used to acquire historical audio data, which includes historical speech data with noise and historical speech data without noise.

[0012] The first input unit is used to input noisy historical speech data into a preset processing model for preprocessing to obtain processed speech data.

[0013] The extraction unit is used to extract features from the processed speech data and the historical speech data without noise, and obtain the first speech feature and the second speech feature, respectively.

[0014] The first calculation unit is used to calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, the preset target audio data is input into the preprocessing model for preprocessing to obtain target speech data. The target audio data is the audio data to be identified.

[0015] The second input unit is used to input the target speech data into a preset speech recognition model for speech recognition and obtain the speech recognition result.

[0016] Thirdly, this application also provides a voice recognition device, comprising:

[0017] Memory, used to store computer programs;

[0018] A processor is used to implement the speech recognition method when executing the computer program.

[0019] Fourthly, this application also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described speech recognition method.

[0020] The beneficial effects of this invention are as follows:

[0021] This invention preprocesses historical speech data carrying noise by inputting it into a model, and then compares the processed data with historical speech data without noise by feature comparison. This continuously optimizes the model's parameters and algorithms, enabling the trained preprocessed model to more effectively identify and remove noise interference. This provides a clearer and more accurate input signal for subsequent speech recognition models, ensuring that only high-quality speech data enters the recognition process and significantly improving the recognition accuracy of speech recognition models.

[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the speech recognition method described in an embodiment of the present invention;

[0025] Figure 2 This is the parameter mapping diagram described in the embodiments of the present invention;

[0026] Figure 3 This is a flowchart illustrating the speech recognition model training process described in this embodiment of the invention.

[0027] Figure 4 This is a schematic diagram of the speech recognition device described in an embodiment of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of the speech recognition device described in an embodiment of the present invention.

[0029] The following are the markings in the diagram: 10, First Acquisition Unit; 20, First Input Unit; 30, Extraction Unit; 40, First Calculation Unit; 50, Second Input Unit; 800, Speech Recognition Device; 801, Processor; 802, Memory; 803, Multimedia Component; 804, I / O Interface; 805, Communication Component. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0031] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0032] Example 1:

[0033] This embodiment provides a speech recognition method.

[0034] See Figure 1 The figure shows that the method includes steps S10, S20, S30, S40 and S50.

[0035] Step S10. Obtain historical audio data, which includes historical speech data with noise and historical speech data without noise;

[0036] Specifically, noisy historical voice data refers to audio data collected in a normal environment by a voice acquisition device. This historical audio data will include environmental noise and irrelevant human voices. Noise-free historical voice data is data obtained by filtering, denoising, and manually adjusting the noisy historical voice data. It is assumed that noise-free historical voice data only contains the required human voice data.

[0037] Step S20. Input the noisy historical speech data into the preset processing model for preprocessing to obtain the processed speech data;

[0038] Specifically, the noisy historical speech data is input into the processing model for preprocessing to obtain processed speech data. By comparing the correlation between the speech features of the processed speech data and the noisy historical speech data, it is determined whether the processing model can perform good denoising processing on the noisy speech data.

[0039] Step S30. Extract features from the processed speech data and the historical speech data without noise to obtain the first speech feature and the second speech feature, respectively;

[0040] Specifically, step S30 includes steps S31, S32, S33, S34, S35, S36, and S37:

[0041] Step S31. Calculate the energy value of the processed speech data in each preset time period to obtain multiple energy values;

[0042] Step S32. When the energy value is less than the preset energy threshold, the time period corresponding to the energy value is taken as the pause time period;

[0043] Step S33. Divide the processed speech data based on all pause time periods to obtain multiple sub-audio data, and calculate the energy value of each sub-audio data to obtain multiple target energy values;

[0044] Specifically, due to different speaking habits, different people may speak the same text in different ways, sometimes with a continuous flow. Therefore, sentences can be divided into multiple sub-audio data by using the energy values ​​of the corresponding speech segments at the start and end points of speech. This makes subsequent processing and analysis more flexible and allows for more detailed feature extraction from different sub-audio data.

[0045] Step S34. Based on the clustering algorithm and multiple target energy values, perform clustering processing on multiple sub-audio data to obtain multiple cluster sets. Each cluster set contains multiple sub-audio data with small differences in target energy values.

[0046] Specifically, distance-based clustering algorithms can be used. This step, by performing more detailed clustering on the sub-audio data, can more accurately identify data sets with similar features, providing more accurate input for subsequent feature extraction and association analysis.

[0047] Step S35. Calculate the energy threshold range of all sub-audio data contained in each cluster set based on the Laida criterion, and take the cluster set corresponding to the smallest threshold range among all threshold ranges as the outlier set;

[0048] Step S36. Delete the outlier sets from all cluster sets to obtain multiple target cluster sets;

[0049] Specifically, by using the Raida criterion and analysis threshold range, abnormal data can be automatically detected and eliminated. For example, when normal speech is collected in a relatively quiet environment, sudden background noise can cause the energy value of some sub-audio data to rise abnormally. Such noise sub-data with high energy values ​​will interfere with subsequent analysis and affect accuracy. Therefore, after detecting these high-energy noise sub-data, they can be directly deleted to improve the accuracy and reliability of subsequent analysis.

[0050] Step S37. Extract features from the sub-audio data in all target cluster sets, and aggregate the extracted features to obtain the first speech feature;

[0051] Specifically, considering that a target cluster set contains multiple sub-audio data, the speech features of each target cluster set should be composed of all audio data within the set. This not only effectively counteracts the influence of occasional noise in a single sub-audio data, but also ensures that the extracted features are more comprehensive and representative.

[0052] Specifically, step S37 includes steps S371, S372, S373, S374, S375, S376, and S377:

[0053] Step S371. Perform a Fourier transform on each sub-audio data, and input the spectrum obtained after the Fourier transform into a Mel filter for processing to obtain multiple energy logarithms. The energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter.

[0054] Step S372. Perform a cosine transform on the logarithm of the energy to obtain the corresponding multiple Mel frequency cepstral coefficients;

[0055] Specifically, by performing Fourier transform and Mel filtering on each sub-audio data, the spectral features of the audio can be extracted. The cosine transform further converts the spectral information into Mel frequency cepstral coefficients, enhancing the noise resistance and expressive power of the features.

[0056] Step S373. Take the first twelve coefficients of the multiple Mel frequency cepstral coefficients corresponding to each sub-audio data as feature coefficients, and encode all feature coefficients to obtain the feature coefficient matrix;

[0057] Specifically, by selecting the top twelve key feature coefficients and encoding them to form a feature coefficient matrix, the system can focus on the most significant frequency features, while minimizing the impact of noise, speech patterns, or emotional changes. The resulting feature coefficient matrix can better represent the speech features of each target cluster set, which helps improve the accuracy and efficiency of subsequent analysis.

[0058] Step S374. Calculate the correlation degree value of each pair of feature coefficients in the feature coefficient matrix based on the grey relational analysis algorithm;

[0059] Step S375. Pass the information of each feature coefficient to the adjacent feature coefficients, and combine the information of the adjacent feature coefficients to update the features, and obtain the updated feature coefficients;

[0060] Specifically, by calculating the correlation between feature coefficients in the feature coefficient matrix through grey relational analysis, the inherent correlation between features can be revealed. Furthermore, the information of each feature coefficient is passed to neighboring features for feature updates. This allows features to retain individual differences while incorporating contextual information from neighboring features, thereby improving the overall expressive power and robustness of the features. This not only optimizes the discriminative power of the features but also enhances their noise resistance and robustness, providing more accurate and stable feature descriptions for subsequent speech recognition.

[0061] The formula for feature update is:

[0062]

[0063] Where, x' s x is the updated s-th characteristic coefficient; s h is the s-th characteristic coefficient; sz x represents the correlation value between the s-th feature coefficient and its adjacent z-th feature coefficient; z is the z-th characteristic coefficient adjacent to the s-th characteristic coefficient; n is the number of characteristic coefficients adjacent to the s-th characteristic coefficient.

[0064] Step S376. Perform pooling operation on all updated feature coefficients to obtain multiple local feature information. The local feature information is obtained by aggregating adjacent updated feature coefficients.

[0065] Step S377. Concatenate multiple local feature information to obtain the first speech feature of the processed speech data;

[0066] Specifically, by performing pooling operations on all updated feature coefficients, the information of adjacent feature coefficients can be aggregated to obtain multiple local feature information, thereby simplifying feature representation, reducing noise interference, and preserving key speech information. Then, by feature concatenation, multiple local feature information are integrated into the first speech feature, enabling the system to form a comprehensive and expressive feature vector, which provides a data foundation for subsequent speech analysis and recognition.

[0067] Step S40. Calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than the first set threshold, input the preset target audio data into the preprocessing model for preprocessing to obtain the target speech data. The target audio data is the audio data to be identified.

[0068] Specifically, by continuously adjusting the model parameters of the processing model, when the similarity between the features of the processed historical audio data and the features of the historical speech data reaches a first set threshold, the current processing model is considered to have good noise removal performance and can be used for the preprocessing of subsequent speech data to be recognized.

[0069] Specifically, step S40 includes steps S41, S42, S43, S44, S45, S46, S47, S49, S410, and S411:

[0070] Step S41. Perform feature transformation on the first speech feature and the second speech feature to obtain the first spectrogram and the second spectrogram respectively;

[0071] Step S42. Divide the first spectrogram into blocks to obtain multiple initial modules;

[0072] Step S43. Calculate the sum of the brightness of each primary color in the three primary colors contained in each initial module, and obtain multiple sums of brightness. The sum of brightness is the sum of the brightness of each primary color in the initial module.

[0073] Step S44. Determine any one of the multiple brightness sums as the first brightness sum;

[0074] Step S45. Determine the second brightness and its neighborhood within the first brightness and the neighborhood based on the preset neighborhood radius;

[0075] Step S46. When the number of second brightness sums is not less than a preset minimum number, divide the first brightness sum and all second brightness sums into a first target set;

[0076] Step S47. Delete multiple brightness sums contained in the first target set from multiple brightness sums, and determine any brightness sum from them. Repeat the above steps until multiple first target sets are obtained.

[0077] Step S48. Repeat the above steps to determine the multiple sets of second targets corresponding to the second spectrogram;

[0078] Specifically, a neighborhood radius is set for the brightness sum, meaning that the speech energy of all brightness sums and their corresponding initial modules within this neighborhood is roughly the same. A minimum number is set to define the number of brightness sums within the minimum neighborhood of a core point. A first brightness sum is randomly selected, and the difference between this brightness sum and all brightness sums in the target cluster set is calculated. Points with differences within the neighborhood radius are marked as second brightness sums. When the number of second brightness sums is greater than the minimum number, the first brightness sum is designated as a core point. The first brightness sum and its adjacent second brightness sums are assigned to the same target set. The above steps are repeated until all brightness sums are assigned to their corresponding target sets, thereby determining all target sets. Speech feature information is then determined using target sets as units.

[0079] Step S49. Determine the first correlation based on the difference between the number of items in the first target set and the number of items in the second target set;

[0080] Specifically, by dividing the neighborhood as described above, the differences between the spectrograms corresponding to the first and second speech features can be effectively compared, i.e., the differences in speech features.

[0081] The formula for calculating the first correlation is:

[0082]

[0083] Where S1 is the first correlation; β is the first control parameter; D1 is the number of the first target set; and D2 is the number of the second target set.

[0084] The control parameters are used to control the rate at which the correlation decreases.

[0085] Step S410. Based on the difference between the sum of brightness in the first target set and the sum of brightness in the second target set, determine the second correlation;

[0086] Specifically, step S410 includes steps S4101, S4102, S4103, S4104, S4105, S4106, and S4107:

[0087] Step S4101. Based on a random function, determine a partial set from multiple first target sets as the third target set;

[0088] Specifically, all first target sets can be encoded, each first target set has a corresponding code, a random function is used to randomly determine some codes from multiple codes, and the set corresponding to the code is used as the third target set.

[0089] Step S4102. Based on the location of the third target set, determine the fourth target set in the second spectrum that is closest to the location of the third target set;

[0090] Specifically, the position coordinates of the third target set in the first spectrum map and the position coordinates of each second target set in the second spectrum map are determined. The distance difference between the position coordinates of the third target set and the position coordinates of each second target set in the second spectrum map is calculated, thereby determining the minimum distance difference. The second target set corresponding to the minimum distance difference is the fourth target set.

[0091] Step S4103. Calculate the positional deviation between each third target set and its corresponding fourth target set to obtain the first deviation, which includes the lateral positional deviation and the longitudinal positional deviation;

[0092] Step S4104. Calculate the difference between the sum of brightness in each third target set and the sum of brightness in the corresponding fourth target set, as the second deviation;

[0093] Step S4105. Calculate the sum of the first deviation and the second deviation to obtain the target deviation;

[0094] Step S4106. Calculate the product of the target deviation and the control parameters to obtain the target product;

[0095] Step S4107. Calculate the value of the exponential function with the natural constant as the base and the target product as the exponent to obtain the second correlation;

[0096] Specifically, the formula for calculating the second correlation is:

[0097]

[0098] Where S2 is the second correlation; γ is the second control parameter; D position For positional deviation; D brightness For brightness and deviation;

[0099] The control parameters are used to control the rate at which the correlation decreases.

[0100] Step S411. Determine the similarity based on the first and second correlations;

[0101] Specifically, the similarity calculation formula is as follows:

[0102] S = S1 + S2

[0103] Where S represents similarity; S1 represents primary relevance; and S2 represents secondary relevance.

[0104] Step S50. Input the target speech data into the preset speech recognition model to perform speech recognition and obtain the speech recognition result;

[0105] Specifically, such as Figure 3 As shown, the training process of the speech recognition model relies on the sparrow optimization algorithm. First, the sparrow population is initialized and its fitness is evaluated. Individual sparrows are divided into different species. Then, the different species of sparrows are optimized generation by generation through simulated natural selection, crossover and mutation operations until the optimal solution is found, thereby constructing the corresponding speech recognition model.

[0106] Specifically, step S50 includes steps S51, S52, S53, S54, S55, S56, S57, S58, S59, S510, and S511:

[0107] Step S51. Obtain the historical voice text data corresponding to the historical voice data;

[0108] Specifically, historical voice-text data refers to text data that is automatically identified and manually adjusted based on historical voice data.

[0109] Step S52. Initialize the sparrow population positions based on chaotic mapping to obtain multiple initial position parameters;

[0110] Specifically, using Logistic chaotic mapping to initialize the sparrow population allows for a non-repeating traversal of the sparrow population's state within a certain range, resulting in a relatively even distribution of the sparrow population throughout the search space. This increases the diversity of the initial sparrow population and avoids the sparrow algorithm getting stuck in local optima during the search process.

[0111] Specifically, step S52 includes steps S521, S522, S523, S524, and S525:

[0112] Step S521. Obtain the initial value and branch parameters of each sparrow in the sparrow population. The initial value is within a first set range, and the branch parameters are within a second set range.

[0113] Step S522. Calculate the product of the initial value and the branch parameter to obtain the first value;

[0114] Step S523. Calculate the difference between the preset value and the initial value to obtain the second value;

[0115] Step S524. Calculate the product of the first value and the second value to obtain the chaotic value;

[0116] Step S525. Repeatedly calculate the product of the initial value and the branch parameter, calculate the difference between the preset value and the initial value, and calculate the product of the first value and the second value until the number of chaotic values ​​reaches the set number of times, and use all chaotic values ​​as the initial position parameters.

[0117] Specifically, the formula for calculating the initial position parameters is:

[0118]

[0119] in, Let be the chaos value of sparrow i at the (t+1)th time. Let t be the chaotic value of sparrow i, with a value range of [0,1]; α is the branch parameter that determines whether the chaotic mapping is in a chaotic state, with a value range of [0,4].

[0120] In this embodiment, the chaotic value obtained after mapping each parameter two thousand times is used as the initial position parameter of the sparrow, such as... Figure 2 The figure shown is the parameter mapping diagram after two thousand iterations of the chaotic mapping.

[0121] Step S53. Construct a model based on the initial position parameters to obtain the initial recognition model;

[0122] Step S54. Determine the initial fitness of all sparrows based on their initial position parameters;

[0123] Specifically, assuming the sparrow population has n sparrows, and each sparrow is a solution in the d-dimensional solution space, the fitness of the sparrow population is:

[0124]

[0125] Where F(X) represents the fitness of the sparrow population; x n,d Let f([x] be the position of sparrow n in the population in the d-th dimension; n, 1x n,2 …x n,d ]) represents the individual fitness of sparrow n.

[0126] Step S55. Based on the initial position parameters, classify all sparrows into discoverers, followers, and watchers;

[0127] Step S56. Update the position parameters of the discoverer, follower and vigilant respectively based on the preset position update formula to obtain the updated position parameters;

[0128] Specifically, classifying sparrows into discoverers, followers, and watchers helps simulate the natural selection process, enhances the diversity and adaptability of the group, and dynamically adjusts the sparrows of different roles through the position update formula, which can optimize the search strategy, improve the global search capability, and thus converge to the optimal solution more effectively.

[0129] Considering that global search relies too heavily on the location of the discoverer, we introduce an adaptive weight factor by drawing on the spiral-like ascending search method of iterative optimization in the whale optimization algorithm. This allows the discoverer in the sparrow algorithm to explore the next region with a larger step size in the early stage and a smaller step size for convergence exploration in the later stage after the location is updated.

[0130] In this embodiment of the application, the location update formula for the discoverer is:

[0131]

[0132] in, Let be the j-th dimension position of sparrow i in the population after the (T+1)th iteration; Let be the j-th position of sparrow i in the population after the T-th iteration; δ is the adaptive weight factor; iter maxis the maximum number of iterations for the population; θ is a uniformly distributed random number with a value range of (0, 1]; R2 is the warning value with a value range of [0, 1]; ST is the warning threshold with a value range of [0.5, 1]; Q is a random number following a normal distribution; Z is a 1×d matrix; is the worst position of sparrow i in the population after the T-th iteration; is the optimal position of sparrow i in the population after the T-th iteration; W(T) is an adaptive weight factor, which is a function that decreases with the increase of the iteration number T; A is the amplitude factor of the spiral search; B is the frequency factor of the spiral search; is the upper bound constraint of sparrow i; [[ID=X]] is the lower bound constraint of sparrow i; iter i is the current iteration number of sparrow i.

[0133] When R2 < ST, it indicates that when the warning value is lower than the safety value, there is no predator in the foraging environment, and at this time, the discoverer can conduct extensive searches; when R2 ≥ ST, it indicates that some sparrows in the population have discovered the predator and issued warnings to other sparrows, and all sparrows need to quickly fly to a safe area to forage.

[0134] [[ID=X]] In the embodiments of this application, the position update of the sparrow follower is as follows;

[0135]

[0136] where is the j-th dimension position of sparrow i in the population after the (T + 1)-th iteration; X w is the position of the worst sparrow in the population after the T-th iteration; is the j-th dimension position of sparrow i in the population after the T-th iteration; Q is a random number following a normal distribution; is the position of sparrow p in the population after the (T + 1)-th iteration; K is a uniformly distributed random number with a value range of [-1, 1]; d is the space dimension; is the j-th dimension position of sparrow p in the population after the T-th iteration; n is the number of individuals in the sparrow population;

[0137] When i > n / 2, it indicates that the i-th follower has not obtained food and is in a hungry state. At this time, it needs to fly to other places to forage to obtain more energy; when i ≤ n / 2, it indicates that the i-th follower can obtain food and does not need to fly to other places to forage to obtain energy.

[0138] In order to improve the convergence speed and accuracy of the sparrow algorithm, the step size factors and B are corrected, and the step size factor correction formula is:

[0139] [[ID=4X]]

[0140] Among them, f g f represents the fitness of the best sparrow in the current population. w t represents the fitness of the worst sparrow in the current population; t is the current iteration number; Si is the initialization parameter; rand is a uniformly distributed random factor. B is the step size factor, and the value of B ranges from [0,1].

[0141] The formula for updating the position of the Sparrow Watcher is:

[0142]

[0143] in, Let be the j-th dimension position of sparrow i in the population after the (T+1)th iteration; Let be the j-th dimension position of sparrow i in the population after the T-th iteration; Let be the position of the best sparrow in the population after the Tth iteration; f represents the position of the worst-performing sparrow in the population after the Tth iteration; i f represents the individual fitness of sparrow i; w f represents the fitness of the worst sparrow in the current population. g ε represents the fitness of the best sparrow in the current population; ε is the minimum constant to prevent the denominator from being zero. B is the step size factor, and the value of B ranges from [0,1].

[0144] When f i ≠f g When f indicates that the sparrow is on the edge of the population and is extremely vulnerable to predators; when f i =f g This indicates that sparrows in the middle of the population are also in danger, and they need to approach other sparrows to reduce the risk of being preyed upon.

[0145] Step S57. Determine the current iteration number based on the number of position updates;

[0146] Step S58. Calculate the mutation rate based on the preset mutation rate, the preset minimum mutation rate, the preset maximum number of iterations, and the current number of iterations;

[0147] Step S59. Based on the mutation rate, identify multiple target sparrows from the sparrow population, and optimize the updated position parameters of the multiple target sparrows based on the Cauchy distribution algorithm to obtain the optimized position parameters;

[0148] Specifically, mutation operations can expand the search space of the sparrow population, but not every individual sparrow should perform this operation in every iteration. The target sparrow to be mutated needs to be determined by the mutation rate, so as to optimize the position parameters. While maintaining appropriate population diversity, excessive random perturbation should be avoided, thereby improving the convergence speed and search efficiency of the algorithm.

[0149] Step S510. When the number of vigilants is less than the second set threshold, the target sparrow is determined from the current sparrow population, along with the optimized position parameters and initial fitness of the target sparrow. The target sparrow is the sparrow with the highest fitness in the sparrow population.

[0150] Step S511. Based on historical speech data, historical speech text data, optimized position parameters of the target sparrow, and the initial fitness of the target sparrow, adjust and optimize the preset initial recognition model to obtain a speech recognition model;

[0151] Specifically, considering that a larger number of watchers is beneficial for the algorithm's global search, while a smaller number of watchers is more conducive to the algorithm's rapid local search within a small area, this embodiment sets a higher proportion of watchers in the early stages of the algorithm to enhance the population's global exploration capability, and gradually reduces the proportion of watchers as the number of iterations increases, thereby improving the algorithm's convergence speed.

[0152] The formula for updating the proportion of vigilant individuals is:

[0153]

[0154] Where P is the proportion of vigilants; P0 is the initial proportion of vigilants; T is the current iteration number; iter max P represents the maximum number of iterations for the population. min The minimum pre-set proportion of vigilant individuals;

[0155] When P>P min When P ≤ P, continue the sparrow position update iteration operation until the condition is met; min When the condition is met, the iteration ends. The optimal position and best fitness value are obtained from the global data. The optimal weights and thresholds of the convolutional neural network are determined. The optimal weights and thresholds are then passed back to the convolutional neural network for retraining to obtain the speech recognition model.

[0156] By inputting target speech data into a speech recognition model, speech can be quickly recognized and converted into text. When applied in judicial scenarios, this can help judicial mediators record and process case information more efficiently, thereby improving the quality and efficiency of mediation work.

[0157] Example 2:

[0158] like Figure 4 As shown, this embodiment provides a voice recognition device, which includes:

[0159] The first acquisition unit 10 is used to acquire historical audio data, which includes historical speech data with noise and historical speech data without noise.

[0160] The first input unit 20 is used to input noisy historical speech data into a preset processing model for preprocessing to obtain processed speech data.

[0161] Extraction unit 30 is used to extract features from the processed speech data and the historical speech data without noise, and obtain the first speech feature and the second speech feature respectively;

[0162] The first calculation unit 40 is used to calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than the first set threshold, the preset target audio data is input into the preprocessing model for preprocessing to obtain the target speech data. The target audio data is the audio data to be identified.

[0163] The second input unit 50 is used to input the target speech data into a preset speech recognition model for speech recognition and obtain the speech recognition result.

[0164] In one specific embodiment disclosed in this application, the extraction unit 30 includes:

[0165] The second calculation unit is used to calculate the energy value of the processed speech data in each preset time period to obtain multiple energy values.

[0166] The first unit is used to pause the time period corresponding to the energy value when the energy value is less than the preset energy threshold.

[0167] The first segmentation unit is used to divide the processed speech data based on all pause time periods to obtain multiple sub-audio data, and calculate the energy value of each sub-audio data to obtain multiple target energy values;

[0168] The clustering unit is used to cluster multiple sub-audio data based on clustering algorithms and multiple target energy values ​​to obtain multiple cluster sets. Each cluster set contains multiple sub-audio data with small differences in target energy values.

[0169] The second unit is used to calculate the energy threshold range of all sub-audio data contained in each cluster set based on the Laida criterion, and to take the cluster set corresponding to the smallest threshold range among all threshold ranges as the outlier set.

[0170] The first deletion unit is used to delete outlier sets from all cluster sets to obtain multiple target cluster sets;

[0171] The aggregation unit is used to extract features from the sub-audio data in all target cluster sets, and to aggregate the extracted features to obtain the first speech feature.

[0172] In one specific embodiment disclosed in this application, the aggregation unit includes:

[0173] The first transformation unit is used to perform Fourier transform on each sub-audio data, and input the spectrum obtained after Fourier transform into a Mel filter for processing to obtain multiple energy logarithms. The energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter.

[0174] The second transformation unit is used to perform a cosine transformation on the logarithm of energy to obtain the corresponding multiple Mel frequency cepstral coefficients.

[0175] The third unit is used to take the first twelve coefficients of the multiple Mel frequency cepstral coefficients corresponding to each sub-audio data as feature coefficients, and encode all feature coefficients to obtain the feature coefficient matrix.

[0176] The third calculation unit is used to calculate the correlation degree value of each pair of feature coefficients in the feature coefficient matrix based on the grey relational analysis algorithm.

[0177] The transmission unit is used to transmit the information of each feature coefficient to the adjacent feature coefficients, and combine the information of the adjacent feature coefficients to update the features and obtain the updated feature coefficients.

[0178] The pooling unit is used to perform pooling operations on all updated feature coefficients to obtain multiple local feature information, which is obtained by aggregating adjacent updated feature coefficients.

[0179] The splicing unit is used to splice multiple local feature information to obtain the first speech feature of the processed speech data.

[0180] In one specific embodiment disclosed in this application, the second input unit 50 includes:

[0181] The second acquisition unit is used to acquire historical speech text data corresponding to historical speech data;

[0182] An initialization unit is used to initialize the position of a sparrow population based on a chaotic mapping, and obtain multiple initial position parameters.

[0183] The building unit is used to construct a model based on the initial position parameters to obtain the initial recognition model;

[0184] The first determining unit is used to determine the initial fitness of all sparrows based on the initial position parameters of all sparrows;

[0185] The second division unit is used to classify all sparrows into discoverers, followers, and watchers based on the initial position parameters.

[0186] The update unit is used to update the position parameters of the discoverer, follower and guard respectively based on the preset position update formula to obtain the updated position parameters;

[0187] The second determining unit is used to determine the current iteration number based on the number of position updates;

[0188] The fourth calculation unit is used to calculate the mutation rate based on the preset mutation rate, the preset minimum mutation rate, the preset maximum number of iterations, and the current number of iterations.

[0189] The third determining unit is used to identify multiple target sparrows from the sparrow population based on the mutation rate, and to optimize the updated position parameters of the multiple target sparrows based on the Cauchy distribution algorithm to obtain the optimized position parameters.

[0190] The fourth determining unit is used to determine the target sparrow from the current sparrow population when the number of vigilants is less than the second set threshold, as well as the optimized position parameters and initial fitness of the target sparrow. The target sparrow is the sparrow with the highest fitness in the sparrow population.

[0191] The optimization unit is used to adjust and optimize the preset initial recognition model based on historical speech data, historical speech-text data, optimized position parameters of the target sparrow, and the initial fitness of the target sparrow, so as to obtain a speech recognition model.

[0192] In one specific embodiment disclosed in this application, the initialization unit includes:

[0193] The third acquisition unit is used to acquire the initial value and branch parameters of each sparrow in the sparrow population. The initial value is within a first set range, and the branch parameters are within a second set range.

[0194] The product unit is used to calculate the product of the initial value and the branch parameter to obtain the first value;

[0195] The fifth calculation unit is used to calculate the difference between the preset value and the initial value to obtain the second value;

[0196] The sixth calculation unit is used to calculate the product of the first value and the second value to obtain the chaotic value;

[0197] The first repeating unit is used to repeatedly calculate the product of the initial value and the branch parameter, calculate the difference between the preset value and the initial value, and calculate the product of the first value and the second value until the number of chaotic values ​​reaches the set number of times, and then use all chaotic values ​​as the initial position parameters.

[0198] In one specific embodiment disclosed in this application, the first computing unit 40 includes:

[0199] The conversion unit is used to convert the first speech feature and the second speech feature to obtain the first spectrogram and the second spectrogram, respectively.

[0200] The processing unit is used to divide the first spectrogram into blocks to obtain multiple initial modules;

[0201] The seventh calculation unit is used to calculate the sum of the brightness of each of the three primary colors contained in each initial module, and obtain multiple sums of brightness, which are the sum of the brightness of each primary color in the initial module;

[0202] The fifth determining unit is used to determine any one of the multiple brightness sums as the first brightness sum;

[0203] The sixth determining unit is used to determine the second brightness and its neighborhood within the first brightness and the neighborhood based on a preset neighborhood radius;

[0204] The third division unit is used to divide the first brightness sum and all second brightness sums into a first target set when the number of the second brightness sums is not less than a preset minimum number.

[0205] The second deletion unit is used to delete multiple brightness sums contained in the first target set from multiple brightness sums, and determine any brightness sum from them, repeating the above steps until multiple first target sets are obtained;

[0206] The second repeating unit is used to repeat the above steps to determine the multiple sets of second targets corresponding to the second spectrogram;

[0207] The seventh determining unit is used to determine the first correlation based on the difference between the number of the first target set and the number of the second target set;

[0208] The eighth determining unit is used to determine the second correlation based on the difference between the sum of brightness in the first target set and the sum of brightness in the second target set;

[0209] The ninth determining unit is used to determine the similarity based on the first correlation and the second correlation.

[0210] In one specific embodiment disclosed in this application, the eighth determining unit includes:

[0211] The tenth determining unit is used to determine a partial set from multiple first target sets based on a random function, which serves as the third target set;

[0212] The eleventh determining unit is used to determine the fourth target set in the second spectrum that is closest to the location of the third target set based on the location of the third target set;

[0213] The eighth calculation unit is used to calculate the positional deviation between each third target set and the corresponding fourth target set to obtain the first deviation, which includes the lateral positional deviation and the longitudinal positional deviation.

[0214] The ninth calculation unit is used to calculate the difference between the sum of brightness in each third target set and the sum of brightness in the corresponding fourth target set, as the second deviation;

[0215] The tenth calculation unit is used to calculate the sum of the first deviation and the second deviation to obtain the target deviation;

[0216] The eleventh calculation unit is used to calculate the product of the target deviation and the control parameters to obtain the target product;

[0217] The twelfth calculation unit is used to calculate the value of the exponential function with the natural constant as the base and the target product as the exponent, to obtain the second correlation.

[0218] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.

[0219] Example 3:

[0220] Corresponding to the above method embodiments, this embodiment also provides a speech recognition device. The speech recognition device described below and the speech recognition method described above can be referred to each other.

[0221] Figure 5 This is a block diagram illustrating a speech recognition device 800 according to an exemplary embodiment. Figure 5 As shown, the voice recognition device 800 may include a processor 801 and a memory 802. The voice recognition device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0222] The processor 801 controls the overall operation of the voice recognition device 800 to complete all or part of the steps in the aforementioned voice recognition method. The memory 802 stores various types of data to support the operation of the voice recognition device 800. This data may include, for example, instructions for any application or method operating on the voice recognition device 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 802 or transmitted via communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 805 is used for wired or wireless communication between the voice recognition device 800 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of these. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0223] In an exemplary embodiment, the speech recognition device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the speech recognition method described above.

[0224] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the speech recognition method described above. For example, the computer-readable storage medium may be the memory 802 including the program instructions described above, which may be executed by the processor 801 of the speech recognition device 800 to complete the speech recognition method described above.

[0225] Example 4:

[0226] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the speech recognition method described above.

[0227] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the speech recognition method described in the above method embodiments.

[0228] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.

[0229] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0230] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A speech recognition method, characterized in that, include: Acquire historical audio data, which includes historical speech data with noise and historical speech data without noise; Noisy historical speech data is input into a preset processing model for preprocessing to obtain processed speech data. Feature extraction is performed on the processed speech data and the historical speech data without noise to obtain the first speech feature and the second speech feature, respectively. Calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, input the preset target audio data into the preprocessing model for preprocessing to obtain target speech data. The target audio data is the audio data to be identified. The target speech data is input into a preset speech recognition model for speech recognition to obtain the speech recognition result; The construction of the speech recognition model includes: Obtain the historical voice text data corresponding to the historical voice data; The sparrow population location was initialized using chaotic mapping to obtain multiple initial location parameters; Based on the initial position parameters, a model is constructed to obtain the initial recognition model; Based on the initial position parameters of all sparrows, the initial fitness of all sparrows is determined; Based on the initial position parameters, all sparrows are divided into discoverers, followers, and watchers; The position parameters of the discoverer, the follower, and the vigilant are updated based on a preset position update formula to obtain the updated position parameters. The current iteration number is determined based on the number of position updates. The mutation rate is calculated based on a preset mutation rate, a preset minimum mutation rate, a preset maximum number of iterations, and the current number of iterations. Based on the mutation rate, multiple target sparrows were identified from the sparrow population, and the updated position parameters of the multiple target sparrows were optimized based on the Cauchy distribution algorithm to obtain optimized position parameters. When the number of the watchers is less than the second set threshold, the target sparrow is determined from the current sparrow population, along with the optimized position parameters and initial fitness of the target sparrow. The target sparrow is the sparrow with the highest fitness in the sparrow population. Based on the historical speech data, the historical speech text data, the optimized position parameters of the target sparrow, and the initial fitness of the target sparrow, the preset initial recognition model is adjusted and optimized to obtain the speech recognition model.

2. The speech recognition method according to claim 1, characterized in that... Feature extraction is performed on the processed speech data and the noise-free historical speech data to obtain the first speech feature and the second speech feature, respectively, including: The energy value of the processed speech data is calculated within each preset time period to obtain multiple energy values; When the energy value is less than the preset energy threshold, the time period corresponding to the energy value is taken as the pause time period; The processed speech data is divided into multiple sub-audio data based on all pause time periods, and the energy value of each sub-audio data is calculated to obtain multiple target energy values. Multiple sub-audio data are clustered based on clustering algorithms and multiple target energy values ​​to obtain multiple cluster sets. Each cluster set contains multiple sub-audio data with small differences in target energy values. Based on the Laida criterion, the energy threshold range of all sub-audio data contained in each cluster set is calculated, and the cluster set corresponding to the smallest threshold range among all threshold ranges is taken as the outlier set. The outlier sets in all cluster sets are deleted to obtain multiple target cluster sets; Feature extraction is performed on the sub-audio data in all target cluster sets, and the extracted features are aggregated to obtain the first speech feature.

3. The speech recognition method according to claim 2, characterized in that... Feature extraction is performed on the sub-audio data in all target cluster sets, and the extracted features are aggregated to obtain the first speech feature, including: A Fourier transform is performed on each sub-audio data, and the spectrum obtained after the Fourier transform is input into a Mel filter for processing to obtain multiple energy logarithms. The energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter. Performing a cosine transform on the logarithm of the energy yields multiple corresponding Mel frequency cepstral coefficients; The first twelve coefficients among the multiple Mel frequency cepstral coefficients corresponding to each sub-audio data are taken as feature coefficients, and all feature coefficients are encoded to obtain a feature coefficient matrix; The correlation degree value of each pair of feature coefficients in the feature coefficient matrix is ​​calculated based on the grey relational analysis algorithm; The information of each feature coefficient is passed to the adjacent feature coefficients, and the feature is updated by combining the information of the adjacent feature coefficients to obtain the updated feature coefficients. All updated feature coefficients are pooled to obtain multiple local feature information, which is obtained by aggregating adjacent updated feature coefficients. Multiple local feature information are concatenated to obtain the first speech feature of the processed speech data.

4. The speech recognition method according to claim 1, characterized in that... Based on chaotic mapping, the sparrow population location is initialized, and multiple initial location parameters are obtained, including: Obtain the initial value and branch parameters for each sparrow in the sparrow population, wherein the initial value is within a first set range and the branch parameters are within a second set range; Calculate the product of the initial value and the branch parameter to obtain the first value; Calculate the difference between the preset value and the initial value to obtain the second value; Calculate the product of the first value and the second value to obtain the chaos value; Repeat the calculation of the product of the initial value and the branch parameter, the calculation of the difference between the preset value and the initial value, and the calculation of the product of the first value and the second value until the number of chaotic values ​​reaches the set number of times, and use all chaotic values ​​as the initial position parameters.

5. The speech recognition method according to claim 4, characterized in that... The formula for calculating the initial position parameters is: ; in, For sparrow The chaos value at the (t+1)th iteration; For sparrow The chaotic value of the t-th time has a range of [0,1]. The branch parameter, which determines whether the chaotic mapping is in a chaotic state, has a value range of [0, 4].

6. The speech recognition method according to claim 1, characterized in that... The formula for updating the discoverer's location is: ; ; in, For the first After the next iteration, sparrows in the population The j-th dimension position; For sparrows in the population after the Tth iteration The j-th dimension position; An adaptive weighting factor; This represents the maximum number of iterations for the population. The values ​​are uniformly randomized and range from 1 to 10. ; This is a warning value, and its range is [value range missing]. ; The warning threshold has a range of values. ; These are random numbers that follow a normal distribution. It is a matrix with 1 row and d columns; For the first After the next iteration, sparrows in the population The worst position; For the first After the next iteration, sparrows in the population The optimal position; As an adaptive weighting factor, it varies with the number of iterations. A function that decreases as the value increases; This is the amplitude factor for the spiral search; This is the frequency factor for the spiral search; For sparrow Upper limit constraint; For sparrow The lower bound constraint; For sparrow Current iteration number.

7. A voice recognition device, characterized in that, include: The first acquisition unit is used to acquire historical audio data, which includes historical speech data with noise and historical speech data without noise. The first input unit is used to input noisy historical speech data into a preset processing model for preprocessing to obtain processed speech data. The extraction unit is used to extract features from the processed speech data and the historical speech data without noise, and obtain the first speech feature and the second speech feature, respectively. The first calculation unit is used to calculate the similarity between the first speech feature and the second speech feature. When the similarity is greater than a first set threshold, the preset target audio data is input into the preprocessing model for preprocessing to obtain target speech data. The target audio data is the audio data to be identified. The second input unit is used to input the target speech data into a preset speech recognition model for speech recognition and obtain the speech recognition result; The construction of the speech recognition model includes: The second acquisition unit is used to acquire historical speech text data corresponding to historical speech data; An initialization unit is used to initialize the position of a sparrow population based on a chaotic mapping, and obtain multiple initial position parameters. The building unit is used to construct a model based on the initial position parameters to obtain the initial recognition model; The first determining unit is used to determine the initial fitness of all sparrows based on the initial position parameters of all sparrows; The second division unit is used to classify all sparrows into discoverers, followers, and watchers based on the initial position parameters. The update unit is used to update the position parameters of the discoverer, follower and guard respectively based on the preset position update formula to obtain the updated position parameters; The second determining unit is used to determine the current iteration number based on the number of position updates; The fourth calculation unit is used to calculate the mutation rate based on the preset mutation rate, the preset minimum mutation rate, the preset maximum number of iterations, and the current number of iterations. The third determining unit is used to identify multiple target sparrows from the sparrow population based on the mutation rate, and to optimize the updated position parameters of the multiple target sparrows based on the Cauchy distribution algorithm to obtain the optimized position parameters. The fourth determining unit is used to determine the target sparrow from the current sparrow population when the number of vigilants is less than the second set threshold, as well as the optimized position parameters and initial fitness of the target sparrow. The target sparrow is the sparrow with the highest fitness in the sparrow population. The optimization unit is used to adjust and optimize the preset initial recognition model based on historical speech data, historical speech-text data, optimized position parameters of the target sparrow, and the initial fitness of the target sparrow, so as to obtain a speech recognition model.

8. The speech recognition device according to claim 7, characterized in that, The extraction unit includes: The second calculation unit is used to calculate the energy value of the processed speech data in each preset time period to obtain multiple energy values. The first unit is used to take the time period corresponding to the energy value as a pause time period when the energy value is less than a preset energy threshold. The first segmentation unit is used to divide the processed speech data based on all pause time periods to obtain multiple sub-audio data, and calculate the energy value of each sub-audio data to obtain multiple target energy values; The clustering unit is used to cluster multiple sub-audio data based on clustering algorithms and multiple target energy values ​​to obtain multiple cluster sets. Each cluster set contains multiple sub-audio data with small differences in target energy values. The second unit is used to calculate the energy threshold range of all sub-audio data contained in each cluster set based on the Laida criterion, and to take the cluster set corresponding to the smallest threshold range among all threshold ranges as the outlier set. The first deletion unit is used to delete the outlier sets in all cluster sets to obtain multiple target cluster sets; The aggregation unit is used to extract features from the sub-audio data in all target cluster sets, and to aggregate the extracted features to obtain the first speech feature.

9. The speech recognition device according to claim 8, characterized in that, The aggregation unit includes: The first transformation unit is used to perform Fourier transform on each sub-audio data, and input the spectrum obtained after Fourier transform into a Mel filter for processing to obtain multiple energy logarithms. The energy logarithm is the logarithm of the energy value of each frequency band output by the Mel filter. The second transformation unit is used to perform a cosine transformation on the energy logarithm to obtain a plurality of corresponding Mel frequency cepstral coefficients. The third unit is used to take the first twelve coefficients of the multiple Mel frequency cepstral coefficients corresponding to each sub-audio data as feature coefficients, and encode all feature coefficients to obtain a feature coefficient matrix. The third calculation unit is used to calculate the correlation degree value of each pair of feature coefficients in the feature coefficient matrix based on the grey relational analysis algorithm. The transmission unit is used to transmit the information of each feature coefficient to the adjacent feature coefficients, and combine the information of the adjacent feature coefficients to update the features and obtain the updated feature coefficients. The pooling unit is used to perform pooling operations on all updated feature coefficients to obtain multiple local feature information, which is obtained by aggregating adjacent updated feature coefficients. The splicing unit is used to splice multiple local feature information to obtain the first speech feature of the processed speech data.

Citation Information

Patent Citations

  • Speech recognition model training method, speech recognition method and electronic equipment

    CN114582330A