Speech transcription acceleration method based on artificial intelligence

Through the speech transfer acceleration method based on artificial intelligence, voice data is preprocessed and enhanced, voice features are extracted, and optimized transcribed text is generated using an adaptive dynamic text optimization algorithm. The problems of background noise interference, low signal processing efficiency and inaccurate semantic understanding in the prior art are solved, and efficient and accurate speech transfer effects are achieved.

CN119091861BActive Publication Date: 2025-05-13NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411149307.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-05-13
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The existing speech transfer system has shortcomings in the problems of background noise interference, low signal processing efficiency and inaccurate semantic understanding, which affects the accuracy and efficiency of speech transfer.

Method used

Using artificial intelligence-based speech transcription acceleration method, voice features are extracted by preprocessing and enhancing speech data, and the optimized transcribed text is generated using an adaptive dynamic text optimization algorithm to improve the translation efficiency and accuracy.

Benefits of technology

It effectively reduces background noise and interference, improves signal quality and translation efficiency, and the generated translation text is coherent in semantics and clear in logic, which significantly improves the accuracy and speed of speech translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091861B_ABST
    Figure CN119091861B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech transcription, and in particular to a speech transcription acceleration method based on artificial intelligence, comprising the following steps: (S1) obtaining original speech data, pre-processing and then enhancing the obtained original speech data, extracting features of the enhanced speech data, obtaining speech features, performing speech recognition based on the speech features, and obtaining recognition results; (S2) generating a preliminary transcription text according to the recognition result, optimizing the preliminary written text by an adaptive dynamic text optimization algorithm, obtaining an optimized transcription text, and optimizing the transcription efficiency by an optimization acceleration algorithm during the transcription process. The speech transcription acceleration method based on artificial intelligence disclosed by the present invention reduces background noise and other interferences, and improves the accuracy and speed of the final written text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech transcription, and in particular to a speech transcription acceleration method based on artificial intelligence. Background Art

[0002] With the rapid development and widespread application of speech recognition technology, speech transcription has become a key technology in many fields. However, existing speech transcription systems still face many technical challenges in practical applications, including background noise interference, inefficient signal processing, inaccurate semantic understanding, etc. These problems seriously affect the accuracy and efficiency of speech transcription, and restrict its promotion and application in a wider range of application scenarios.

[0003] In practical applications, speech signals are often accompanied by a large amount of background noise and interference, which makes it difficult for speech recognition systems to accurately extract effective speech features. Traditional signal processing methods, such as bandpass filters and fast Fourier transforms (FFTs), have limited effectiveness in dealing with complex noise environments. In addition, the noise spectrum estimation of silent segments and initial segments is inaccurate, resulting in poor noise reduction, which further affects the accuracy of speech recognition. In addition, when generating preliminary transcriptions, speech recognition systems usually only focus on local features and ignore the semantic relationship between words, resulting in incoherent semantics and unclear logic in the generated transcriptions.

[0004] In addition to the technical problems mentioned above, the existing technology also has technical problems such as poor accuracy and low efficiency in speech transcription. Summary of the invention

[0005] In order to overcome the deficiencies of the prior art in the field of speech transcription mentioned in the background technology, the present invention provides a speech transcription acceleration method based on artificial intelligence.

[0006] To achieve the above purpose, the present invention discloses an artificial intelligence-based speech transcription acceleration method, comprising the following steps:

[0007] (S1) obtaining original speech data, preprocessing and then enhancing the obtained original speech data, performing feature extraction on the enhanced speech data to obtain speech features, performing speech recognition based on the speech features to obtain a recognition result;

[0008] (S2) generating a preliminary transcription text according to the recognition result, optimizing the preliminary transcription text by using an adaptive dynamic text optimization algorithm to obtain an optimized transcription text, and optimizing the transcription efficiency by using an optimization acceleration algorithm during the transcription process.

[0009] Preferably, in step (S1), the method for preprocessing the original speech data comprises the following steps:

[0010] (A1) dividing the original speech signal into frames of fixed length to obtain a framed speech signal;

[0011] (A2) applying a windowing function to each frame of speech signal to obtain a windowed speech frame signal;

[0012] (A3) performing a fast Fourier transform on each frame of the windowed speech frame signal to convert the time domain signal into the frequency domain to obtain a speech frame represented in the frequency domain;

[0013] (A4) in a silent segment or an initial segment, calculating a noise spectrum to obtain an estimated background noise spectrum;

[0014] (A5) filtering out noise from the speech frame and the background noise spectrum to obtain a frequency domain representation after noise reduction;

[0015] (A6) performing an inverse fast Fourier transform on the denoised frequency domain signal, converting the frequency domain signal back to the time domain, obtaining a denoised time domain speech frame signal, and re-joining the denoised time domain speech frame signal into a continuous speech signal to form pre-processed speech data.

[0016] Preferably, in step (S1), the speech data enhancement processing method comprises the following steps:

[0017] (B1) performing short-time spectrum conversion on the preprocessed speech data to convert the time domain signal into a frequency domain signal;

[0018] (B2) Adaptively adjust the weight of the frequency domain signal to eliminate echo and residual noise;

[0019] (B3) using the signal after adaptive weight adjustment to suppress echo, remove the echo component in the signal, and obtain the echo suppressed signal by subtracting it from the original signal;

[0020] (B4) performing enhancement processing on the signal after echo suppression to further enhance the clarity of the speech signal by reducing background noise and enhancing signal details;

[0021] (B5) The enhanced frequency domain signal is converted back to the time domain to obtain the enhanced speech signal.

[0022] Preferably, in step (S2), the method for generating a preliminary transcription text according to the recognition result comprises the following steps:

[0023] (C1) Mapping the speech recognition result into a word embedding vector according to the recognition result;

[0024] (C2) Using a network of sequence units to extract temporal features of the mapped feature vector;

[0025] (C3) construct a multi-layer structure to optimize the extracted temporal features;

[0026] (C4) A log-likelihood optimization method is used to generate preliminary transcription text.

[0027] Preferably, in step (C3), the method for optimizing the extracted time features comprises the following steps:

[0028] (D1) Processing time series features through multi-head self-attention mechanism;

[0029] (D2) Use a feedforward neural network to further process the attention features;

[0030] (D3) Generate text based on output features.

[0031] Preferably, the method for optimizing the preliminary transcription text using the adaptive text optimization algorithm comprises the following steps:

[0032] (E1) For each word, a context window is introduced and word embedding is used to convert the word into a vector representation for each context window;

[0033] (E2) performing adaptive weighted averaging on each context vector;

[0034] (E3) dynamically adjusting the context vector after adaptive weighted averaging through multi-layer perception to generate an optimized word vector;

[0035] (E4) Re-map the optimized word vector back to the word space.

[0036] The present invention has the following beneficial effects:

[0037] 1. Convert the time domain signal into a frequency domain signal, further reduce the background noise and other interference through bandpass filter and adaptive filtering technology, dynamically adjust the filter weight, optimize the signal quality, and make the frequency domain representation clearer after noise reduction; the time domain speech frame signal after noise reduction is re-spliced ​​to ensure the smooth transition and continuity of the signal.

[0038] 2. By dividing the preliminary transcription text into multiple subsequences, the parallel computing processor processes multiple subtasks simultaneously, effectively improving the overall processing efficiency of the system, expanding and boundary processing the subsequences, ensuring the integrity of the context information, and thus improving the robustness of the processing.

[0039] 3. Through the adaptive dynamic text optimization algorithm, the preliminary transcription text is analyzed in context and dynamically adjusted, so that the semantics and position of each word in the context are fully understood, thereby generating an optimized transcription text; by dynamically adjusting the preliminary transcription text, the recognition results are further optimized, the errors generated in the recognition process are reduced, and the accuracy and speed of the final transcription text are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flow chart of the artificial intelligence-based speech transcription acceleration method of the present invention. DETAILED DESCRIPTION

[0041] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0043] The specific scheme of the speech transcription acceleration method based on artificial intelligence provided by the present invention is described in detail below with reference to the accompanying drawings.

[0044] See attached Figure 1 , which shows a method for accelerating speech transcription based on artificial intelligence provided by an embodiment of the present invention, the method comprising the following steps:

[0045] (S1) obtaining original speech data, preprocessing and then enhancing the obtained original speech data, performing feature extraction on the enhanced speech data to obtain speech features, performing speech recognition based on the speech features to obtain recognition results.

[0046] Voice data is collected using a data collection device such as a microphone to obtain raw voice data, and the voice data is preprocessed to obtain preprocessed voice data.

[0047] The method for preprocessing the original speech data comprises the following steps:

[0048] (A1) The original speech signal is divided into frames of fixed length (for example, each frame contains 25 milliseconds of speech data, and the frames overlap by 10 milliseconds) to obtain a framed speech signal.

[0049] (A2) Applying a windowing function (such as a Hamming window or a Hanning window) to each frame of speech signal to reduce spectrum leakage and obtain a windowed speech frame signal.

[0050] (A3) Performing a fast Fourier transform on each windowed speech frame signal to convert the time domain signal into the frequency domain to obtain a speech frame represented in the frequency domain.

[0051] (A4) In the silent segment or the initial segment, the noise spectrum is calculated using the average value or minimum value method to obtain an estimated background noise spectrum.

[0052] (A5) Based on the speech frames represented in the frequency domain and the estimated background noise spectrum, a bandpass filter is applied to the frequency domain representation of each frame to filter out the noise frequency band and retain the speech frequency band. An adaptive filtering technique can be used to adjust the filter parameters according to the noise spectrum to obtain a frequency domain representation after noise reduction.

[0053] (A6) performing an inverse fast Fourier transform on the denoised frequency domain signal, converting the frequency domain signal back to the time domain, obtaining a denoised time domain speech frame signal, re-joining the denoised time domain speech frame signal into a continuous speech signal, and in particular, performing a weighted average process on the overlapping parts between frames to ensure a smooth transition of the signal, thereby obtaining denoised speech data. After the above process, the pre-processed speech data is obtained.

[0054] The speech data enhancement processing method comprises the following steps:

[0055] (B1) Perform short-time spectrum conversion on the pre-processed speech data to convert the time domain signal into a frequency domain signal for subsequent processing. The specific formula is as follows:

[0056]

[0057] In the above formula, Y(f, t) is the frequency domain signal, which represents the signal amplitude and phase at time t and frequency f, y(n) is the time domain signal, which is the value of the preprocessed speech data at time point n, h(nt) is the window function, which is used to process the signal in segments to reduce the boundary effect, N is the number of points of short-time spectrum conversion, which determines the resolution of the frequency domain signal, and j is the imaginary unit; is the Gaussian weighting function, which indicates the weighting effect of the preprocessed speech data near time t. σ is the standard deviation, which indicates the width of the time window, determines the width of the Gaussian weighting function, and affects the range of the weighting effect. It is a parameter selected according to the application scenario. M is the number of points in the Gaussian weighting part, which indicates the number of sampling points used for Gaussian weighting calculation.

[0058] (B2) Adaptively adjust the weight of the frequency domain signal to eliminate echo and residual noise. Optimize signal quality by dynamically adjusting the filter weight. The specific weight formula is as follows:

[0059]

[0060] Among them, W k is the current filter weight matrix, η is the step size factor, which is used to control the speed of weight update, Y(f, t) is the frequency domain signal, ∈(t) is the error signal, that is, the difference between the expected signal and the actual signal, α is the stability factor, which prevents the denominator from being zero, and Y * (f, t) represents the conjugate complex number of Y(f, t), and δ is an additional adjustment factor used to adjust the amplitude of signal processing.

[0061] (B3) Using the adaptive weighted signal W k Y(f, t) is used to suppress the echo and remove the echo component in the signal. The echo-suppressed signal y is obtained by subtracting it from the original signal. res (t). The formula is as follows:

[0062]

[0063] Among them, y res (t) is the signal after echo suppression, which represents the clear speech signal after filtering and echo suppression processing, λ is the adjustment factor used to adjust the sum of weight squares, and I is the cumulative weight and the number of input signals.

[0064] (B4) The signal after echo suppression is enhanced to further enhance the clarity of the speech signal by reducing background noise and enhancing signal details. The specific formula is as follows:

[0065]

[0066] Among them, Y aug (f, t) is the enhanced frequency domain signal, |N(f, t)| is the noise estimate, which represents the frequency domain representation of the background noise, β is the control parameter in log spectrum subtraction, which is used to adjust the strength of noise suppression, and γ is the threshold parameter, which determines under what circumstances the enhancement processing is performed. is a stabilizing factor to prevent the denominator from being zero, is the number of previous and next frames considered to estimate the average energy, is the calculation range of the sum of squares of weights.

[0067] (B5) Convert the enhanced frequency domain signal back to the time domain to obtain the enhanced speech signal y aug (t), the specific formula is as follows:

[0068]

[0069] Among them, y aug(t) is the time domain signal after enhancement processing, which represents the clear speech signal obtained after the final processing.

[0070] The existing convolutional neural network is used to extract features of the enhanced speech data. The existing Mel-frequency cepstral coefficients or Mel-frequency cepstral coefficients are used to extract frequency domain features of the enhanced speech data, generate speech feature vectors, and obtain speech features.

[0071] The extracted speech features are input into the long short-term memory network and the Transformer-based hybrid model, and then undergo feature preprocessing, convolution layer processing, LSTM layer time series processing, Transformer layer global feature extraction processing, and fully connected layer classification processing to achieve speech recognition and output the recognition results.

[0072] (S2) generating a preliminary transcription text according to the recognition result, optimizing the preliminary transcription text by using an adaptive dynamic text optimization algorithm to obtain an optimized transcription text, and optimizing the transcription efficiency by using an optimization acceleration algorithm during the transcription process.

[0073] This method of generating a preliminary transcription text based on the recognition result includes the following steps:

[0074] (C1) According to the recognition result R, R = {r1, r2, ..., r n}, where r i Represents the i-th recognized word or phoneme. The word mapping technology is used to map the speech recognition result into a word embedding vector, so as to convert the recognition result into a feature vector form that is easy to process. The specific implementation formula is as follows:

[0075]

[0076] Among them, F is the mapped feature vector; We is the word mapping matrix, which is used to convert the speech recognition result into the feature vector; b e is a bias vector used to adjust the nonlinear characteristics of the mapping.

[0077] (C2) In order to capture the temporal dependency of speech data, a network of sequence units is used for processing to extract the temporal features of the mapped feature vector, which can be expressed as:

[0078]

[0079] Among them, S t is the time series feature at time t, F t is the eigenvector at time t, h t-1 It is the hidden state of the previous moment. The sequence unit can capture the dependency of the speech signal in the time dimension, thereby generating an ordered feature S.

[0080] (C3) In order to further optimize the temporal features, a multi-layer structure is constructed, which includes a multi-head self-attention mechanism and a feedforward neural network. Specifically,

[0081] (D1) The time series features are processed through the multi-head self-attention mechanism, and the formula is:

[0082]

[0083] Among them, Q is the query matrix, which represents the input features for which attention weights need to be calculated; K is the key matrix, which represents the reference features used to calculate attention weights; V is the value matrix, which represents the features QK that need to be weighted summed. T ; is the dot product of the query matrix and the key matrix, used to calculate the attention score, T is the transpose; ||Q|| is the norm of the query matrix, indicating the size or magnitude of the query matrix; ||K|| is the norm of the key matrix, indicating the size or magnitude of the key matrix; is the scaling factor, d k is the dimension size of the key matrix, which is used to prevent the dot product result from being too large, causing the gradient of the softmax function to disappear; ||V|| is the norm of the value matrix, which indicates the size or amplitude of the value matrix; It is a nonlinear transformation used to adjust the magnitude of the value matrix; Head i is the output of the ith attention head, representing the calculation result of a single attention head; Concat is a concatenation operation that concatenates the outputs of all attention heads together to form a long vector; W O is the output weight matrix, which is used to perform linear transformation on the concatenated vector; h is the number of attention heads, which indicates the number of attention heads calculated in parallel in the multi-head attention mechanism; Log(||head i ||+1) is a logarithmic transformation, which is used to adjust the output amplitude of each attention head and increase nonlinear processing; the multi-head self-attention mechanism can capture key features in different contexts, thereby generating attention features A=MultiHead(Q, K, V).

[0084] (D2) The attention features are further processed using a feedforward neural network, as follows:

[0085]

[0086] Among them, A is the attention feature, W1 is the first layer weight matrix of the feedforward neural network, W2 is the second layer weight matrix of the feedforward neural network, b1 is the bias term of the first layer of the feedforward neural network, and b2 is the bias term of the second layer of the feedforward neural network. The feedforward neural network can map the attention feature to a higher dimensional feature space, thereby generating the output feature O = FFN (A).

[0087] (D3) Generate text based on output feature O.

[0088] (C4) The log-likelihood optimization method is used to generate preliminary transcription text, and the formula is:

[0089]

[0090] Among them, C * is the generated preliminary transcription text, P(C|O) is the probability of generating the entire text C based on the output feature O, ||P(C|O)|| is the norm C of the generation probability P(C|O), is the text representation in the probability distribution, P(c i |O) is to generate the i-th word c based on the output feature O i The probability of |C| represents the number of text elements.

[0091] After the preliminary transcription is generated, it is further optimized through the adaptive dynamic text optimization algorithm to obtain the optimized transcription. The adaptive dynamic text optimization algorithm performs context analysis and dynamic adjustment on the preliminary transcription to achieve the optimized transcription. The specific implementation process is as follows:

[0092] (E1) In order to use context information to more comprehensively understand the semantics and position of each word for context analysis, a context window L is introduced to analyze the position of each word C. i Perform a comprehensive analysis of the context information; suppose the context window size is 2k+1 (including the current word and the k words before and after it), then for the i-th word, its context window L i It is expressed as:

[0093] L i =[l i-k , ..., l i , ..., l i+k ];

[0094] If ik<1 or i+k>|C * |, then fill in the window, |C * | is the length of the preliminary transcription. For each context window L i , use word embeddings (such as Word2Vec or GloVe) to convert words into vector representations.

[0095] (E2) Obtain the context vector representation, and perform adaptive weighted averaging on each context vector to highlight the influence of important words. The adaptive weighted context vector U i It is expressed as:

[0096]

[0097] Among them, Ui is the adaptively weighted context vector; is the weight of the i+jth word, which is based on the inner product and distance weight, combining similarity and distance information; is the secondary weight of the i+jth word, combining the inner product and the logarithmic transformation of the distance; It is word c i+j The word embedding vector of .

[0098]

[0099] in, It is the dimension of the word embedding vector, which is used to scale the inner product result, adjust the importance of words in the context through adaptive weights, and enhance the influence of key words.

[0100] (E3) Adaptively weight the context vector U through a multi-layer perceptron (MLP) i Perform dynamic adjustments to generate optimized word vector U′ i :

[0101] U′ i =MLP(U i ;θ);

[0102] Here, θ is the set of parameters of the MLP, including the weight matrix and bias.

[0103] (E4) The optimized word vector U′ i Re-map back to the word space to generate optimized transcription text (C * )′. Assume that the mapping function is g (using Softmax function), then the optimized transcription text (C * )′ is expressed as:

[0104]

[0105] During the transcription process, the transcription efficiency is optimized by optimizing the acceleration algorithm. The implementation of the optimization acceleration algorithm includes:

[0106] The preliminary transcription text is divided into P subsequences, and each processor processes a subsequence independently. Let P be the number of processors, then the length of the subsequence processed by each processor is Where |C *| is the length of the preliminary transcribed text. Since the context information of words at the subsequence boundary may be lost during the text segmentation process, it is necessary to perform boundary processing on each subsequence after segmentation to ensure the integrity of the context information. Add q words before and after each subsequence to form an extended subsequence. If the extension exceeds the text boundary, zero padding is performed. The q value is determined by the specific scenario and will not be elaborated on in detail. This process ensures the integrity of the context information of each subsequence during boundary processing, thereby improving the optimization effect.

[0107] After the above processing, an adaptive dynamic text optimization algorithm is used for optimization processing to achieve parallel optimization and further improve transcription efficiency and text quality.

[0108] In summary, a speech transcription acceleration method based on artificial intelligence has been completed.

[0109] The order of the embodiments of the invention is for description only and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0110] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0111] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A speech transcription acceleration method based on artificial intelligence, characterized in that: The following steps are involved: (S1) obtaining original speech data, preprocessing and then enhancing the obtained original speech data, performing feature extraction on the enhanced speech data to obtain speech features, performing speech recognition based on the speech features to obtain a recognition result; (S2) generating a preliminary transcription text according to the recognition result, optimizing the preliminary transcription text by using an adaptive dynamic text optimization algorithm to obtain an optimized transcription text, and optimizing the transcription efficiency by using an optimization acceleration algorithm during the transcription process; The method for optimizing the preliminary transcription text using the adaptive dynamic text optimization algorithm comprises the following steps: (E1) For each word, a context window is introduced and word embedding is used to convert the word into a vector representation for each context window; (E2) performing adaptive weighted averaging on each context vector; (E3) dynamically adjusting the context vector after adaptive weighted averaging through multi-layer perception to generate an optimized word vector; (E4) Re-map the optimized word vector back to the word space.

2. The artificial intelligence-based speech transcription acceleration method according to claim 1, characterized in that: In step (S1), the method for preprocessing the original speech data comprises the following steps: (A1) dividing the original speech signal into frames of fixed length to obtain a framed speech signal; (A2) applying a windowing function to each frame of speech signal to obtain a windowed speech frame signal; (A3) performing a fast Fourier transform on each frame of the windowed speech frame signal to convert the time domain signal into the frequency domain to obtain a speech frame represented in the frequency domain; (A4) in a silent segment or an initial segment, calculating a noise spectrum to obtain an estimated background noise spectrum; (A5) filtering out noise from the speech frame and the background noise spectrum to obtain a frequency domain representation after noise reduction; (A6) performing an inverse fast Fourier transform on the denoised frequency domain signal, converting the frequency domain signal back to the time domain, obtaining a denoised time domain speech frame signal, and re-joining the denoised time domain speech frame signal into a continuous speech signal to form pre-processed speech data.

3. The artificial intelligence-based speech transcription acceleration method according to claim 1, characterized in that: In step (S1), the speech data enhancement processing method comprises the following steps: (B1) performing short-time spectrum conversion on the preprocessed speech data to convert the time domain signal into a frequency domain signal; (B2) Adaptively adjust the weight of the frequency domain signal to eliminate echo and residual noise; (B3) using the signal after adaptive weight adjustment to suppress echo, remove the echo component in the signal, and obtain the echo suppressed signal by subtracting it from the original signal; (B4) performing enhancement processing on the signal after echo suppression to further enhance the clarity of the speech signal by reducing background noise and enhancing signal details; (B5) The enhanced frequency domain signal is converted back to the time domain to obtain the enhanced speech signal.

4. The artificial intelligence-based speech transcription acceleration method according to claim 1, characterized in that: In step (S2), the method for generating a preliminary transcription text according to the recognition result comprises the following steps: (C1) Mapping the speech recognition result into a word embedding vector according to the recognition result; (C2) Using a network of sequence units to extract temporal features of the mapped feature vector; (C3) construct a multi-layer structure to optimize the extracted temporal features; (C4) A log-likelihood optimization method is used to generate preliminary transcription text.

5. The artificial intelligence-based speech transcription acceleration method according to claim 4 is characterized in that: In step (C3), the method for optimizing the extracted time features comprises the following steps: (D1) Processing time series features through multi-head self-attention mechanism; (D2) Use a feedforward neural network to further process the attention features; (D3) Generate text based on output features.

Citation Information

Patent Citations

  • Speech recognition result correction method, device and equipment, and storage medium

    CN107293296A

  • Self-adaptive speech recognition method and system

    CN117558278A