Text generation method and device, equipment and storage medium

CN117831536BActive Publication Date: 2026-09-25IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311766209.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2026-09-25
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

[0004]本申请提供一种文本生成方法、装置、设备及存储介质,用以解决现有技术中在进行辅助语言学习时产生的效率低、效果差的缺陷,实现了快速、准确的完成语言学习

Benefits of technology

[0042]本申请实施例提供了一种文本生成方法、装置、设备及存储介质,通过获取包含特定语种语音内容的音频数据,并提取音频数据的音频特征;基于音频特征,识别特定语种语音内容对应的语音文本以及与语音文本对应的音标,可见,本申请通过提取音频数据的语音文本和语音文本对应的音标,以辅助用户如何正确的发音;进而,基于语音文本和音标,生成语言学习讲解文本,语言学习讲解文本包括语音文本、音标,以及学习语音文本和音标的提示信息,可见,本申请通过生成包括有语音文本、音标,以及学习语音文本和音标的提示信息的语言学习讲解文本,使用户充分的理解音频数据,并同时快速、准确的完成语言学习,提高了学习效率,达到了良好的学习效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117831536B_ABST
    Figure CN117831536B_ABST
Patent Text Reader

Abstract

The application provides a text generation method and device, equipment and a storage medium. The method comprises: acquiring audio data containing specific language voice content, and extracting audio features of the audio data; based on the audio features, identifying voice text corresponding to the specific language voice content and phonetic symbols corresponding to the voice text; and based on the voice text and the phonetic symbols, generating language learning explanation text, the language learning explanation text comprising the voice text, the phonetic symbols, and prompt information for learning the voice text and the phonetic symbols. The application is used to solve the defects of low efficiency and poor effect in the prior art when auxiliary language learning is performed, and fast and accurate language learning is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of speech recognition and natural language processing technology, and in particular to a text generation method, apparatus, device and storage medium. Background Technology

[0002] With the development of science and technology, computer-assisted language learning has become an important way for humans to learn languages, such as in the fields of learning less commonly taught languages ​​and early childhood language learning.

[0003] Currently, there are two main types of assisted language learning: one is to help learners master the language through detailed explanations by humans, which is time-consuming and labor-intensive, and may not necessarily yield good results; the other is to use related speech evaluation systems to judge whether the pronunciation is standard. This method does not take into account whether the learner understands the process of pronunciation, but directly evaluates the quality of pronunciation, which is not very learner-friendly. Summary of the Invention

[0004] This application provides a text generation method, apparatus, device, and storage medium to address the shortcomings of low efficiency and poor results in the prior art when conducting assisted language learning, thereby achieving fast and accurate language learning.

[0005] According to a first aspect of the embodiments of this application, a text generation method is provided, comprising:

[0006] Acquire audio data containing speech content in a specific language, and extract the audio features of the audio data;

[0007] Based on the audio features, identify the speech text corresponding to the speech content of the specific language and the phonetic symbols corresponding to the speech text;

[0008] Based on the spoken text and the phonetic symbols, a language learning explanation text is generated, which includes the spoken text, the phonetic symbols, and prompts for learning the spoken text and the phonetic symbols.

[0009] According to the text generation method provided in the first aspect of the embodiments of this application, the step of identifying the speech text corresponding to the speech content of the specific language and the phonetic symbols corresponding to the speech text based on the audio features includes:

[0010] The audio features are input into the speech recognition model to obtain the speech text corresponding to the speech content of the specific language and the phonetic symbols aligned with the speech text.

[0011] The speech recognition model is obtained through training by aligning speech text with phonetic symbols and training by recognizing speech text and phonetic symbols.

[0012] According to the text generation method provided in the first aspect of the embodiments of this application, the training process of the speech recognition model includes:

[0013] Obtain audio feature samples, text samples, and phonetic symbol samples;

[0014] The audio feature samples and the phonetic symbol samples are input into the speech recognition model, so that the speech recognition model processes the audio feature samples and the phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol samples, and generates predicted text and predicted phonetic symbols based on the audio feature samples, the aligned audio feature samples and the phonetic symbol samples;

[0015] The text sample and the predicted text are compared, and the phonetic symbol sample and the predicted phonetic symbol are compared to obtain the comparison result;

[0016] Based on the comparison results, the computational parameters of the speech recognition model are adjusted.

[0017] According to the text generation method provided in the first aspect of the embodiments of this application, the step of processing the audio feature samples and the phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol feature samples includes:

[0018] The phonetic symbol samples are encoded to obtain phonetic symbol feature samples;

[0019] Calculate the product of the audio feature sample and the transpose of the phonetic symbol feature sample to obtain the first fusion result;

[0020] Calculate the product of the first fusion result and the phonetic feature sample to obtain a phonetic feature sample aligned with the audio feature sample.

[0021] According to the text generation method provided in the first aspect of the embodiments of this application, the step of generating predicted text and predicted phonetic symbols based on the audio feature samples, the aligned audio feature samples, and the phonetic symbol feature samples includes:

[0022] Based on the audio feature samples, generate the predicted text and the first predicted phonetic symbol corresponding to the audio feature samples;

[0023] Based on the aligned audio feature samples and phonetic symbol feature samples, a second predicted phonetic symbol is generated;

[0024] The comparison of the text sample and the predicted text, and the comparison of the phonetic symbol sample and the predicted phonetic symbol, to obtain the comparison results include:

[0025] By comparing the text sample with the predicted text, a first comparison result is obtained;

[0026] By comparing the phonetic symbol sample with the first predicted phonetic symbol, a second comparison result is obtained;

[0027] By comparing the phonetic symbol sample with the second predicted phonetic symbol, a third comparison result is obtained;

[0028] The step of adjusting the computational parameters of the speech recognition model based on the comparison results includes:

[0029] The calculation parameters are adjusted based on the first comparison result, the second comparison result, and the third comparison result.

[0030] According to the text generation method provided in the first aspect of the embodiments of this application, the step of generating language learning explanation text based on the speech text and the phonetic symbols includes:

[0031] Create key information, wherein the key information is used to indicate the structural and / or content requirements for generating language learning explanation texts;

[0032] Based on the key information, the speech text, and the phonetic symbols, a task prompt instruction is constructed; the task prompt instruction is used to prompt the large language model to generate language learning explanation text containing the speech text and the phonetic symbols according to the key information;

[0033] The task prompts are input into a pre-trained large language model, which then generates language learning explanation text.

[0034] According to the text generation method provided in the first aspect of the embodiments of this application, the language learning explanation text includes each word of the speech text and the phonetic symbols corresponding to each word;

[0035] Learning the audio text and the phonetic symbols includes learning the pronunciation of the entire audio text by practicing the pronunciation of each word segment.

[0036] According to a second aspect of the embodiments of this application, a text generation apparatus is provided, comprising:

[0037] The feature extraction module is used to acquire audio data containing speech content in a specific language and extract the audio features of the audio data.

[0038] The speech recognition module is used to identify the speech text corresponding to the speech content of the specific language and the phonetic symbols corresponding to the speech text based on the audio features;

[0039] The text generation module is used to generate language learning explanation text based on the speech text and the phonetic symbols. The language learning explanation text includes the speech text, the phonetic symbols, and prompts for learning the speech text and the phonetic symbols.

[0040] According to a third aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the text generation method as described above.

[0041] According to a fourth aspect of the embodiments of this application, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the text generation method as described above.

[0042] This application provides a text generation method, apparatus, device, and storage medium. It acquires audio data containing speech content in a specific language and extracts audio features from the audio data. Based on these audio features, it identifies the speech text corresponding to the speech content in the specific language and the corresponding phonetic symbols. Thus, this application assists users in correctly pronouncing speech by extracting the speech text and corresponding phonetic symbols from the audio data. Furthermore, based on the speech text and phonetic symbols, it generates language learning explanation text, which includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols. Therefore, by generating language learning explanation text that includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols, this application enables users to fully understand the audio data and simultaneously complete language learning quickly and accurately, improving learning efficiency and achieving good learning results. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0044] Figure 1 One of the flowcharts illustrating the text generation method provided in this application;

[0045] Figure 2 The second flowchart illustrating the text generation method provided in this application;

[0046] Figure 3 A schematic diagram of the structure of the speech recognition model provided in this application;

[0047] Figure 4 The third flowchart illustrating the text generation method provided in this application;

[0048] Figure 5 A schematic diagram of the text generation device provided in this application;

[0049] Figure 6 A schematic diagram of the interface of the text generation system provided in this application;

[0050] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] In order to clearly illustrate this application, the following further elaborates on the problems of speech recognition in the prior art.

[0053] Existing speech evaluation systems mostly focus on the accuracy of text recognition, rarely considering the recognition of phonetic symbols other than text. Because current technology does not consider phonetic symbols in text, it reduces the efficiency and accuracy for users learning languages.

[0054] Furthermore, existing technologies for recognizing text and phonetic symbols in speech mainly employ two approaches: one is a cascaded architecture combining a Hidden Markov Model (HMM), a Gaussian Mixture Model (GMM), and a language model; the other is an end-to-end recognition architecture based on an encoder-decoder (ED) model, incorporating neural networks (RNN / Transformer / Conformer) to identify text or phonetic symbols. However, these approaches do not consider the contribution of phonetic symbols (other than text) to the model, resulting in lower accuracy for both text and phonetic symbols.

[0055] Exemplary Implementation Environment

[0056] Taking the learning of less commonly taught languages ​​as an example, it is very difficult for learners to quickly master the pronunciation skills of a completely new language. Therefore, it is very important to find the fastest and best way to learn.

[0057] From a learning environment perspective, due to the learner's actual surroundings, prolonged listening to audio or watching videos makes it difficult to guarantee learning effectiveness. For example, learning a less commonly taught language via audio or video on a subway or airplane is significantly hampered. From a time perspective, learning a less commonly taught language through audio and video is also very time-consuming; in the same amount of time, pure text-based learning can acquire more knowledge. From an effectiveness perspective, listening to audio and watching videos makes it difficult to obtain phonetic and textual information for the current language, and it is also very unfriendly to learners' rapid comprehension. For example, if a learner uses a lecture video as learning material, converting the video into corresponding pronunciation and text, and then using a large model to convert it into easily understandable sentences, would be much more user-friendly.

[0058] Therefore, this application proposes a text generation method that combines phonetic recognition with a large language model to assist language learning. Combining phonetic recognition with a large language model to assist language learning can help language learners quickly master the pronunciation skills of the target language, thereby completing the learning of the target language.

[0059] Specifically, when inputting audio or video (such as a speech from the presidential election in Country A), this application will first identify the text content in the audio or video, and at the same time identify the phonetic symbols corresponding to the text. Then, combining relevant key information and a large language model, it will output a language learning explanation text that is easy for learners to understand and master, so as to help learners quickly learn the standard pronunciation of the target language and quickly improve their speech ability.

[0060] Exemplary methods

[0061] This application provides a text generation method that can be applied to smart terminals, such as mobile phones, computers, and tablets, as well as servers. The following description uses the application of this method to a server as an example; however, it should be noted that this is merely illustrative and not intended to limit the scope of protection of this application. Other descriptions in the embodiments of this application are also illustrative and not intended to limit the scope of protection of this application, and will not be described in detail thereafter.

[0062] Please see Figure 1 In one exemplary embodiment, a text generation method is provided.

[0063] Step 101: Obtain audio data containing speech content in a specific language and extract the audio features of the audio data.

[0064] The audio data can be standalone audio or audio extracted from video data.

[0065] Specifically, the data to be identified is obtained; when the data to be identified is only audio data, the audio features of the audio data can be extracted directly; when the data to be identified is video data, the audio data in the video data is extracted, and then the audio features of the audio data are extracted.

[0066] The data to be identified includes speech content in a specific language.

[0067] Specifically, the necessary steps of speech recognition are generally to extract audio features from audio data. By extracting audio features, relevant information that is beneficial to recognition is retained, while relevant information that is detrimental to or has no impact on the recognition results is discarded, such as background noise.

[0068] Specifically, the specific implementation of extracting audio features from audio data is as follows:

[0069] Using video editing software (moviepy) and an audio processing library (librosa), waveforms of the audio data were extracted, with a sampling rate of 16000. Then, a frame-segmentation and windowing strategy was employed to obtain the audio features corresponding to the waveforms.

[0070] Common audio feature extraction algorithms include Mel frequency cepstral coefficients (MFCC), filter bank-based acoustic feature extraction algorithm (Fbank), and speech acoustic feature extraction algorithm (LogFBank).

[0071] The following explanation uses MFCC as an example to illustrate the audio feature extraction process:

[0072] A waveform is divided into multiple segments of fixed length using a specific method; this is known as frame segmentation. Specifically, given the rapid changes in speech signals, the duration of each frame is typically 10-30 milliseconds to ensure sufficient periods within a frame and prevent overly drastic changes. This makes it more suitable for Fourier transforms, which are applicable to analyzing stationary signals. Due to varying sampling rates in digital audio, the dimension of each frame vector obtained from frame segmentation also differs.

[0073] To avoid information loss due to time window boundaries, a certain overlap is needed between frames when offsetting the time window for each frame extracted from audio features. This time window offset is typically half the frame length, meaning that each step offsets the window by approximately half the frame's length, which serves as the final position for the next frame. This approach avoids excessive characteristic changes between frames.

[0074] For example, select a time window length of 25 milliseconds and a time window offset of 10 milliseconds.

[0075] In the previous framing process, directly dividing a continuous speech signal into several segments would cause spectral leakage due to truncation. The purpose of windowing is to eliminate the signal discontinuity problem that occurs at the edges of the short-time signal of each frame. Hamming windows can be used for windowing.

[0076] After the above series of processing steps, the audio data is obtained in the time domain, but the amount of speech information directly obtainable in the time domain is relatively small. Further audio feature extraction requires converting the time domain signal of each frame to its corresponding frequency domain signal. For example, this can be done through Fourier transform to obtain the frequency domain signal.

[0077] After Fourier transform, the resulting speech features are a complex matrix, which is an energy spectrum. Since the phase spectrum in the energy spectrum contains less information, the phase spectrum can be discarded and the amplitude spectrum retained.

[0078] The energy spectrum after preserving the amplitude spectrum is filtered using a filter to obtain audio features.

[0079] Since the extraction of audio features using a frame-segmentation and windowing strategy is an existing technique, this application will not provide further details.

[0080] Step 102: Based on audio features, identify the speech text corresponding to the speech content in a specific language and the phonetic symbols corresponding to the speech text.

[0081] Specifically, a speech recognition library is pre-created, a large amount of user audio data is collected, and the corresponding speech text and phonetic symbols are obtained. The audio features, speech text, phonetic symbols, and the first correspondence between the audio features and the speech text and phonetic symbols corresponding to the collected audio data are stored in the speech recognition library. Then, the obtained audio features are matched in the speech recognition library, and based on the first correspondence, the speech text corresponding to the audio features and the phonetic symbols corresponding to the speech text are obtained.

[0082] In some embodiments, a speech recognition model can be pre-trained to recognize the text and phonetic symbols corresponding to the speech. Inputting the aforementioned audio features into the speech recognition model yields the output speech text and the corresponding phonetic symbols.

[0083] Step 103: Generate language learning explanation text based on the speech text and phonetic symbols.

[0084] The language learning explanation texts include audio texts, phonetic symbols, and prompts for learning the audio texts and phonetic symbols.

[0085] The language learning explanation text includes each word of the audio text, as well as the phonetic symbols corresponding to each word.

[0086] Specifically, a language learning explanation text generation library is pre-created, and a large amount of speech text, phonetic symbols, and language learning explanation text are collected. A second correspondence relationship is established between the speech text, phonetic symbols, and language learning explanation text. The speech text, phonetic symbols, language learning explanation text, and the second correspondence relationship are stored in the language learning explanation text generation library. Then, the obtained speech text and phonetic symbols are matched in the language learning explanation text generation library, and based on the second correspondence relationship, the language learning explanation text corresponding to the speech text and phonetic symbols is obtained.

[0087] For example, a language learning explanation text generation model can be pre-trained to generate language learning explanation text. This language learning explanation text generation model is trained based on speech text and phonetic symbols. By inputting the speech text and phonetic symbols obtained in step S102 into the above-mentioned language learning explanation text generation model, a language learning explanation text containing the speech text, phonetic symbols, and prompts containing learning speech text and phonetic symbols can be obtained.

[0088] This application provides a text generation method. It acquires audio data containing speech content in a specific language and extracts audio features from the audio data. Based on these audio features, it identifies the speech text corresponding to the speech content in the specific language and the corresponding phonetic symbols. Thus, this application assists users in correctly pronouncing speech by extracting the speech text and corresponding phonetic symbols from the audio data. Furthermore, based on the speech text and phonetic symbols, it generates language learning explanation text, which includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols. Therefore, by generating language learning explanation text that includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols, this application enables users to fully understand the audio data and simultaneously complete language learning quickly and accurately, improving learning efficiency and achieving good learning results.

[0089] In one specific embodiment, the specific implementation of recognizing speech text and the phonetic symbols corresponding to the speech text includes: inputting audio features into a speech recognition model to obtain speech text corresponding to the speech content of a specific language output by the speech recognition model and the phonetic symbols aligned with the speech text.

[0090] The speech recognition model is obtained through training by aligning speech text with phonetic symbols and training by recognizing speech text and phonetic symbols.

[0091] This application treats the speech recognition model as a black box, inputting the obtained audio features into it, and can quickly and accurately obtain the corresponding speech text and the phonetic symbols aligned with the speech text. By aligning the speech text and phonetic symbols, the synchronization between the speech text and pronunciation is ensured, providing a data foundation for the subsequent output of aligned speech text and phonetic symbols.

[0092] This application aligns audio features with speech text and phonetic symbols, then uses the aligned speech text and phonetic symbols to train a speech recognition model, optimizing the model's computational parameters to ensure the accuracy of the output (speech text and phonetic symbols aligned with the speech text). Furthermore, this application leverages the positive assistance of phonetic symbols in the speech recognition training process, improving the accuracy of the model's output. Moreover, the final output of the speech recognition model, consisting of speech text and phonetic symbols aligned with the speech text, is more convenient for learners and enhances the user experience.

[0093] In one specific embodiment, a speech recognition model is pre-trained using a large amount of sample data (sample dataset). The training process of the speech recognition model includes:

[0094] Acquire audio feature samples, text samples, and phonetic symbol samples; input the audio feature samples and phonetic symbol samples into the speech recognition model so that the speech recognition model processes the audio feature samples and phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol samples, and generate predicted text and predicted phonetic symbols based on the audio feature samples, aligned audio feature samples, and phonetic symbol samples; compare the text samples and predicted text, and compare the phonetic symbol samples and predicted phonetic symbols to obtain the comparison results; adjust the operation parameters of the speech recognition model based on the comparison results.

[0095] Specifically, a sample dataset is created in advance, which includes audio feature samples, text samples, and phonetic symbol samples. This sample dataset can be downloaded from the internet or an existing database, or it can be created independently.

[0096] Specifically, annotations are added to the existing audio feature samples and phonetic symbols in the sample dataset to obtain annotated audio feature samples and phonetic symbols. Furthermore, this application does not necessarily require annotation of every frame of the audio feature samples and phonetic symbols; annotation can be performed based on word segmentation of the text sample and the pronunciation of each word. This annotation method ensures both the accuracy of the model training results and improves training efficiency.

[0097] Furthermore, in order to obtain more sample data and improve the output accuracy of the speech recognition model, noise is added to the labeled audio feature samples and phonetic symbols to obtain more sample data for training the speech recognition model.

[0098] This application trains the speech recognition model with a large amount of sample data, which effectively improves the accuracy of the speech recognition model's output results.

[0099] In one specific embodiment, the speech recognition model processes audio feature samples and phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol feature samples as follows:

[0100] The phonetic symbols are encoded to obtain phonetic symbol feature samples; the product of the audio feature samples and the transpose of the phonetic symbol feature samples is calculated to obtain the first fusion result; the product of the first fusion result and the phonetic symbol feature samples is calculated to obtain the phonetic symbol feature samples aligned with the audio feature samples.

[0101] Specifically, the phonetic symbol sample is represented by the letter P, and a dimensionality reduction operation (Embedding) is performed on P to obtain the phonetic symbol feature sample P. E Through AV F This represents an audio feature sample.

[0102] Specifically, before aligning the audio feature samples and the phonetic symbol feature samples, it is necessary to convert the audio feature samples and the phonetic symbol feature samples to the same dimension.

[0103] For example, AV F The dimensions are [B, T1, 80], P E The dimensions are [B, T2, 80].

[0104] Where B is the number of input samples each time when training the speech recognition model, T1 is the frame length corresponding to the speech feature sample, T2 is the frame length corresponding to the phonetic symbol feature sample, and 80 is a constant, i.e. 80 dimensions. Of course, users can set this dimension according to the actual situation. This is only an example and is not intended to limit this application.

[0105] Specifically, regarding B, let me clarify again: when training the speech recognition model, multiple training sessions will be conducted using sample data. Each input sample data session can consist of one set or multiple sets, which the user can choose according to their specific needs. One set of sample data includes one audio feature sample, one text sample, and one phonetic symbol sample.

[0106] Specifically, the audio feature samples and phonetic symbol feature samples are mixed to better integrate them and achieve alignment.

[0107] Specifically, the product of the audio feature sample and the transpose of the phonetic symbol feature sample is calculated to obtain the first fusion result, as shown in formula (1):

[0108]

[0109] Among them, AB E This indicates the first fusion result.

[0110] That is, for AV F (dimensions [B,T1,80]) and P E The transpose of (dimension [B, T2, 80]) is obtained by matrix multiplication to get AB. E (Dimensions are [B,T1,T2]).

[0111] Specifically, the product of the first fusion result and the phonetic feature sample is calculated to obtain the phonetic feature sample aligned with the audio feature sample, as shown in formula (2):

[0112] P F =AB E *P E ……………………………………………(2)

[0113] Among them, P F These are phonetic symbol feature samples aligned with the audio feature samples.

[0114] That is, for AB E (dimensions [B,T1,T2]) and P E (Dimensions [B, T2, 80]) Matrix multiplication is used to obtain the phonetic symbol feature samples P aligned with the audio feature sample frame level. F (Dimensions are [B,T1,80]).

[0115] Specifically, the alignment process of the aforementioned audio feature samples and phonetic symbol feature samples can be found in [reference needed]. Figure 2 .

[0116] This application ensures synchronization between speech text and pronunciation by aligning audio feature samples and phonetic symbol feature samples in the latent space, providing a data foundation for subsequent output of aligned speech text and phonetic symbols. Furthermore, the alignment operation effectively fuses the features of both in the latent space, improving the performance of the speech recognition model in recognizing speech text and phonetic symbols.

[0117] In one specific embodiment, the speech recognition model generates predicted text and predicted phonetic symbols based on the audio feature samples, the aligned audio feature samples, and the phonetic symbol feature samples. Furthermore, the text samples and the predicted text are compared, and the phonetic symbol samples and the predicted phonetic symbols are compared to obtain a comparison result. The specific implementation process of adjusting the operational parameters of the speech recognition model based on the comparison result is as follows:

[0118] The speech recognition model generates predicted text and a first predicted phonetic symbol based on audio feature samples; and generates a second predicted phonetic symbol based on the aligned audio feature samples and phonetic symbol feature samples.

[0119] The first comparison result is obtained by comparing the text sample and the predicted text; the second comparison result is obtained by comparing the phonetic symbol sample and the first predicted phonetic symbol; the third comparison result is obtained by comparing the phonetic symbol sample and the second predicted phonetic symbol; the operation parameters of the speech recognition model are adjusted based on the first comparison result, the second comparison result and the third comparison result.

[0120] Specifically, the following section will first introduce the structure of a speech recognition model:

[0121] The network layer in the speech recognition model that implements the above-mentioned aligned audio feature samples and phonetic symbol feature samples is named the alignment network.

[0122] The speech recognition model also includes DenseNet, Conformer, Long Short-Term Memory (LSTM), and Multilayer Perceptron (MLP) networks.

[0123] DenseNet uses aligned audio feature samples, phonetic symbol feature samples, and audio feature samples as input to form a dense connection structure, thereby increasing feature reuse and gradient flow, which improves the performance and generalization ability of the speech recognition model.

[0124] The following illustration uses a 16-layer Conformer convolutional network:

[0125] The Conformer consists of four parts: a feedforward module (composed of two non-linear activation functions), a multi-head self-attention module, a convolution module, and a text output module. The outputs of both feedforward modules are multiplied by one-half.

[0126] Specifically, the aligned audio feature samples and phonetic symbol feature samples are processed by the Feed Forward Module, Multi-head self attention Module, and Convolution Module to obtain the intermediate features of the speech text and the intermediate features of the phonetic symbols corresponding to the speech data.

[0127] In order to enhance the expressive power of the model, this application utilizes LSTM to learn intermediate features of speech text and intermediate features of phonetic symbols.

[0128] In this model, MLP is used to connect all neurons between layers. Therefore, it can also be defined as a fully connected layer to improve the performance and generalization ability of the speech recognition model.

[0129] Specifically, the structure diagram of this speech recognition model can be found in [reference needed]. Figure 3 In order to clearly illustrate the structure of the speech recognition model, in Figure 3 In the diagram, the convolutional network is illustrated using a dashed box.

[0130] In addition, during the training of the speech recognition model, the processing through multiple layers of networks will generate losses, that is, each network will generate a certain loss. Therefore, it is necessary to calculate the loss function of each network.

[0131] Specifically, based on the first comparison result, the first consistency rate between the text sample and the predicted text is determined; based on the second comparison result, the second consistency rate between the phonetic symbol sample and the second predicted phonetic symbol is determined; based on the third comparison result, the third consistency rate between the phonetic symbol sample and the second predicted phonetic symbol is determined; based on the first consistency rate, the second consistency rate, the third consistency rate and the loss function, the step of adjusting the operation parameters is repeatedly executed until the first consistency rate, the second consistency rate and the third consistency rate all reach their corresponding values, then the speech training model is determined to be trained successfully.

[0132] Specifically, a loss function can be assigned to each of the first, second, and third consistency rates. This loss function is used to adjust the computational parameters of the speech recognition model, allowing it to better fit the training process and improve its performance. Alternatively, one loss function can be assigned to each of the first, second, and third consistency rates; these three loss functions can be used to adjust the computational parameters of the speech recognition model, further improving its performance. This application does not restrict the specific settings of the loss functions; users can configure them according to their actual needs.

[0133] Specifically, the speech training model in this application has two training branches. The first training branch takes audio feature samples and phonetic symbol feature samples as input parameters and outputs a second predicted phonetic symbol. The second training branch takes audio feature samples as input parameters and outputs predicted text and a first predicted phonetic symbol. The two branches do not affect each other; they only provide positive assistance to the speech recognition model through phonetic symbol samples, making the final output of the speech recognition model more accurate. Specifically, when the first training branch takes audio feature samples and phonetic symbol feature samples as input parameters, it needs to perform an alignment operation between the semantic text and the phonetic symbols, ensuring synchronization between the speech text and pronunciation, and providing a valid data foundation for the subsequent output of aligned speech text and phonetic symbols.

[0134] In one specific embodiment, in order to generate language learning explanation texts that are easy for learners to learn and enable learners to quickly master the language they are learning, this application achieves this by indicating key information on the structural and / or content requirements for generating language learning explanation texts.

[0135] Specifically, key information is created; based on the key information, speech text, and phonetic symbols, task prompts are constructed; the task prompts are input into a pre-trained large language model so that the large language model can generate language learning explanation text.

[0136] Among them, the task prompting instruction is used to prompt the large language model to generate language learning explanation text containing speech text and phonetic symbols according to key information.

[0137] The task instruction contains key information, voice text, and phonetic symbols.

[0138] For example, in a scenario where a less commonly taught language is being taught, if a piece of audio data is input, the key information is as follows: Based on the obtained phonetic symbols, tell me the easiest way to learn to read the audio data aloud.

[0139] Among them, large language models refer to deep learning models trained with a large amount of text data, which can generate natural language text or understand the meaning of language text.

[0140] This application only requires optimizing the large language model based on key information, speech text, and phonetic symbols before generating the text. By utilizing key information, this application enables the large language model to better understand the input speech text and phonetic symbols, thereby generating language learning explanation text from the speech text and phonetic symbols. By using key information and the large language model to output speech text and phonetic symbols, this application makes the final language learning explanation text more aligned with learners' learning requirements, improving the user experience.

[0141] In one specific embodiment, the large language model learning prompts specifically include prompts for learning the pronunciation of the entire sentence of the speech text by practicing the pronunciation of each word segment of the speech text.

[0142] This application ensures the accuracy of language learning by segmenting the text into words and learning the pronunciation of the segmented words.

[0143] The specific examples of language learning explanation texts can be seen in the following illustrations. The audio data "Carbohydrates are a necessary food for humans" is used as an example for illustration.

[0144] First, divide this passage into several smaller segments, each containing one or more syllables. Then, practice the pronunciation of each segment one by one, paying attention to the initial consonant and tone of each syllable. You can use pinyin, Chinese characters, or other methods to mark the pronunciation of each syllable.

[0145] The following are the paragraph breaks:

[0146] 1. Carbohydrate (tan4 shuei3)

[0147] 2 is (shi4)

[0148] 3. Human beings (ren2lei4)

[0149] 4. Must (bi4xv1)

[0150] 5. Supplement (bu3chong1)

[0151] 6. of (de1)

[0152] 7. Food (shi2_uu4)

[0153] By practicing the pronunciation of each small segment, you can gradually combine them to form the pronunciation of the whole passage.

[0154] During practice, you can pay attention to the following:

[0155] 1. Ensure that each syllable is pronounced accurately.

[0156] 2. Pay attention to the changes in initial consonants, such as "i" and "ei", "u" and "uei", etc.

[0157] 3. Pay attention to the changes in tone, such as the first, second, fourth, sixth, and eighth tones.

[0158] 4. When reading aloud, pay attention to the pace and pauses to make the sentences flow more smoothly and naturally.

[0159] Below, through Figure 4 The text generation method of this application is illustrated by example:

[0160] Step 401: Obtain audio data containing speech content in a specific language and extract audio features from the audio data.

[0161] Step 402: Input the audio features into the speech recognition model to obtain the speech text corresponding to the audio data and the phonetic symbols corresponding to the speech text output by the speech recognition model.

[0162] Step 403: Create key information and generate task prompts based on the key information, voice text, and phonetic symbols.

[0163] Step 404: Input the task prompt instruction into the pre-trained large language model so that the large language model can generate language learning explanation text.

[0164] This application utilizes audio and video understanding technology, speech recognition technology, and text generation technology to identify the corresponding text and phonetic symbols based on the currently learned audio or video files. Combined with the powerful text generation capabilities of a large language model, it generates language learning explanation texts that are easy for learners to understand and master, helping learners quickly learn the standard pronunciation of the target language and rapidly improve their speech skills.

[0165] In addition, this application converts audio or video files into language learning explanation texts that are easy for learners to understand and master, which is very user-friendly and improves the user experience.

[0166] Exemplary device

[0167] Accordingly, please refer to Figure 5 This application also provides a text generation apparatus.

[0168] The feature extraction module 501 is used to acquire audio data containing speech content in a specific language and extract audio features from the audio data.

[0169] The speech recognition module 502 is used to recognize the speech text corresponding to the speech content of a specific language and the phonetic symbols corresponding to the speech text based on audio features;

[0170] The text generation module 503 is used to generate language learning explanation text based on speech text and phonetic symbols. The language learning explanation text includes speech text, phonetic symbols, and prompts for learning speech text and phonetic symbols.

[0171] This application provides a text generation device. A feature extraction module 501 acquires audio data containing speech content in a specific language and extracts audio features from the audio data. A speech recognition module 502, based on the audio features, identifies the speech text corresponding to the speech content in the specific language and the corresponding phonetic symbols. Thus, this application assists users in correctly pronouncing speech by extracting the speech text and corresponding phonetic symbols from the audio data. Furthermore, a text generation module 503 generates language learning explanation text based on the speech text and phonetic symbols. The language learning explanation text includes the speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols. Therefore, this application, by generating language learning explanation text including speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols, enables users to fully understand the audio data and simultaneously complete language learning quickly and accurately, improving learning efficiency and achieving good learning results.

[0172] In one specific embodiment, the speech recognition module 502 is specifically used to input audio features into the speech recognition model to obtain the speech text corresponding to the speech content of a specific language output by the speech recognition model and the phonetic symbols aligned with the speech text; wherein, the speech recognition model is obtained through speech text and phonetic symbol alignment training and speech text and phonetic symbol recognition training.

[0173] This application treats the speech recognition model as a black box, inputting the obtained audio features into it, and can quickly and accurately obtain the corresponding speech text and the phonetic symbols aligned with the speech text. By aligning the speech text and phonetic symbols, the synchronization between the speech text and pronunciation is ensured, providing a data foundation for the subsequent output of aligned speech text and phonetic symbols.

[0174] This application aligns audio features with speech text and phonetic symbols, then uses the aligned speech text and phonetic symbols to train a speech recognition model, optimizing the model's computational parameters to ensure the accuracy of the output (speech text and phonetic symbols aligned with the speech text). Furthermore, this application leverages the positive assistance of phonetic symbols in the speech recognition training process, improving the accuracy of the model's output. Moreover, the final output of the speech recognition model, consisting of speech text and phonetic symbols aligned with the speech text, is more convenient for learners and enhances the user experience.

[0175] In one specific embodiment, the device further includes a model training module for training a speech recognition model, specifically for acquiring audio feature samples, text samples, and phonetic symbol samples; inputting the audio feature samples and phonetic symbol samples into the speech recognition model so that the speech recognition model processes the audio feature samples and phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol sample samples, and generating predicted text and predicted phonetic symbols based on the audio feature samples, aligned audio feature samples, and phonetic symbol sample samples; comparing the text samples and predicted text, and comparing the phonetic symbol samples and predicted phonetic symbols to obtain comparison results; and adjusting the operation parameters of the speech recognition model based on the comparison results.

[0176] This application trains the speech recognition model with a large amount of sample data, which effectively improves the accuracy of the speech recognition model's output results.

[0177] In one specific embodiment, the model training module is specifically used to encode the phonetic symbol samples to obtain phonetic symbol feature samples; calculate the product of the audio feature samples and the transpose of the phonetic symbol feature samples to obtain a first fusion result; and calculate the product of the first fusion result and the phonetic symbol feature samples to obtain phonetic symbol feature samples aligned with the audio feature samples.

[0178] This application ensures synchronization between speech text and pronunciation by aligning audio feature samples and phonetic symbol feature samples in the latent space, providing a data foundation for subsequent output of aligned speech text and phonetic symbols. Furthermore, the alignment operation effectively fuses the features of both in the latent space, improving the performance of the speech recognition model in recognizing speech text and phonetic symbols.

[0179] In one specific embodiment, the model training module is specifically used to generate predicted text and a first predicted phonetic symbol corresponding to the audio feature sample based on the audio feature sample; generate a second predicted phonetic symbol based on the aligned audio feature sample and phonetic symbol feature sample; compare the text sample and the predicted text to obtain a first comparison result; compare the phonetic symbol sample and the first predicted phonetic symbol to obtain a second comparison result; and adjust the operation parameters of the speech recognition model based on the first comparison result, the second comparison result, and the third comparison result.

[0180] The speech training model in this application has two training branches. The first training branch takes audio feature samples and phonetic symbol feature samples as input parameters and outputs a second predicted phonetic symbol. The second training branch takes audio feature samples as input parameters and outputs predicted text and a first predicted phonetic symbol. The two branches do not affect each other; they only provide positive assistance to the speech recognition model through phonetic symbol samples, making the final output of the speech recognition model more accurate. Specifically, when the first training branch takes audio feature samples and phonetic symbol feature samples as input parameters, it needs to perform an alignment operation between semantic text and phonetic symbols to ensure synchronization between speech text and pronunciation, providing an effective data foundation for the subsequent output of aligned speech text and phonetic symbols.

[0181] In one specific embodiment, the text generation module 503 is specifically used to create key information, wherein the key information is used to indicate the structural requirements and / or content requirements for generating language learning explanation text; based on the key information, speech text, and phonetic symbols, a task prompt instruction is constructed; the task prompt instruction is used to prompt the large language model to generate language learning explanation text containing speech text and phonetic symbols according to the key information; the task prompt instruction is input into the pre-trained large language model so that the large language model generates language learning explanation text.

[0182] This application only requires optimizing the large language model based on key information, speech text, and phonetic symbols before generating the text. By utilizing key information, this application enables the large language model to better understand the input speech text and phonetic symbols, thereby generating language learning explanation text from the speech text and phonetic symbols. By using key information and the large language model to output speech text and phonetic symbols, this application makes the final language learning explanation text more aligned with learners' learning requirements, improving the user experience.

[0183] In one specific embodiment, the language learning explanation text includes each segment of the spoken text and the phonetic symbol corresponding to each segment. The text generation module 503 is specifically used to learn the pronunciation prompts of the entire sentence of the spoken text by practicing the pronunciation of each segment of the spoken text.

[0184] This application utilizes audio and video understanding technology, speech recognition technology, and text generation technology to identify the corresponding text and phonetic symbols based on the currently learned audio or video files. Combined with the powerful text generation capabilities of a large language model, it generates language learning explanation texts that are easy for learners to understand and master, helping learners quickly learn the standard pronunciation of the target language and rapidly improve their speech skills.

[0185] In addition, this application converts audio or video files into language learning explanation texts that are easy for learners to understand and master, which is very user-friendly and improves the user experience.

[0186] The text generation apparatus provided in this embodiment belongs to the same concept as the text generation method provided in the above embodiments of this application. It can execute the text generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the text generation method. Technical details not described in detail in this embodiment can be found in the specific processing content of the text generation method provided in the above embodiments of this application, and will not be repeated here.

[0187] Exemplary System

[0188] Optionally, embodiments of this application also provide a text generation system, which includes a system interface and a speech recognition architecture, wherein the speech recognition architecture is deployed on a service provider running the text generation system.

[0189] The system interface includes a "+Upload Video / Audio" button, allowing learners to upload audio or video files. After receiving the audio or video files, the system extracts the audio data and executes the steps of the text generation method described above, outputting language learning explanation text for learners to study. See the attached screenshot for a detailed interface diagram. Figure 6 , Figure 6 The audio data "Carbohydrates are a necessary food for humans" is used as an example for illustration.

[0190] The system architecture includes a feature extraction module 501, used to acquire audio data containing speech content in a specific language and extract audio features from the audio data; a speech recognition module 502, used to identify the speech text corresponding to the speech content in a specific language and the corresponding phonetic symbols based on the audio features; a text generation module 503, used to generate language learning explanation text based on the speech text and phonetic symbols, the language learning explanation text including speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols; a key information construction module, used to create key information, wherein the key information is used to indicate the structural requirements and / or content requirements for generating the language learning explanation text; and a task prompt instruction based on the key information, speech text, and phonetic symbols, the task prompt instruction being used to prompt the large language model to generate language learning explanation text containing speech text and phonetic symbols according to the key information.

[0191] The text generation system provided in this application allows users to quickly obtain language learning explanation texts by uploading the files they need to learn using the "+Upload Video / Audio" button. The entire operation is simple and fast, improving the user experience.

[0192] This application provides a text generation system that acquires audio data containing speech content in a specific language through a system architecture and extracts audio features from the audio data. Based on the audio features, it identifies the speech text corresponding to the speech content in the specific language and the corresponding phonetic symbols. Thus, this application assists users in pronouncing correctly by extracting the speech text and corresponding phonetic symbols from the audio data. Furthermore, based on the speech text and phonetic symbols, it generates language learning explanation text, which includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols. Therefore, by generating language learning explanation text that includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols, this application enables users to fully understand the audio data and simultaneously complete language learning quickly and accurately, improving learning efficiency and achieving good learning results.

[0193] The text generation system provided in this embodiment belongs to the same concept as the text generation method provided in the above embodiments of this application. It can execute the text generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the text generation method. Technical details not described in detail in this embodiment can be found in the specific processing content of the text generation method provided in the above embodiments of this application, and will not be repeated here.

[0194] Exemplary electronic devices

[0195] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 7 As shown, the device includes:

[0196] Memory 200 and processor 210;

[0197] The memory 200 is connected to the processor 210 and is used to store programs;

[0198] The processor 210 is configured to implement the text generation method disclosed in any of the above embodiments by running a program stored in the memory 200. The method specifically includes: acquiring audio data containing speech content in a specific language and extracting audio features from the audio data; identifying the speech text corresponding to the speech content in the specific language and the phonetic symbols corresponding to the speech text based on the audio features; and generating language learning explanation text based on the speech text and phonetic symbols. The language learning explanation text includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols.

[0199] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0200] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:

[0201] A bus can include a pathway for transmitting information between various components of a computer system.

[0202] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0203] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.

[0204] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0205] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0206] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0207] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0208] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement the various steps of any of the text generation methods provided in the above embodiments of this application.

[0209] Specifically, an electronic device capable of executing any of the above text generation methods can assist users in pronouncing words correctly by extracting the speech text and corresponding phonetic symbols from audio data; and by generating language learning explanation text that includes speech text, phonetic symbols, and prompts for learning speech text and phonetic symbols, users can fully understand the audio data and simultaneously complete language learning quickly and accurately, thereby improving learning efficiency and achieving good learning results.

[0210] Exemplary computer program products and storage media

[0211] In addition to the methods and devices described above, embodiments of this application may also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the text generation methods according to various embodiments of this application described in the "Exemplary Methods" section of this specification. Specifically, the method includes: acquiring audio data containing speech content in a specific language and extracting audio features from the audio data; identifying, based on the audio features, speech text corresponding to the speech content in the specific language and phonetic symbols corresponding to the speech text; and generating language learning explanation text based on the speech text and phonetic symbols, the language learning explanation text including speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols.

[0212] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0213] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor in the steps of the text generation method according to various embodiments of this application described in the "Exemplary Methods" section above. This method specifically includes: acquiring audio data containing speech content in a specific language and extracting audio features from the audio data; identifying the speech text corresponding to the speech content in the specific language and the phonetic symbols corresponding to the speech text based on the audio features; and generating language learning explanation text based on the speech text and phonetic symbols, wherein the language learning explanation text includes the speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols.

[0214] Specifically, products and storage media capable of executing any of the above text generation methods can assist users in pronouncing audio data correctly by extracting the speech text and corresponding phonetic symbols from the audio data; and by generating language learning explanation text that includes speech text, phonetic symbols, and prompts for learning the speech text and phonetic symbols, users can fully understand the audio data and simultaneously complete language learning quickly and accurately, thereby improving learning efficiency and achieving good learning results.

[0215] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0216] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0217] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0218] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0219] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0220] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0221] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0222] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0223] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0224] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0225] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text generation method, characterized in that, include: Acquire audio data containing speech content in a specific language, and extract the audio features of the audio data; Based on the audio features, identify the speech text corresponding to the speech content of the specific language and the phonetic symbols corresponding to the speech text; Based on the spoken text and the phonetic symbols, a language learning explanation text is generated. The language learning explanation text includes each word of the spoken text, the phonetic symbols corresponding to each word, and prompts to learn the pronunciation of the whole sentence of the spoken text by practicing the pronunciation of each word.

2. The text generation method according to claim 1, characterized in that, The step of identifying the speech text corresponding to the specific language speech content and the phonetic symbols corresponding to the speech text based on the audio features includes: The audio features are input into the speech recognition model to obtain the speech text corresponding to the speech content of the specific language and the phonetic symbols aligned with the speech text. The training process of the speech recognition model includes: Obtain audio feature samples, text samples, and phonetic symbol samples; The audio feature samples and the phonetic symbol samples are input into the speech recognition model, so that the speech recognition model processes the audio feature samples and the phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol sample, and generates predicted text and predicted phonetic symbols based on the audio feature samples, the aligned audio feature samples and the phonetic symbol sample; the text sample and the predicted text are compared, and the phonetic symbol sample and the predicted phonetic symbols are compared to obtain the comparison result; the operation parameters of the speech recognition model are adjusted based on the comparison result.

3. The text generation method according to claim 2, characterized in that, The process of processing the audio feature samples and the phonetic symbol samples to obtain aligned audio feature samples and phonetic symbol feature samples includes: The phonetic symbol samples are encoded to obtain phonetic symbol feature samples; Calculate the product of the audio feature sample and the transpose of the phonetic symbol feature sample to obtain the first fusion result; Calculate the product of the first fusion result and the phonetic feature sample to obtain a phonetic feature sample aligned with the audio feature sample.

4. The text generation method according to claim 2, characterized in that, The process of generating predicted text and predicted phonetic symbols based on the audio feature samples, the aligned audio feature samples, and the phonetic symbol feature samples includes: Based on the audio feature samples, generate the predicted text and the first predicted phonetic symbol corresponding to the audio feature samples; Based on the aligned audio feature samples and phonetic symbol feature samples, a second predicted phonetic symbol is generated; The comparison of the text sample and the predicted text, and the comparison of the phonetic symbol sample and the predicted phonetic symbol, to obtain the comparison results include: By comparing the text sample with the predicted text, a first comparison result is obtained; By comparing the phonetic symbol sample with the first predicted phonetic symbol, a second comparison result is obtained; By comparing the phonetic symbol sample with the second predicted phonetic symbol, a third comparison result is obtained; The step of adjusting the computational parameters of the speech recognition model based on the comparison results includes: adjusting the computational parameters based on the first comparison result, the second comparison result, and the third comparison result.

5. The text generation method according to any one of claims 1-4, characterized in that, The process of generating language learning explanation text based on the spoken text and the phonetic symbols includes: Create key information, wherein the key information is used to indicate the structural and / or content requirements for generating language learning explanation texts; Based on the key information, the speech text, and the phonetic symbols, a task prompt instruction is constructed; the task prompt instruction is used to prompt the large language model to generate language learning explanation text containing the speech text and the phonetic symbols according to the key information; The task prompts are input into a pre-trained large language model, which then generates language learning explanation text.

6. A text generation device, characterized in that, include: The feature extraction module is used to acquire audio data containing speech content in a specific language and extract the audio features of the audio data. The speech recognition module is used to identify the speech text corresponding to the speech content of the specific language and the phonetic symbols corresponding to the speech text based on the audio features; The text generation module is used to generate language learning explanation text based on the speech text and the phonetic symbols. The language learning explanation text includes each word of the speech text, the phonetic symbols corresponding to each word, and prompts to learn the pronunciation of the whole sentence of the speech text by practicing the pronunciation of each word.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the text generation method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text generation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for learning word pronunciation in continuous context

    CN101425230A

  • Method and device for obtaining information

    CN109410918A