Mixed speech pronunciation sequence generation method, model training method and related device
By predicting and analyzing the original audio signal, a high-quality mixed speech pronunciation sequence is generated, which solves the problem of poor multilingual mixed audio recognition effect in the prior art, and significantly improves the multilingual mixed audio recognition effect of the language recognition model.
Patent Information
- Application Number
- CN202510036970.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The existing automatic speech recognition model has poor recognition effect when processing multilingual mixed audio. This is mainly due to the small number of multilingual mixed audio samples, especially the lack of pronunciation of professional phrases, which makes it difficult to accurately identify during training.
By performing pronunciation prediction based on the original audio signal, a predicted pronunciation sequence is generated and the predicted pronunciation sequence is analyzed to obtain the actual pronunciation sequence of the target phrase. Then, the actual pronunciation sequence and the initial text are used to generate mixed pronunciation sequences to improve the quality of multilingual mixed audio pronunciation sequences.
By generating high-quality mixed speech pronunciation sequences, the recognition effect of the language recognition model in a multilingual mixed audio environment is improved, especially in the case of serious lack of relevant audio data, the recall and recognition effect of the target phrase can be quickly improved.
Smart Images

Figure CN120071893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio signal processing, in particular to a method for generating a mixed speech pronunciation sequence, a method for model training, and related devices. Background Art
[0002] Currently, the existing Automatic Speech Recognition (ASR) model has a good recognition effect when processing audio in a single language, but has a poor recognition effect for multi-language mixed audio. In the prior art, during the training process of the multi-language mixed audio recognition ability of the automatic speech recognition model, the multi-language mixed audio samples used are few. For example, when training an automatic speech recognition model for Chinese-English mixed audio of certain professional English phrases, the audio data containing these professional English phrases is scarce, and the pronunciation of these professional English phrases is lacking, resulting in difficult accurate recognition during the training process. Summary of the Invention
[0003] This application provides at least a method for generating a mixed speech pronunciation sequence, a method for model training, and related devices, which can improve the quality of the pronunciation sequence of multi-language mixed audio.
[0004] In a first aspect of this application, a method for generating a mixed speech pronunciation sequence is provided. The method includes: performing pronunciation prediction on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal, where the original audio signal is an audio signal corresponding to a first language and carries a target phrase corresponding to a second language, and the first language and the second language are different languages; analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target phrase; and generating an initial text containing the target phrase; using the actual pronunciation sequence of the target phrase and the initial text to generate a mixed speech pronunciation sequence corresponding to the initial text, where in the mixed speech pronunciation sequence, the target phrase is in the second language, and other text parts in the initial text except the target phrase are in a third language, and the second language and the third language are different languages.
[0005] Among them, performing pronunciation prediction on the original audio signal to obtain a predicted pronunciation sequence of the original audio signal includes: extracting features from the original audio signal to obtain a plurality of filter bank features; predicting the plurality of filter bank features to obtain a predicted pronunciation sequence of the original audio signal.
[0006] Among them, extracting features from the original audio signal to obtain filter bank features includes: extracting features from the original audio signal to obtain audio features in a preset format; performing pre-preprocessing on the audio features in the preset format to obtain filter bank features.
[0007] Among them, before predicting a plurality of filter bank features to obtain a predicted pronunciation sequence of the original audio signal, it includes: performing global correlation processing on each filter bank feature to obtain a plurality of filter-related features; predicting a plurality of filter bank features to obtain a predicted pronunciation sequence of the original audio signal, including: predicting a plurality of filter-related features to obtain a predicted pronunciation sequence.
[0008] Among them, analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target phrase includes: based on the pronunciation reference data of the first language, taking the pronunciation subsequence that does not belong to the first language in the predicted pronunciation sequence as the actual pronunciation sequence of the target phrase.
[0009] Among them, the target phrase in the initial text is in the second language, and other text parts are in the third language; and / or, generating a mixed speech pronunciation sequence corresponding to the initial text by using the actual pronunciation sequence of the target phrase and the initial text, including: generating other pronunciation sequences corresponding to other text parts in the third language; combining the actual pronunciation sequence and the other pronunciation sequences to obtain a mixed speech pronunciation sequence.
[0010] The second aspect of this application provides a model training method, including: using the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio; and using the text recognition module of the language recognition model to perform text conversion on the mixed speech pronunciation sequence to obtain the recognition text corresponding to the mixed speech pronunciation sequence, where the mixed speech pronunciation sequence is obtained by the method of any item in the above first aspect; adjusting the network parameters of the language recognition model by using the first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio, and the second difference between the recognition text and the reference text corresponding to the mixed speech pronunciation sequence.
[0011] Among them, using the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio includes: using the pronunciation prediction module to extract features from the training audio to obtain training audio features; recognizing the training audio features to obtain a sample predicted pronunciation sequence; and / or, using the text recognition module of the language recognition model to perform text conversion on the mixed speech pronunciation sequence to obtain the recognition text corresponding to the mixed speech pronunciation sequence, including: using the text recognition module to extract features from the mixed speech pronunciation sequence to obtain mixed pronunciation features; performing text recognition on the mixed pronunciation features to obtain the recognition text.
[0012] The third aspect of this application provides an electronic device, including a memory and a processor coupled to each other, and the processor is used to execute program instructions stored in the memory to implement the mixed speech pronunciation sequence generation method in the above first aspect, and / or implement the model training method in the above second aspect.
[0013] The fourth aspect of this application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the hybrid speech pronunciation sequence generation method in the first aspect above is implemented, and / or the model training method in the second aspect above is implemented.
[0014] In the above solution, first, pronunciation prediction is performed on the original audio signal to obtain a predicted pronunciation sequence corresponding to the original audio signal. Then, the predicted pronunciation sequence is analyzed to obtain the actual pronunciation sequence of the target phrase, and at the same time, an initial text containing the target phrase is generated. Then, using the actual pronunciation sequence of the target phrase and the initial text, a hybrid speech pronunciation sequence corresponding to the initial text is generated, so that the hybrid speech pronunciation sequence also contains the actual pronunciation sequence of the target phrase in the second language, thereby improving the quality of the hybrid pronunciation sequence corresponding to the initial text.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application. Description of the Drawings
[0016] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with this application and are used together with the specification to explain the technical solutions of this application.
[0017] Figure 1 It is a schematic flowchart of an embodiment of the hybrid speech pronunciation sequence generation method of this application;
[0018] Figure 2 It is a schematic flowchart of another embodiment of the hybrid speech pronunciation sequence generation method of this application;
[0019] Figure 3 It is a schematic flowchart of an embodiment of the model training method of this application;
[0020] Figure 4 It is a schematic framework diagram of an embodiment of the speech recognition model of this application;
[0021] Figure 5 It is a schematic framework diagram of an embodiment of the electronic device of this application;
[0022] Figure 6 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of this application. Detailed Embodiments
[0023] The following combines the drawings in the specification to detail the solutions of the embodiments of this application.
[0024] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand this application.
[0025] As used herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article means any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C may mean including any one or more elements selected from the set composed of A, B, and C.
[0026] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the method for generating a mixed voice pronunciation sequence in the drawings of the present application. Specifically, it may include the following steps:
[0027] Step S110: Perform pronunciation prediction based on the original audio signal to obtain the predicted pronunciation sequence of the original audio signal.
[0028] Among them, the original audio signal is an audio signal corresponding to the first language and carries a target phrase corresponding to the second language, and the first language and the second language are different languages. For example, the first language may be Chinese, and the second language may be English, French, Persian, etc. Therefore, the first language and the second language are not specifically limited.
[0029] The present application mainly relates to the fields of audio signal processing, natural language processing, and deep learning. It can quickly predict the pronunciation of some professional target phrases using a small amount of existing audio signals, and use the LLM (Large Language Model) to generate a large number of multilingual mixed texts containing these professional target phrases. Finally, the constructed multilingual mixed text and the training audio information are used together to assist in training the ASR model, so that even in the case of severe lack of relevant audio, the recall rate and recognition effect of the target phrase can be quickly improved.
[0030] In some embodiments, the first language is the main language, and the pronunciation dictionary corresponding to the first language is known. Therefore, when predicting the pronunciation of the original audio signal, the original audio signal can be first recognized to determine the segmentation points in the original audio signal, and the original audio signal can be divided using the segmentation points to obtain a number of sub-audio signals, where each sub-audio signal corresponds to a word or a phrase in the first language. After that, the similarity between the pronunciation dictionary corresponding to the first language and each sub-audio signal is calculated. After obtaining the multiple similarities corresponding to each sub-audio signal, the pronunciation sequence in the pronunciation dictionary with the highest similarity to the sub-audio signal is selected as the predicted sub-pronunciation sequence of the sub-audio signal. Then, according to the position order of the sub-audio signals in the original audio signal, the predicted sub-pronunciation sequences are combined to obtain the predicted pronunciation sequence of the original audio signal.
[0031] In some other embodiments, the audio features in the original audio signal can be recognized and analyzed to obtain the predicted pronunciation sequence of the original audio signal. Specifically, reference can be made to steps S111 to S112.
[0032] Step S111: Extract features from the original audio signal to obtain a number of filter bank features.
[0033] In some embodiments, the original audio signal can be first subjected to feature extraction to obtain audio features in a preset format. Then, preprocessing is performed on the audio features in the preset format to obtain filter bank features. Among them, the filter bank features can correspond to a phrase in the original audio signal; the audio features in the preset format can be one-dimensional WAV (Waveform Audio File Format, a lossless compressed audio file format) audio features, etc., which are not specifically limited here. These filter bank features can reflect the important characteristics of the original audio signal and provide a basis for subsequent pronunciation prediction.
[0034] For example, by performing preprocessing on one-dimensional WAV audio features, fb40 features, that is, filter bank features, can be obtained. Among them, "fb" is the abbreviation of fbank, fbank is the filter bank, and "40" represents the byte length of the filter bank features.
[0035] Step S112: Predict the number of filter bank features to obtain the predicted pronunciation sequence of the original audio signal.
[0036] In some embodiments, after obtaining a number of filter bank features, global correlation processing can be performed on each filter bank feature to obtain a number of filter correlation features, so as to strengthen the correlation between features and improve the accuracy of pronunciation prediction. Specifically, a number of filter bank features can be input into a pre-trained pronunciation prediction model, and the N Conformer modules in the pronunciation prediction model are used to perform global correlation processing on the number of filter bank features. Then, a number of filter correlation features are predicted to obtain a predicted pronunciation sequence. For example, the original audio signal is an audio signal of "extracting fbank features". After feature processing, a number of filter correlation features corresponding to "extracting fbank features" are predicted to obtain a predicted pronunciation sequence of "t i2q u3ei4fb an4k t e4zh eng1". In addition, methods such as deep learning can be used for prediction, and the trained model is used to process the filter correlation features to obtain the final predicted pronunciation sequence.
[0037] Step S120: Analyze the predicted pronunciation sequence to obtain the actual pronunciation sequence of the target phrase.
[0038] In some embodiments, according to the pronunciation reference data of the first language, the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language can be used as the actual pronunciation sequence of the target phrase. Specifically, the pronunciation reference data can be a pronunciation dictionary. Using the pronunciation dictionary of the first language, the predicted pronunciation sequence can be analyzed to determine the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language, and this pronunciation subsequence is used as the actual pronunciation sequence of the target phrase.
[0039] For example, the similarity between the pronunciation dictionary and the pronunciation subsequence in the predicted pronunciation sequence is calculated, and the dictionary pronunciation sequence with the highest similarity in the pronunciation dictionary is selected. Then, the highest dictionary pronunciation sequence is compared with the similarity threshold. If the similarity of the dictionary pronunciation sequence is greater than or equal to the similarity threshold, it indicates that the pronunciation subsequence corresponding to this dictionary pronunciation sequence belongs to the first language; if the similarity of the dictionary pronunciation sequence is less than the similarity threshold, it indicates that the pronunciation subsequence corresponding to this dictionary pronunciation sequence does not belong to the first language. Furthermore, the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language is used as the actual pronunciation sequence of the target phrase.
[0040] In some other embodiments, the correlation between pronunciation subsequences in the predicted pronunciation sequence can be calculated, and the pronunciation subsequence in the predicted pronunciation sequence that is least correlated with other pronunciation subsequences is used as the actual pronunciation sequence of the target phrase. Specifically, the features of each pronunciation subsequence in the predicted pronunciation sequence are extracted, the correlation is calculated based on the features, and the correlation coefficients between the features of one pronunciation subsequence and the features of other pronunciation subsequences are fused to obtain the final correlation coefficient of this pronunciation subsequence. If the final correlation coefficient is less than the median of the final correlation coefficients of all pronunciation subsequences, the pronunciation subsequence corresponding to this final correlation coefficient does not belong to the pronunciation subsequences of the first language. Therefore, this pronunciation subsequence is used as the actual pronunciation sequence of the target phrase.
[0041] It can be understood that methods for calculating correlation can include methods such as Pearson correlation and Spearman correlation, and specific limitations are not made here.
[0042] Step S130: Generate an initial text containing the target phrase.
[0043] In some embodiments, a large language model can be used to generate an initial text containing the target phrase. To generate a high-quality initial text, appropriate prompt words need to be created. To enhance the domain diversity of the generated initial text, different domains of text need to be specified in the prompt words, such as "finance", "technology", "sports", "history and humanities", and so on. The large language model adopts the mainstream multi-layer transformer structure and has strong context modeling capabilities. The generated sentences have good effects in terms of diversity, instruction following, semantic smoothness, etc.
[0044] For example, the prompt word is that the text contains the target phrase "negotiation strategy", and it is specified to generate text in the financial domain. Therefore, after receiving this prompt word, the large language model will generate a large number of initial texts similar to "In this negotiation, we need to adopt an effective negotiation strategy to achieve a satisfactory result for both parties."
[0045] In some other embodiments, first translate the target phrase in the second language to obtain the target phrase text corresponding to the first language. Search in the database according to the target phrase text, and use the text containing the target phrase text as the candidate text. Then replace the target phrase text in the candidate text with the target phrase in the second language to obtain the initial text.
[0046] In one implementation scenario, the above steps S120 and S130 can be executed in sequence. For example, step S120 is executed first, and then step S130; or step S130 is executed first, and then step S120. In another implementation scenario, the above steps S120 and S130 can also be executed simultaneously, which can be specifically set according to the actual application and will not be limited here.
[0047] Step S140: Generate a mixed speech pronunciation sequence corresponding to the initial text by using the actual pronunciation sequence of the target phrase and the initial text.
[0048] Among them, in the mixed speech pronunciation sequence, the target phrase is in the second language, and the other text parts in the initial text except the target phrase are in the third language. The second language and the third language are different languages. The third language can be the same as the first language or different from the first language.
[0049] For example, in a speech assistance system for international business negotiations, when it is necessary to synthesize a Chinese text containing specific professional terms (the target phrase is in English, such as "negotiation strategy"), the initial text can be generated through a preset text template or according to the real-time input theme information, such as "In this negotiation, we need to adopt an effective negotiation strategy to achieve a satisfactory result for both parties." Here, the target phrase "negotiation strategy" is in the second language English, and the other text parts in the initial text except the target phrase are in Chinese, that is, the third language.
[0050] In some embodiments, the pronunciation reference data of the third language can be used to generate other pronunciation sequences corresponding to the other text parts in the third language. Then, the actual pronunciation sequence and the other pronunciation sequences are combined to obtain the mixed speech pronunciation sequence.
[0051] Taking the speech synthesis method based on deep learning as an example, when the third language is Chinese and the second language is English, the Chinese text in the initial text can be first subjected to feature extraction to obtain audio features in a preset format, and then pre-preprocessing is performed to obtain filter bank features. Then, the filter bank features are predicted to obtain the pronunciation sequence of the Chinese part. For example, for the Chinese text "In this negotiation, we need to adopt an effective", a corresponding Chinese pronunciation sequence can be obtained after a series of processes.
[0052] After obtaining the actual pronunciation sequence of the target phrase (such as the pronunciation sequence of the English phrase "negotiation strategy") and the pronunciation sequence of the Chinese part, they can be combined to obtain a mixed speech pronunciation sequence. In the generation of the mixed speech pronunciation sequence, the actual pronunciation sequence of the English phrase and the pronunciation sequence of the Chinese part are combined according to certain rules, so that in the mixed speech pronunciation sequence, the target phrase is in the second language, and the other text parts in the original text except the target phrase are in the third language. For example, in the above example, the finally obtained mixed speech pronunciation sequence is the combination of the Chinese pronunciation sequence and the pronunciation sequence of the English phrase "negotiation strategy".
[0053] In a specific implementation scenario, the first language and the third language are both Chinese, the second language is English, the original audio signal is "extract fbank features", and the target phrase is "fbank". Please refer to Figure 2 , input the original audio signal "extract fbank features" into the pronunciation prediction model. After receiving "extract fbank features", the pronunciation prediction model extracts features from the original audio signal "extract fbank features" to obtain one-dimensional WAV audio features, and then performs preprocessing on the one-dimensional WAV audio features to obtain several fb40 features. Then, the obtained several fb40 features are sent to N Conformer modules in the pronunciation prediction model for global correlation processing to obtain several filtered correlation features. After that, the prediction layer in the pronunciation prediction model is used to predict the several filtered correlation features to obtain the predicted pronunciation sequence "ti2qu3ei4fb an4k t e4zh eng1" corresponding to the original audio signal. Finally, by reasoning the predicted pronunciation sequence using the known Chinese pronunciation reference data, the actual pronunciation sequence "ei4f b an4k" of the target phrase can be obtained, thus determining the pronunciation sequence of the target phrase.
[0054] At the same time, set the prompt words and specify different fields, so that the large language model can generate the original text containing the target phrase "fbank" according to the prompt words. Combine the generated original text with the known Chinese pronunciation reference data and the actual pronunciation sequence of the target phrase "fbank" to obtain the mixed speech pronunciation sequence corresponding to the original text. Then, use the mixed speech pronunciation sequence corresponding to this original text and the training audio to train the language recognition model.
[0055] This solution can utilize a small amount of existing audio signals to determine the actual pronunciation sequence of the target phrase corresponding to the second language, and combine the actual pronunciation sequence of the target phrase and the initial text for conversion to obtain the mixed speech pronunciation sequence corresponding to the initial text, thereby improving the quality of the generated mixed speech pronunciation sequence and laying a good foundation for the training of subsequent language recognition models to improve the training effect of the language recognition models.
[0056] For the training method of the language recognition model, reference can be made to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of the model training method of this application. Specifically, it may include the following steps:
[0057] Step S310: Use the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain the sample prediction pronunciation sequence of the training audio.
[0058] Please refer to Figure 4 , the language recognition model 400 includes a pronunciation prediction module 410 and a text recognition module 420. The pronunciation prediction module 410 is used to perform pronunciation prediction on the training audio to obtain the sample prediction pronunciation sequence of the training audio. The text recognition module 420 is used to perform text conversion on the mixed speech pronunciation sequence to obtain the recognition text corresponding to the mixed speech pronunciation sequence.
[0059] Therefore, after receiving the training audio, the language recognition model 400 uses the pronunciation prediction module 410 to extract features from the training audio to obtain the training audio features; and recognizes the training audio features to obtain the sample prediction pronunciation sequence.
[0060] For example, use the pronunciation prediction module 410 to extract features from the training audio to obtain the training audio features, then perform downsampling and frame length reduction processing on the training audio features through 2D convolution, and then perform global correlation processing in the 16-layer conformer module in the pronunciation prediction module 410 to obtain the sample prediction pronunciation sequence.
[0061] Step S320: Use the text recognition module of the language recognition model to perform text conversion on the mixed speech pronunciation sequence to obtain the recognition text corresponding to the mixed speech pronunciation sequence.
[0062] Among them, the mixed speech pronunciation sequence is obtained according to the above-mentioned mixed speech pronunciation sequence generation method.
[0063] In some embodiments, use the text recognition module 420 to extract features from the mixed speech pronunciation sequence to obtain the mixed pronunciation features; and perform text recognition on the mixed pronunciation features to obtain the recognition text.
[0064] For example, after the text recognition module 420 extracts features from the mixed speech pronunciation sequence to obtain the mixed pronunciation features, the mixed pronunciation features are sent to the independent conformer module in the text recognition module 420 for processing, so as to obtain the recognized text.
[0065] Step S330: Adjust the network parameters of the language recognition model by using the first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio, and the second difference between the recognized text and the reference text corresponding to the mixed speech pronunciation sequence.
[0066] In some embodiments, the first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio can be used to determine the first loss value; then the second difference between the recognized text and the reference text corresponding to the mixed speech pronunciation sequence can be used to determine the second loss value. The first loss value and the second loss value are fused and weighted to obtain the final loss value, and the network parameters of the language recognition model 400 are adjusted according to the final loss value. After multiple iterative processes, when the final loss value reaches a certain convergence condition, the final language recognition model 400 can be obtained. This model can accurately convert the audio signal into the corresponding multi-language mixed text. Among them, the calculation methods of the first loss value and the second loss value can adopt Mean Squared Error (MSE), Cross Entropy, etc., which are not specifically limited here.
[0067] Compared with the prior art, this case uses the LLM model to generate a large number of multi-language mixed texts containing target phrases in other languages, and uses the multi-language mixed texts to assist in the training of multi-language speech recognition, so as to improve the recognition effect of the training audio of the target phrases in other languages. And in the process of generating the multi-language mixed text, multiple different fields are specified in the indicator words, so that the LLM model generates multi-language mixed texts containing target phrases in other languages in different fields, making the domain generalization ability of this solution stronger.
[0068] In addition, for the target phrases in unknown other languages, a small amount of multi-language mixed audio containing target phrases in other languages can be used to analyze the pronunciation of the target phrases, and the pronunciation of the target phrases can be constructed for multi-language mixed text assistance, further improving the recall rate of the target phrases and the overall recognition effect.
[0069] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0070] Please refer to Figure 5 ,Figure 5 It is a schematic diagram of the framework of an embodiment of the electronic device 50 of the present application. The electronic device 50 includes a memory 51 and a processor 52 that are coupled to each other. The processor 52 is configured to execute program instructions stored in the memory 51 to implement the steps of any of the above-described embodiments of the hybrid speech pronunciation sequence generation method, or to implement the steps of any of the above-described embodiments of the model training method. In a specific implementation scenario, the electronic device 50 may include, but is not limited to, a microcomputer, a server. In addition, the electronic device 50 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0071] Specifically, the processor 52 is configured to control itself and the memory 51 to implement the steps of any of the above-described embodiments of the hybrid speech pronunciation sequence generation method, or to implement the steps of any of the above-described embodiments of the model training method. The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip having signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 52 may be implemented jointly by integrated circuit chips.
[0072] Please refer to Figure 6 , Figure 6 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium 60 of the present application. The computer-readable storage medium 60 stores program instructions 601 that can be run by a processor. The program instructions 601 are configured to implement the steps of any of the above-described embodiments of the hybrid speech pronunciation sequence generation method, or to implement the steps of any of the above-described embodiments of the model training method.
[0073] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0074] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated herein.
[0075] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in electrical, mechanical or other forms.
[0076] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.
Claims
1. A method for generating a mixed speech pronunciation sequence, characterized in that: include: Performing pronunciation prediction based on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal, wherein the original audio signal is an audio signal corresponding to a first language and carries a target phrase corresponding to a second language, and the first language and the second language are different languages; Analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target phrase; and, generating an initial text containing the target phrase; A mixed speech pronunciation sequence corresponding to the initial text is generated by using the actual pronunciation sequence of the target phrase and the initial text, wherein in the mixed speech pronunciation sequence, the target phrase is in the second language, and other text parts of the initial text except the target phrase are in a third language, and the second language and the third language are different languages.
2. The method according to claim 1, characterized in that: The method of performing pronunciation prediction based on the original audio signal to obtain a predicted pronunciation sequence of the original audio signal includes: Extracting features from the original audio signal to obtain a number of filter bank features; Predicting a number of the filter bank features to obtain a predicted pronunciation sequence of the original audio signal.
3. The method according to claim 2, characterized in that The extracting features of the original audio signal to obtain filter bank features includes: Extracting features from the original audio signal to obtain audio features in a preset format; The preset format audio feature is pre-processed to obtain the filter bank feature.
4. The method according to claim 2, characterized in that: Before predicting the plurality of filter bank features to obtain the predicted pronunciation sequence of the original audio signal, the method comprises: Performing global correlation processing on each of the filter group features to obtain a number of filter correlation features; The step of predicting the plurality of filter bank features to obtain a predicted pronunciation sequence of the original audio signal comprises: Predicting a number of the filter-related features to obtain the predicted pronunciation sequence.
5. The method according to claim 1, characterized in that The step of analyzing the predicted pronunciation sequence to obtain the actual pronunciation sequence of the target phrase includes: According to the pronunciation reference data of the first language, a pronunciation subsequence that does not belong to the first language in the predicted pronunciation sequence is used as an actual pronunciation sequence of the target phrase.
6. The method according to claim 1, characterized in that The target phrase in the initial text is in the second language, and the other text parts are in the third language; And / or, the step of using the actual pronunciation sequence of the target phrase and the initial text to generate a mixed speech pronunciation sequence corresponding to the initial text comprises: Generating other pronunciation sequences of the other text parts corresponding to the third language; The actual pronunciation sequence is combined with the other pronunciation sequences to obtain the mixed speech pronunciation sequence.
7. A model training method, characterized in that: include: Using the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio; as well as, Using a text recognition module of a language recognition model to perform text conversion on a mixed speech pronunciation sequence to obtain a recognition text corresponding to the mixed speech pronunciation sequence, wherein the mixed speech pronunciation sequence is obtained by the method according to any one of claims 1 to 6 above; The network parameters of the language recognition model are adjusted using a first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio, and a second difference between the recognized text and a reference text corresponding to the mixed speech pronunciation sequence.
8. The method according to claim 7, characterized in that The method of using the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio includes: Using the pronunciation prediction module to extract features from the training audio to obtain training audio features; Identifying the training audio features to obtain the sample predicted pronunciation sequence; And / or, the text recognition module of the language recognition model is used to convert the mixed speech pronunciation sequence into text to obtain the recognition text corresponding to the mixed speech pronunciation sequence, including: Using the text recognition module to extract features from the mixed speech pronunciation sequence to obtain mixed pronunciation features; Perform text recognition on the mixed pronunciation features to obtain the recognized text.
9. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the mixed speech pronunciation sequence generation method described in any one of claims 1 to 6, and / or the model training method described in claims 7 to 8.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by the processor, the mixed speech pronunciation sequence generation method described in any one of claims 1 to 6 and / or the model training method described in claims 7 to 8 are implemented.
Citation Information
Patent Citations
Speech synthesis method and device, electronic equipment and storage medium
CN112750419A
Mixed speech recognition method and device, storage medium and electronic device
CN113160804A
Voice processing method and device and electronic equipment
CN113539233A
Speech synthesis method and device, medium and electronic equipment
CN114242035A
Pronunciation evaluation method, training method of pronunciation evaluation system, device and equipment
CN114283788A