Hybrid speech pronunciation sequence generation method and model training method, and related devices

By performing pronunciation prediction on the original audio signal and generating initial text using a large language model, the problem of poor recognition performance of multilingual mixed audio is solved, achieving efficient recognition and improved training results in the absence of data.

CN120071893BActive Publication Date: 2025-11-07IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510036970.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-11-07
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing automatic speech recognition models perform poorly when dealing with mixed multilingual audio, mainly because there are few mixed multilingual audio samples during training, especially lacking pronunciation of specialized phrases, making accurate recognition difficult.

Method used

By predicting pronunciation based on the original audio signal, a mixed speech pronunciation sequence containing the target word is generated. The initial text is generated using a large language model, and the training audio information is combined to assist in training the language recognition model, thereby improving the pronunciation sequence quality of multilingual mixed audio.

Benefits of technology

It improves the recognition performance and recall rate of multilingual mixed audio, enhances the training effect of language recognition models, and is able to quickly identify professional phrases, especially in the absence of relevant audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071893B_ABST
    Figure CN120071893B_ABST
Patent Text Reader

Abstract

The application discloses a mixed voice pronunciation sequence generation method and model training method, and related devices. The mixed voice pronunciation sequence generation method comprises the following steps: performing pronunciation prediction based on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal, wherein the original audio signal is an audio signal corresponding to a first language, and carries a target word group corresponding to a second language, and the first language and the second language are different languages; performing analysis on the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target word group; and generating an initial text containing the target word group; and generating a mixed voice pronunciation sequence corresponding to the initial text by using the actual pronunciation sequence of the target word group and the initial text. The above scheme can improve the quality of the pronunciation sequence of the multi-language mixed audio.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, in particular to a mixed speech pronunciation sequence generation method and model training method, and related devices. BACKGROUND

[0002] Current automatic speech recognition (ASR) models have good recognition effects when processing audio of a single language, but have poor recognition effects for multi-language mixed audio. In the prior art, the training process of the multi-language mixed audio recognition capability of the automatic speech recognition model uses few multi-language mixed audio samples. For example, when training an automatic speech recognition model for mixed Chinese and English audio of some professional English word groups, there is little audio data containing these professional English word groups, and there is a lack of pronunciation of these professional English word groups, which makes it difficult to accurately recognize during the training process. SUMMARY

[0003] The present application provides a mixed speech pronunciation sequence generation method and model training method, and related devices, which can improve the quality of multi-language mixed audio pronunciation sequences.

[0004] The first aspect of the present application provides a mixed speech pronunciation sequence generation method, which comprises: performing pronunciation prediction based on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal, wherein the original audio signal is an audio signal corresponding to a first language and carries a target word group corresponding to a second language, and the first language and the second language are different languages; analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target word group; and generating an initial text containing the target word group; and generating a mixed speech pronunciation sequence corresponding to the initial text using the actual pronunciation sequence of the target word group and the initial text, wherein in the mixed speech pronunciation sequence, the target word group is in the second language, and the other text part in the initial text except the target word group is in a third language, and the second language and the third language are different languages.

[0005] The method comprises: performing feature extraction on the original audio signal to obtain a plurality of filter bank features; and performing prediction on the plurality of filter bank features to obtain the predicted pronunciation sequence of the original audio signal.

[0006] The method comprises: performing feature extraction on the original audio signal to obtain a plurality of filter bank features; and performing prediction on the plurality of filter bank features to obtain the predicted pronunciation sequence of the original audio signal.

[0007] The method comprises: performing global correlation processing on the filter bank features to obtain filter correlation features; and predicting the filter bank features to obtain a predicted pronunciation sequence of the original audio signal, which comprises predicting the filter correlation features to obtain the predicted pronunciation sequence.

[0008] The method comprises: performing global correlation processing on the filter bank features to obtain filter correlation features; and predicting the filter bank features to obtain a predicted pronunciation sequence of the original audio signal, which comprises predicting the filter correlation features to obtain the predicted pronunciation sequence.

[0009] The method comprises: performing global correlation processing on the filter bank features to obtain filter correlation features; and predicting the filter bank features to obtain a predicted pronunciation sequence of the original audio signal, which comprises predicting the filter correlation features to obtain the predicted pronunciation sequence.

[0010] The second aspect of the application provides a model training method, which comprises: performing pronunciation prediction on training audio by using a pronunciation prediction module of a language recognition model to obtain a sample predicted pronunciation sequence of the training audio; and performing text conversion on a mixed voice pronunciation sequence by using a text recognition module of the language recognition model to obtain recognized text corresponding to the mixed voice pronunciation sequence, wherein the mixed voice pronunciation sequence is obtained by the method of any one of the first aspect; and adjusting network parameters of the language recognition model by using a first difference between the sample predicted pronunciation sequence and a standard pronunciation sequence of the training audio and a second difference between the recognized text and reference text corresponding to the mixed voice pronunciation sequence.

[0011] The second aspect of the application provides a model training method, which comprises: performing pronunciation prediction on training audio by using a pronunciation prediction module of a language recognition model to obtain a sample predicted pronunciation sequence of the training audio; and performing text conversion on a mixed voice pronunciation sequence by using a text recognition module of the language recognition model to obtain recognized text corresponding to the mixed voice pronunciation sequence, wherein the mixed voice pronunciation sequence is obtained by the method of any one of the first aspect; and adjusting network parameters of the language recognition model by using a first difference between the sample predicted pronunciation sequence and a standard pronunciation sequence of the training audio and a second difference between the recognized text and reference text corresponding to the mixed voice pronunciation sequence.

[0012] The third aspect of the application provides an electronic device, which comprises a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the mixed voice pronunciation sequence generation method in the first aspect and / or the model training method in the second aspect.

[0013] The fourth aspect of the present application provides a computer-readable storage medium, which stores program instructions. The program instructions are executed by a processor to implement the mixed speech pronunciation sequence generation method in the first aspect and / or implement the model training method in the second aspect.

[0014] The above scheme first predicts the pronunciation of the original audio signal to obtain a predicted pronunciation sequence corresponding to the original audio signal, then analyzes the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target word group, simultaneously generates an initial text containing the target word group, and then uses the actual pronunciation sequence of the target word group and the initial text to generate a mixed speech pronunciation sequence corresponding to the initial text, so that the mixed speech pronunciation sequence also contains the actual pronunciation sequence of the target word group in the second language, thereby improving the quality of the mixed pronunciation sequence corresponding to the initial text.

[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present application. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.

[0017] Figure 1 is a flowchart of an embodiment of the mixed speech pronunciation sequence generation method of the present application;

[0018] Figure 2 is a flowchart of another embodiment of the mixed speech pronunciation sequence generation method of the present application;

[0019] Figure 3 is a flowchart of an embodiment of the model training method of the present application;

[0020] Figure 4 is a framework diagram of an embodiment of the speech recognition model of the present application;

[0021] Figure 5 is a framework diagram of an embodiment of the electronic device of the present application;

[0022] Figure 6 is a framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0023] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0024] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to thoroughly understand the present application, but are not intended to limit the present application.

[0025] The term "and / or", used herein, merely describes an associated relationship, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents that the associated objects before and after are in an "or" relationship. In addition, "multiple" herein means two or more than two. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple, for example, including at least one of A, B and C can mean including any one or more elements selected from the set consisting of A, B and C.

[0026] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the method for generating a mixed language pronunciation sequence of the application. Specifically, it can include the following steps:

[0027] Step S110: Perform pronunciation prediction based on the original audio signal to obtain a predicted pronunciation sequence of the original audio signal.

[0028] The original audio signal is an audio signal corresponding to a first language, and carries a target phrase corresponding to a second language, and the first language and the second language are different languages. For example, the first language can be Chinese, and the second language can be English, French, Persian, etc., so the first language and the second language are not specifically limited.

[0029] The present application mainly relates to the fields of audio signal processing, natural language processing and deep learning, and can quickly predict the pronunciation of some professional target phrases using a small amount of existing audio signals, and can generate a large amount of multi-language mixed text containing these professional target phrases using LLM (Large Language Model). Finally, the constructed multi-language mixed text and the training audio information are used to jointly assist in training the ASR model, so that even in the case of serious lack of related audio, the recall rate and recognition effect of the target phrase can be quickly improved.

[0030] In some embodiments, the first language is a major language, and a pronunciation dictionary corresponding to the first language is known. Therefore, when performing pronunciation prediction on the original audio signal, the original audio signal can be recognized first to determine segmentation points in the original audio signal, and the original audio signal is divided by using the segmentation points to obtain a plurality of sub-audio signals, wherein each sub-audio signal corresponds to a word or a phrase in the first language. Then, the pronunciation dictionary corresponding to the first language and each sub-audio signal are subjected to similarity calculation, after obtaining a plurality of similarities corresponding to each sub-audio signal, the pronunciation sequence in the pronunciation dictionary with the highest similarity to the sub-audio signal is selected as the predicted pronunciation sequence of the sub-audio signal. Then, according to the position order of the sub-audio signal in the original audio signal, the predicted pronunciation sequences are combined to obtain the predicted pronunciation sequence of the original audio signal.

[0031] In other embodiments, the audio features in the original audio signal can be recognized and analyzed to obtain the predicted pronunciation sequence of the original audio signal. Specifically, reference can be made to steps S111 to S112.

[0032] Step S111: performing feature extraction on the original audio signal to obtain a plurality of filter bank features.

[0033] In some embodiments, the original audio signal can be subjected to feature extraction first to obtain audio features in a preset format. Then, the audio features in the preset format are subjected to pre-processing to obtain filter bank features. The filter bank features can correspond to a phrase in the original audio signal; the audio features in the preset format can be one-dimensional WAV (Waveform Audio File Format, lossless compressed audio file format) audio features, etc., which are not specifically limited herein. These filter bank features can reflect important characteristics of the original audio signal, and provide a basis for subsequent pronunciation prediction.

[0034] For example, the one-dimensional WAV audio features are subjected to pre-processing to obtain fb40 features, i.e., filter bank features, wherein "fb" is the abbreviation of fbank, fbank is a filter bank, and "40" represents the byte length of the filter bank features.

[0035] Step S112: performing prediction on the plurality of filter bank features to obtain the predicted pronunciation sequence of the original audio signal.

[0036] In some embodiments, after obtaining the filter bank features, global correlation processing can be performed on each filter bank feature to obtain filter correlation features, so as to strengthen the correlation between the features and improve the accuracy of pronunciation prediction. Specifically, the filter bank features can be input into a pre-trained pronunciation prediction model, and the N Conformer modules in the pronunciation prediction model can be used to perform global correlation processing on the filter bank features. Then, the filter correlation features are predicted to obtain a predicted pronunciation sequence. For example, the original audio signal is an audio signal of “extracting fbank features”, and after feature processing, the filter correlation features corresponding to “extracting fbank features” are predicted to obtain a predicted pronunciation sequence “t i2q u3ei4fb an4k t e4zh eng1”.

[0037] Step S120: analyzing the predicted pronunciation sequence to obtain the actual pronunciation sequence of the target phrase.

[0038] In some embodiments, according to the pronunciation reference data of the first language, the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language can be taken as the actual pronunciation sequence of the target phrase. Specifically, the pronunciation reference data can be a pronunciation dictionary. By using the pronunciation dictionary of the first language, the predicted pronunciation sequence can be analyzed to determine the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language, and the pronunciation subsequence can be taken as the actual pronunciation sequence of the target phrase.

[0039] For example, the pronunciation dictionary and the pronunciation subsequence in the predicted pronunciation sequence are used to calculate the similarity, the dictionary pronunciation sequence with the highest similarity to the pronunciation subsequence in the pronunciation dictionary is selected, and the highest dictionary pronunciation sequence is compared with a similarity threshold. If the similarity of the dictionary pronunciation sequence is greater than or equal to the similarity threshold, it indicates that the pronunciation subsequence corresponding to the dictionary pronunciation sequence belongs to the first language. If the similarity of the dictionary pronunciation sequence is less than the similarity threshold, it indicates that the pronunciation subsequence corresponding to the dictionary pronunciation sequence does not belong to the first language, and then the pronunciation subsequence in the predicted pronunciation sequence that does not belong to the first language is taken as the actual pronunciation sequence of the target phrase.

[0040] In some embodiments, the correlation between the sub-sequences of the predicted pronunciation sequence can be calculated, and the sub-sequence of the predicted pronunciation sequence that is least correlated with other sub-sequences is taken as the actual pronunciation sequence of the target word group. Specifically, the features of each sub-sequence of the predicted pronunciation sequence are extracted, the correlation is calculated based on the features, and the correlation coefficient of the features of a sub-sequence with the features of other sub-sequences is fused to obtain the final correlation coefficient of the sub-sequence. If the final correlation coefficient is less than the median of the final correlation coefficients of all sub-sequences, the sub-sequence corresponding to the final correlation coefficient does not belong to the sub-sequences of the first language. Therefore, the sub-sequence is taken as the actual pronunciation sequence of the target word group.

[0041] It can be understood that the correlation calculation method can use Pearson correlation, Spearman correlation, etc., which are not specifically limited here.

[0042] Step S130: generating an initial text containing the target word group.

[0043] In some embodiments, a large language model can be used to generate an initial text containing the target word group. In order to generate an initial text with high quality, a suitable prompt word needs to be created, and in order to enhance the domain diversity of the generated initial text, the prompt word needs to specify the generation of text in different domains, such as "finance", "technology", "sports", "history and culture", etc. The large language model adopts the mainstream multi-layer transformer structure and has strong context modeling capability. The generated sentences have good effects in terms of diversity, instruction following, semantic fluency, etc.

[0044] For example, the prompt word is a text containing the target word group "negotiation strategy", and it is specified to generate a text in the finance domain. Therefore, after receiving the prompt word, the large language model will generate a large number of initial texts similar to "In this negotiation, we need to use effective negotiation strategies to achieve a satisfactory result for both parties."

[0045] In some other embodiments, the target word group of the second language is first translated to obtain the corresponding target word group text of the first language. According to the target word group text, a search is performed in the database, and the text containing the target word group text is taken as a candidate text. Then, the target word group text in the candidate text is replaced with the target word group of the second language, thereby obtaining the initial text.

[0046] In one implementation scenario, the above steps S120 and S130 can be executed in a sequential order, for example, step S120 is executed first and then step S130 is executed, or step S130 is executed first and then step S120 is executed. In another implementation scenario, the above steps S120 and S130 can also be executed simultaneously, which can be set according to actual application and is not limited herein.

[0047] Step S140: generating a mixed speech pronunciation sequence corresponding to the initial text by using the actual pronunciation sequence of the target word group and the initial text.

[0048] In the mixed speech pronunciation sequence, the target word group is in a second language, the other text part in the initial text except the target word group is in a third language, and the second language and the third language are different languages. The third language can be the same as the first language or different from the first language.

[0049] For example, in an international business negotiation speech assistance system, when a Chinese text containing a specific professional term (the target word group is English, such as "negotiation strategy") needs to be synthesized, an initial text can be generated by a preset text template or according to real-time input theme information, such as "In this negotiation, we need to use an effective negotiation strategy to achieve a satisfactory result for both parties." Here, the target word group "negotiation strategy" is in the second language English, and the other text part in the initial text except the target word group is in the third language Chinese.

[0050] In some embodiments, other pronunciation sequences corresponding to the third language for the other text part can be generated by using pronunciation reference data of the third language. Then, the actual pronunciation sequence and the other pronunciation sequence are combined to obtain the mixed speech pronunciation sequence.

[0051] Taking a speech synthesis method based on deep learning as an example, the third language is Chinese and the second language is English. The Chinese text in the initial text can be first feature-extracted to obtain audio features in a preset format, and then preprocessed to obtain filter bank features. Then, the filter bank features are predicted to obtain the pronunciation sequence of the Chinese part. For example, for the Chinese text "in this negotiation, we need to use an effective", the corresponding Chinese pronunciation sequence can be obtained after a series of processing.

[0052] After obtaining the actual pronunciation sequence of the target phrase (such as the pronunciation sequence of the English phrase "negotiation strategy") and the pronunciation sequence of the Chinese part, they can be combined to obtain a mixed speech pronunciation sequence. In the mixed speech pronunciation sequence generation, the actual pronunciation sequence of the English phrase and the pronunciation sequence of the Chinese part are combined according to certain rules, so that in the mixed speech pronunciation sequence, the target phrase is the second language, and the other text part in the initial text except the target phrase is the third language. For example, in the above example, the final mixed speech pronunciation sequence obtained is the combination of the Chinese pronunciation sequence and the pronunciation sequence of the English phrase "negotiation strategy".

[0053] In a specific implementation scenario, the first language and the third language are both Chinese, the second language is English, the original audio signal is "extract fbank feature", and the target phrase is "fbank". Please refer to Figure 2 The original audio signal "extract fbank feature" is input into the pronunciation prediction model, and after receiving "extract fbank feature", the pronunciation prediction model performs feature extraction on the original audio signal "extract fbank feature" to obtain one-dimensional WAV audio features. Then, the one-dimensional WAV audio features are preprocessed to obtain a plurality of fb40 features. Then, the obtained plurality of fb40 features are sent to the N Conformer modules in the pronunciation prediction model for global correlation processing to obtain a plurality of filter correlation features. Then, the prediction layer in the pronunciation prediction model is used to predict the plurality of filter correlation features to obtain the predicted pronunciation sequence "ti2qu3ei4fb an4k t e4zh eng1" corresponding to the original audio signal. Finally, the predicted pronunciation sequence is inferred using known Chinese pronunciation reference data to obtain the actual pronunciation sequence "ei4fb an4k" of the target phrase, so that the pronunciation sequence of the target phrase is determined.

[0054] At the same time, the prompt word is set and different domains are specified, so that the large language model can generate the initial text containing the target phrase "fbank" according to the prompt word. The generated initial text, the known Chinese pronunciation reference data, and the actual pronunciation sequence of the target phrase "fbank" are combined to obtain the mixed speech pronunciation sequence corresponding to the initial text. Then, the mixed speech pronunciation sequence corresponding to the initial text is used to train the language recognition model in combination with the training audio.

[0055] The scheme can utilize a small amount of existing audio signals to determine the actual pronunciation sequence of the target word group corresponding to the second language, and convert the actual pronunciation sequence of the target word group and the initial text to obtain the mixed speech pronunciation sequence corresponding to the initial text, thereby improving the quality of the generated mixed speech pronunciation sequence and laying a good foundation for the training of the subsequent language recognition model to improve the training effect of the language recognition model.

[0056] The training method of the language recognition model can be referred to in Figure 3 , Figure 3 is a flowchart of an embodiment of the model training method of the present application. Specifically, it can include the following steps:

[0057] Step S310: using the pronunciation prediction module of the language recognition model to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio.

[0058] Please refer to Figure 4 , the language recognition model 400 includes a pronunciation prediction module 410 and a text recognition module 420. The pronunciation prediction module 410 is used to perform pronunciation prediction on the training audio to obtain a sample predicted pronunciation sequence of the training audio. The text recognition module 420 is used to perform text conversion on the mixed speech pronunciation sequence to obtain a recognized text corresponding to the mixed speech pronunciation sequence.

[0059] Therefore, after receiving the training audio, the language recognition model 400 uses the pronunciation prediction module 410 to perform feature extraction on the training audio to obtain training audio features; and performs recognition on the training audio features to obtain a sample predicted pronunciation sequence.

[0060] For example, the pronunciation prediction module 410 is used to perform feature extraction on the training audio to obtain training audio features, and then 2D convolution is used to perform down-sampling, frame length reduction processing on the training audio features, and then global correlation processing is performed in the 16-layer conformer module in the pronunciation prediction module 410, thereby obtaining a sample predicted pronunciation sequence.

[0061] Step S320: using the text recognition module of the language recognition model to perform text conversion on the mixed speech pronunciation sequence to obtain a recognized text corresponding to the mixed speech pronunciation sequence.

[0062] The mixed speech pronunciation sequence is obtained according to the mixed speech pronunciation sequence generation method described above.

[0063] In some embodiments, the text recognition module 420 is used to perform feature extraction on the mixed speech pronunciation sequence to obtain mixed pronunciation features; and the mixed pronunciation features are subjected to text recognition to obtain a recognized text.

[0064] For example, after the text recognition module 420 performs feature extraction on the mixed speech pronunciation sequence to obtain mixed pronunciation features, the mixed pronunciation features are sent to an independent conformer module in the text recognition module 420 for processing, so as to obtain the recognized text.

[0065] Step S330: Adjust the network parameters of the language recognition model by using the first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio, and the second difference between the recognized text and the reference text corresponding to the mixed speech pronunciation sequence.

[0066] In some embodiments, the first loss value can be determined by using the first difference between the sample predicted pronunciation sequence and the standard pronunciation sequence of the training audio, and the second loss value can be determined by using the second difference between the recognized text and the reference text corresponding to the mixed speech pronunciation sequence. The first loss value and the second loss value are fused and weighted to obtain a final loss value, and the network parameters of the language recognition model 400 are adjusted according to the final loss value. After multiple iterations, when the final loss value reaches a certain convergence condition, the final language recognition model 400 can be obtained. This model can accurately convert the audio signal into the corresponding multilingual mixed text. The calculation method of the first loss value and the second loss value can use Mean Squared Error (MSE), Cross Entropy, etc., which is not limited here.

[0067] Compared with the prior art, the present case generates a large number of multilingual mixed texts containing target word groups of other languages by using the LLM model, and assists the training of multilingual speech recognition by using the multilingual mixed texts, thereby improving the recognition effect of the training audio containing target word groups of other languages. And in the process of generating multilingual mixed texts, multiple different fields are specified in the demonstrative words, so that the LLM model generates multilingual mixed texts containing target word groups of other languages in different fields, so that the domain generalization ability of the present scheme is stronger.

[0068] In addition, for unknown target word groups of other languages, a small amount of multilingual mixed audio containing target word groups of other languages can be used to analyze the pronunciation of the target word groups, and the pronunciation of the target word groups is constructed to assist the multilingual mixed text, which further improves the recall rate of the target word groups and the overall recognition effect.

[0069] Those skilled in the art can understand that in the above method of the specific implementation, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0070] Please refer to Figure 5 ,Figure 5 is a schematic diagram of a framework of an embodiment of the electronic device 50 of the present application. The electronic device 50 comprises a memory 51 and a processor 52 coupled with each other. The processor 52 is configured to execute program instructions stored in the memory 51 to implement the steps of any of the above-described embodiments of the method for generating a mixed speech pronunciation sequence, or implement the steps of any of the above-described embodiments of the method for training a model. In a specific implementation scenario, the electronic device 50 can include, but is not limited to, a microcomputer, a server, and in addition, the electronic device 50 can also include a notebook computer, a tablet computer, and other mobile devices, which are not limited herein.

[0071] Specifically, the processor 52 is configured to control itself and the memory 51 to implement the steps of any of the above-described embodiments of the method for generating a mixed speech pronunciation sequence, or implement the steps of any of the above-described embodiments of the method for training a model. The processor 52 can also be referred to as a CPU (Central Processing Unit). The processor 52 can be an integrated circuit chip having a processing capability of signals. The processor 52 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. In addition, the processor 52 can be jointly implemented by integrated circuit chips.

[0072] Please refer to Figure 6 , Figure 6 is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 60 of the present application. The computer-readable storage medium 60 stores program instructions 601 capable of being executed by a processor, and the program instructions 601 are configured to implement the steps of any of the above-described embodiments of the method for generating a mixed speech pronunciation sequence, or implement the steps of any of the above-described embodiments of the method for training a model.

[0073] In some embodiments, the apparatus provided by the embodiments of the present application has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, they will not be described here again.

[0074] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For the sake of brevity, they will not be described here again.

[0075] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0076] In addition, each function unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist alone physically, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of a software function unit.

[0077] If the integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various other media that can store program codes.

Claims

1. A method of generating a mixed-voice phonetic sequence, the method comprising: The method comprises: performing pronunciation prediction based on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal, wherein the original audio signal is an audio signal corresponding to a first language and carrying a target phrase corresponding to a second language, and the first language and the second language are different languages; analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target phrase; and generating an initial text containing the target phrase; using the actual pronunciation sequence of the target phrase and the initial text to generate a mixed voice pronunciation sequence corresponding to the initial text, wherein in the mixed voice pronunciation sequence, the target phrase is in the second language, and other text parts in the initial text except the target phrase are in a third language, and the second language and the third language are different languages.

2. The method of claim 1, wherein, The method of performing pronunciation prediction based on an original audio signal to obtain a predicted pronunciation sequence of the original audio signal comprises: performing feature extraction on the original audio signal to obtain a plurality of filter bank features; performing prediction on a plurality of the filter bank features to obtain the predicted pronunciation sequence of the original audio signal.

3. The method of claim 2, wherein, The method of performing feature extraction on the original audio signal to obtain a filter bank feature comprises: performing feature extraction on the original audio signal to obtain a preset format audio feature; performing pre-processing on the preset format audio feature to obtain the filter bank feature.

4. The method of claim 2, wherein, Before the method of performing prediction on a plurality of the filter bank features to obtain the predicted pronunciation sequence of the original audio signal, the method comprises: performing global correlation processing on each of the filter bank features to obtain a plurality of filter correlation features; The method of performing prediction on a plurality of the filter bank features to obtain the predicted pronunciation sequence of the original audio signal comprises: performing prediction on a plurality of the filter correlation features to obtain the predicted pronunciation sequence.

5. The method of claim 1, wherein, The method of analyzing the predicted pronunciation sequence to obtain an actual pronunciation sequence of the target phrase comprises: According to pronunciation reference data of the first language, taking a pronunciation sub-sequence in the predicted pronunciation sequence that does not belong to the first language as the actual pronunciation sequence of the target phrase.

6. The method of claim 1, wherein, The target phrase in the initial text is in the second language, and the other text parts are in the third language. And / or, the method of using the actual pronunciation sequence of the target phrase and the initial text to generate a mixed voice pronunciation sequence corresponding to the initial text comprises: generating other pronunciation sequences corresponding to the third language for the other text parts; combining the actual pronunciation sequence and the other pronunciation sequences to obtain the mixed voice pronunciation sequence.

7. A model training method, comprising: The method comprises: using a pronunciation prediction module of a language recognition model to perform pronunciation prediction on training audio to obtain a sample predicted pronunciation sequence of the training audio; and using a text recognition module of the language recognition model to perform text conversion on a mixed voice pronunciation sequence to obtain a recognized text corresponding to the mixed voice pronunciation sequence, wherein the mixed voice pronunciation sequence is obtained by the method of any one of claims 1 to 6. The first difference between the sample pronunciation sequence and a standard pronunciation sequence of the training audio and the second difference between the recognized text and a reference text corresponding to the mixed speech pronunciation sequence are used to adjust network parameters of a language recognition model.

8. The method of claim 7, wherein, The pronunciation prediction module using the language recognition model performs pronunciation prediction on the training audio to obtain a sample pronunciation sequence of the training audio, including: The pronunciation prediction module using the language recognition model performs pronunciation prediction on the training audio to obtain a sample pronunciation sequence of the training audio, including: The training audio features are recognized to obtain the sample pronunciation sequence; And / or, the text recognition module using the language recognition model performs text conversion on the mixed speech pronunciation sequence to obtain a recognized text corresponding to the mixed speech pronunciation sequence, including: The text recognition module using the language recognition model performs text conversion on the mixed speech pronunciation sequence to obtain a recognized text corresponding to the mixed speech pronunciation sequence, including: The mixed pronunciation features are recognized to obtain the recognized text.

9. An electronic device, comprising: The processor is configured to execute program instructions stored in the memory to implement the mixed speech pronunciation sequence generation method of any one of claims 1 to 6, and / or implement the model training method of claims 7 to 8.

10. A computer-readable storage medium having stored thereon program instructions, wherein, The program instructions, when executed by the processor, implement the mixed speech pronunciation sequence generation method of any one of claims 1 to 6, and / or implement the model training method of claims 7 to 8.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN112750419A

  • Mixed speech recognition method and device, storage medium and electronic device

    CN113160804A