Audio generation method and device

By obtaining and processing the verbal corpus in the target business scenario and generating audio data sets that are in line with the business scenario, the problem that traditional speech synthesis models cannot generate specific speeches for actual business scenarios is solved, and high-quality speech synthesis model training is achieved.

CN115171646BActive Publication Date: 2025-05-13DINGFU NEW POWER (BEIJING) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210792253.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-05-13
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Traditional speech synthesis models cannot generate specific speeches for actual business scenarios during training, resulting in low quality of generated speech.

Method used

By obtaining the term corpus in the target business scenario, performing smoothness checks and combinations, a second set of speeches containing the target corpus is generated, and recording in a preset recording environment, the initial audio data set is obtained. Then, the initial audio dataset is normalized with the public dataset to generate the target audio dataset.

Benefits of technology

The generated audio data set is consistent with the actual business scenario, with high recording quality, and can improve the training accuracy and generation quality of the speech synthesis model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171646B_ABST
    Figure CN115171646B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an audio generation method and device, the method comprising obtaining a first speech set, checking the smoothness of the speech language in the first speech set based on a preset language model, determining the target language, and generating a second speech set containing the target language, and recording the second speech set under a preset recording environment to obtain an initial audio data set. The initial audio data set and the preset public data set are normalized to obtain a target audio data set. The present application can generate a second speech set based on a target business scenario, so that the speech language in the second speech set fits the target business scenario. The second speech set can also be recorded under a preset recording environment to ensure the recording effect. In addition, a target audio data set can be generated based on the initial audio data set and the public data set, and the target audio data set can be applied to the speech synthesis model training process to ensure the accuracy of the training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to an audio generation method and device. Background Art

[0002] Speech synthesis is a technology that generates artificial speech, such as speech for marketing. With the rapid development of artificial intelligence, higher requirements are placed on speech synthesis technology.

[0003] At present, speech synthesis tasks generally use known speech datasets to train speech synthesis models. Known speech datasets are public datasets, such as aishell-3 and Biaobei data. The generation process of known speech datasets includes determining the speech script, the speaker making sounds based on the speech script, and recording the sounds made by the speaker.

[0004] However, the speech used to generate known speech data sets has a certain degree of randomness, and the content of the speech is single and has a low relevance to actual business scenarios. Therefore, the speech synthesis model trained based on the known speech data set is usually of low quality and cannot generate specific speech for actual business scenarios. Summary of the invention

[0005] The embodiments of the present application provide an audio generation method and device to solve the problem that when traditional audio data is used to train a speech synthesis model, the speech synthesis model cannot generate specific speech for actual business scenarios.

[0006] In a first aspect, an embodiment of the present application provides an audio generation method, the method comprising: obtaining a first speech set, the first speech set comprising a plurality of speech corpora, the speech corpora comprising a first speech corpus and a second speech corpus, the first speech corpus being collected from a target business scenario, and the second speech corpus being composed of a plurality of language elements with different sentence components; performing a smoothness check on the speech corpus in the first speech set based on a preset language model, determining a target corpus, and generating a second speech set containing the target corpus, the smoothness of the target corpus being greater than a preset threshold; recording the second speech set under a preset recording environment to obtain an initial audio data set, the initial audio data set including audio data corresponding to each speech corpus in the second speech set; normalizing the initial audio data set and a preset public data set to obtain a target audio data set; the normalization process is used to adjust the amplitudes of the initial audio data set and the public data set to within a preset range.

[0007] In one feasible method, an initial audio data set and a preset public data set are normalized to obtain a target audio data set, including: converting the initial audio data set and the public data set into a matrix form to obtain a first matrix; determining multiple matrix parameters of the first matrix, the matrix parameters including the median, mean and / or mode of the first matrix; using each matrix parameter, respectively normalizing the first part of the audio data to obtain a normalized result corresponding to each matrix parameter; the first part of the audio data includes part of the data in the initial audio data set and part of the data in the public data set; based on the audition effect of the normalized results corresponding to each matrix parameter, determining the normalization parameter from the multiple matrix parameters, the normalization parameter being the matrix parameter corresponding to the optimal audition effect; using the normalization parameter to normalize the initial audio data set and the public data set to obtain the target audio data set.

[0008] In one achievable method, the smoothness of the speech utterances in the first speech set is checked based on a preset language model, the target corpus is determined, and a second speech set containing the target corpus is generated, including: inputting the speech utterances in the first speech set into the preset language model to obtain the probability value of each speech utterance, the probability value being used to indicate the smoothness of the speech utterance; and determining the speech utterances having a probability value greater than a preset threshold as the target corpus.

[0009] In one feasible method, it also includes: randomly extracting the target corpus in the second speech set according to a preset sampling ratio; if the extracted target corpus contains preset grammatical defects, removing the extracted target corpus from the second speech set and randomly extracting again; if the target corpus extracted for N consecutive times does not contain the preset grammatical defects, then the extraction is terminated.

[0010] In one achievable manner, the second speech material is obtained by the following steps: determining multiple first candidate sets, the first candidate sets including multiple language elements corresponding to sentence components, the language elements in different first candidate sets having different sentence components, the language elements being determined based on speech material collected from the target business scenario; extracting one or more language elements corresponding to sentence components from the multiple first candidate sets; and combining the extracted language elements to obtain the second speech material.

[0011] In one achievable method, the second speech material is obtained by the following steps: obtaining a second candidate set, the second candidate set including a plurality of pre-set first speech samples, and a plurality of second speech samples determined based on speech material collected from business scenarios; extracting at least one first speech sample and at least one second speech sample from the second candidate set; combining the extracted at least one first speech sample and at least one second speech sample to obtain the second speech material.

[0012] In one achievable manner, it also includes: before recording the second speech set, determining the recording format of the initial audio data set, the recording format including the number of channels and / or the sampling rate.

[0013] In one achievable manner, the preset language model is an n-gram model.

[0014] In one achievable manner, recording the second speech set in a preset recording environment includes: having a specific speaker record the second speech set in a preset recording environment.

[0015] In a second aspect, an embodiment of the present application provides an audio generation device, which includes: an acquisition module, used to acquire a first speech set, the first speech set includes multiple speech corpora, the speech corpora include first speech corpora and second speech corpora, the first speech corpora are collected from a target business scenario, and the second speech corpora are composed of multiple language elements with different sentence components; a smoothness check module, used to perform a smoothness check on the speech corpora in the first speech set based on a preset language model, determine the target corpus, and generate a second speech set containing the target corpus, the smoothness of the target corpus is greater than a preset threshold; a recording module, used to record the second speech set under a preset recording environment to obtain an initial audio data set, the initial audio data set including audio data corresponding to each speech corpus in the second speech set; a normalization module, used to perform normalization processing on the initial audio data set and a preset public data set to obtain a target audio data set; the normalization processing is used to adjust the amplitude of the initial audio data set and the public data set to within a preset range.

[0016] It can be seen from the above technical solutions that the embodiment of the present application provides an audio generation method and device, the method includes obtaining a first speech set, checking the smoothness of the speech language in the first speech set based on a preset language model, determining the target corpus, and generating a second speech set containing the target corpus, and recording the second speech set under a preset recording environment to obtain an initial audio data set. The initial audio data set and the preset public data set are normalized to obtain a target audio data set. The present application can generate a second speech set based on a target business scenario, so that the speech language in the second speech set fits the target business scenario. The second speech set can also be recorded under a preset recording environment to ensure the recording effect. In addition, a target audio data set can be generated based on the initial audio data set and the public data set, and the target audio data set can be applied to the speech synthesis model training process to ensure the accuracy of the training. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of an audio generation method provided in an embodiment of the present application;

[0018] Figure 2 A schematic diagram of a process for determining second speech material provided in an embodiment of the present application;

[0019] Figure 3 Another schematic diagram of a process for determining second speech material provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of a smoothness inspection process provided in an embodiment of the present application;

[0021] Figure 5 A schematic diagram of a process for randomly extracting a second speech set provided in an embodiment of the present application;

[0022] Figure 6 A schematic diagram of a normalization process provided in an embodiment of the present application;

[0023] Figure 7 A schematic diagram of the structure of an audio generating device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application.

[0025] Speech synthesis is a technology that generates artificial speech, such as speech for marketing. With the rapid development of artificial intelligence, higher requirements are placed on speech synthesis technology.

[0026] At present, speech synthesis tasks generally use known speech datasets to train speech synthesis models. Known speech datasets are public datasets, such as aishell-3 and Biaobei data. The generation process of known speech datasets includes determining the speech script, the speaker making sounds based on the speech script, and recording the sounds made by the speaker.

[0027] However, the speech used to generate known speech data sets has a certain degree of randomness, and the content of the speech is single and has a low relevance to actual business scenarios. Therefore, the speech synthesis model trained based on the known speech data set is usually of low quality and cannot generate specific speech for actual business scenarios.

[0028] An embodiment of the present application provides an audio generation method, which can generate an audio data set based on a specific business scenario. The audio data set includes audio data that is highly relevant to the business scenario, and can solve the problem that the speech synthesis model cannot generate specific speech for the actual business scenario.

[0029] The audio data generated by the audio generation method according to the embodiment of the present application can be used to train a speech synthesis model, specifically a TTS (Text To Speech) model. In addition, it can also be used to train other speech synthesis models, which is not specifically limited in the present application.

[0030] Figure 1 Schematic diagram of the flow of the audio generation method provided in the embodiment of the present application. Figure 1 As shown, the audio generation method provided in the embodiment of the present application includes the following steps:

[0031] S101: Obtain a first speech set, the first speech set includes multiple speech literals, the speech literals include first speech literals and second speech literals, the first speech literals are collected from a target business scenario, and the second speech literals are composed of multiple language elements with different sentence components.

[0032] The first speech set provided in the embodiment of the present application is closely related to the real business scenario. The business scenario can be a scenario selected according to the actual situation, such as: marketing scenario, usage scenario, after-sales scenario, etc. When it is necessary to generate audio data related to any of the above business scenarios, this business scenario is the target business scenario. For example: when it is necessary to generate some audio data related to the marketing scenario, the marketing scenario is the target business scenario. At this time, the speech material collected based on the marketing scenario can be used to generate the first speech set. The speech material collected based on the marketing scenario can be asking the user's name, gender, and purchase intention, etc. For example, "Hello, what is your name?" and "Does your family need to buy an air conditioner?"

[0033] In an embodiment of the present application, the first speech set may include a first speech material and a second speech material. Among them, the first speech material may be a speech material directly collected from a real target business scenario, so that the first speech material is highly consistent with the target business scenario. The second speech material may be composed of a plurality of language elements with different sentence components, and the language elements may be determined by the speech material collected from the target business scenario, so that the second speech material is consistent with the target business scenario. At the same time, since the second speech material is randomly combined from a plurality of language elements and the combination process is flexible, the second speech material includes richer content. In this way, the speech material in the first speech set is rich in content and highly relevant to the target business scenario.

[0034] S102: Check the fluency of the speech corpus in the first speech set based on a preset language model, determine the target corpus, and generate a second speech set including the target corpus, wherein the fluency of the target corpus is greater than a preset threshold.

[0035] The smoothness check of the dialogue language material can determine whether the dialogue language material conforms to the language expression habits, whether there are grammatical defects, whether it can accurately express the meaning, etc. Only the dialogue language material whose semantic smoothness meets the preset conditions can be selected into the second dialogue set.

[0036] Performing a smoothness check on the speech language materials in the first speech language set to obtain the second speech language set can improve the quality of the speech language materials and ensure the accuracy of the speech language materials.

[0037] S103: Record the second speech set in a preset recording environment to obtain an initial audio data set, wherein the initial audio data set includes audio data corresponding to each speech material in the second speech set, and control parameters of the recording environment include noise level and / or reverberation level.

[0038] When recording the second speech material, the recording environment will directly affect the quality of the audio data. Therefore, the embodiment of the present application also includes recording the second speech set under a preset recording environment. Among them, the noise level and the reverberation size of the recording environment will affect the recording effect, so the noise level and the reverberation size can be used as control parameters for controlling the recording environment, and the recording environment can be changed by adjusting the control parameters. The noise level can be specifically less than 30 decibels, and the reverberation size can be determined by the needs of the actual business scenario, generally not less than the minimum requirements of the business. This embodiment does not make specific restrictions on the reverberation size. In this way, the appearance of background sound, noise, etc. in the initial audio data set can be avoided, which is conducive to improving the clarity of the recorded audio data, and then improving the training effect of the speech synthesis model.

[0039] In some implementations, the embodiments of the present application may also add a step of determining the recording format of the initial audio data set before step S103. Specifically, before recording the second speech set, the recording format of the initial audio data set is determined, and the recording format may include the number of channels and / or the sampling rate. The recording format of the initial audio data set will also affect the recording effect of the initial audio data set, so the recording format can be set before recording. The recording format can be specifically determined by the specific scenario that the audio data needs to reflect. For example, in order to make the audio data reflect the scenario of telephone communication, the recording format of the initial audio data can be set to mono and the sampling rate can be 8kHz. In this way, the recording effect of the initial audio data set can be guaranteed.

[0040] In some implementations, a specific speaker may also record the second speech set in a preset recording environment. In this way, the initial audio data set includes the voice features of the specific speaker, and after the initial audio data is applied to the training process of the speech synthesis model, the speech synthesis model can also generate speech that matches the voice features of the specific speaker.

[0041] In some implementations, the format of the initial audio data set may also be set to a wav format.

[0042] In addition, the initial audio data set includes audio data corresponding to each speech utterance in the second speech set. In this way, the consistency between the audio data and the text content of the speech utterance can be guaranteed, so that when the text content of the speech utterance and the audio data are applied to the training process of the speech synthesis model, the training results are more accurate.

[0043] It is understandable that the initial audio data set is audio data obtained by the speaker reading the text content. Therefore, the initial audio data set has text content corresponding to it. In this embodiment, the text content is also called the second speech set.

[0044] S104: normalizing the initial audio data set and the preset public audio data set to obtain a target audio data set. The normalization process is used to adjust the amplitudes of the initial audio data set and the public audio data set to within a preset range.

[0045] The preset public data set can be a data set that can be obtained arbitrarily through public channels. Specifically, a certain proportion of a public data set can be selected, such as 30% of the standard data. The target audio data set can be determined by using the initial audio data set and the public data set together, which can enrich the content contained in the data set and increase the universality of the speech materials contained in the data set, so that the data set can have both the speech materials of the target business scenario and the universal public speech materials.

[0046] The embodiment of the present application also includes normalizing the initial audio data set and the public data set together to finally obtain the target audio data set. In this way, the amplitude ranges of the initial audio data set and the public data set are more concentrated, the target audio data set is more conducive to the training of the speech synthesis model, and the trained speech synthesis model has higher quality. Among them, the specific amplitude preset range can be determined according to the actual situation, and generally an amplitude range with good listening sense is selected.

[0047] From the above content, it can be seen that the audio generation method provided in the embodiment of the present application forms a first speech set based on the first speech material and the second speech material, and then performs a smoothness check on the first speech set to form a second speech set. The method also includes recording the second speech set to obtain an initial audio data set. The initial audio data set and the public data set are then normalized to obtain a target audio data set. In this way, the target audio data set includes audio corresponding to the target business scenario, and the recording quality is good, which can clearly correspond to the text content of the speech material in the second speech set, and can reflect the actual situation in the target business scenario with high accuracy.

[0048] Figure 2 A schematic diagram of a process for determining second speech material provided in an embodiment of the present application.

[0049] like Figure 2 As shown, in the embodiment of the present application, the following steps may be further included before step S101:

[0050] S201: Determine a plurality of first candidate sets, where the first candidate sets include a plurality of language elements corresponding to sentence components, where the sentence components of the language elements in different first candidate sets are different, and where the language elements are determined based on discourse material collected from a target business scenario.

[0051] The first candidate set is a set consisting of multiple language elements, and the language elements corresponding to each first candidate set are different, and the difference is reflected in the different sentence components corresponding to the language elements. The language elements may include subject, predicate and object, and may also include attributive, adverbial and complement.

[0052] For example, in the target business scenario, the following conversational data is collected: "This is xx Fitness Center". In this conversational data, "here" can be selected into the first candidate set corresponding to the subject, "xx" can be selected into the first candidate set corresponding to the attributive, and "fitness center" can be selected into the first candidate set corresponding to the object. For another example, "Mr. Zhang, do you like fitness?" In this conversational data, "Mr. Zhang" can be selected into the first candidate set corresponding to the subject, "like" can be selected into the first candidate set corresponding to the predicate, and "fitness" can be selected into the first candidate set corresponding to the object.

[0053] S202: Extracting language elements corresponding to one or more sentence components from multiple first candidate sets.

[0054] For example, "this product" can be extracted from the first candidate set corresponding to the subject, "can relieve" can be extracted from the first candidate set corresponding to the predicate, "muscle" can be extracted from the first candidate set corresponding to the attributive, "fatigue" can be extracted from the first candidate set corresponding to the object, etc. For another example, "I" can be extracted from the first candidate set corresponding to the subject, "give" can be extracted from the first candidate set corresponding to the predicate, "you" and "contact information" can be extracted from the first candidate set corresponding to the object, and "one" can be extracted from the first candidate set corresponding to the attributive.

[0055] It is understandable that, during the extraction process, multiple language elements can be extracted from the candidate set corresponding to the language element of a sentence component, and the extraction is not limited to only one.

[0056] S203: Combining the extracted language elements to obtain second discourse material.

[0057] Among them, the extracted language elements can be combined into sentences according to the order relationship between the various components that constitute the sentence to obtain the second discourse material. Specifically, a sentence pattern can be formed based on the order relationship between the various components of the sentence, and then the language elements can be combined according to the sentence pattern to obtain the second discourse material. For example, according to the sentence pattern of subject-predicate-attributive-object, the above-extracted language elements can be combined into "This product can relieve muscle fatigue." For another example, according to the sentence pattern of subject-predicate-object-attributive-object, the above-extracted language elements can be combined into "I will give you a contact method."

[0058] In the embodiment of the present application, the method of combining multiple language elements to form the second speech corpus is flexible, and a large number of second speech corpus can be obtained, and the second speech corpus can correspond to the actual communication situation of the target business scenario. Therefore, the first speech set determined based on the second speech corpus is rich in corpus and highly fits the target business scenario.

[0059] Figure 3 Another flowchart of determining second speech material is provided in an embodiment of the present application.

[0060] like Figure 3 As shown, in the embodiment of the present application, the following steps may be further included before step S101:

[0061] S301: Obtain a second candidate set, where the second candidate set includes a preset first speech sample and a plurality of second speech samples determined based on speech materials collected from business scenarios.

[0062] Among them, the first speech sample can be a fixed speech, for example, it can be some general speech, such as: "Hello", "Please wait", "You have been waiting for a long time", "Goodbye", "I wish you a happy life", etc. The second speech sample is determined based on the speech material collected from the business scenario, and the business scenario is the target business scenario. The second speech sample can be a variable speech, that is, the content in the second speech sample can be replaced. Specifically, the second speech sample can include multiple categories, and the category can be determined based on the target business scenario. For example, member business can be a category in the second speech sample, and the member business can include two contents: member type and member name. The text corresponding to the member type and member name can be replaced. For example, the text corresponding to the member type can be xxAPP member, xx store member, etc. collected in the marketing scenario, and the text corresponding to the member name can be "Mr. Wang", "Mr. Li", etc. determined based on the Hundred Family Names.

[0063] S302: Extract at least one first speech sample and at least one second speech sample from the second candidate set.

[0064] In the embodiment of the present application, the first speech sample and the second speech sample can be randomly selected from the second speech set. For example, the first speech sample selected from the second candidate set can be "Hello", and the second speech sample selected can be "xxAPP member" or "Mr. Wang".

[0065] S303: Combine the extracted at least one first speech sample and at least one second speech sample to obtain second speech material.

[0066] In the embodiment of the present application, the first speech sample and the second speech sample can be randomly combined. For example, the second speech material formed by combining the first speech sample and the second speech sample can be: "Hello, Mr. Wang, a member of xxAPP". The second speech material formed by combining the first speech sample and the second speech sample can also be: "Hello, Mr. Wang, a member of xxAPP".

[0067] In the embodiment of the present application, the method of combining the first speech sample and the second speech sample to form speech corpus can combine the general speech and the speech for the target business scenario, making the speech corpus richer and more practical. Therefore, the first speech set determined based on the second speech corpus is rich in corpus and highly consistent with the target business scenario.

[0068] Figure 4 A schematic diagram of the smoothness inspection process provided in an embodiment of the present application.

[0069] like Figure 4 As shown, in the embodiment of the present application, step S102 may include the following steps:

[0070] S1021: Input the speech language in the first speech set into a preset language model to obtain a probability value of each speech language, where the probability value is used to indicate the smoothness of the speech language.

[0071] The preset language model may be an n-gram model, which may be used to determine the probability value of the spoken language material. The numerical value of the probability value may reflect whether the spoken language material conforms to the language expression habits.

[0072] Specifically, the language model (such as the n-gram model) measures the degree to which a word sequence conforms to language expression habits through the probability P (word sequence) of the word sequence. The calculation formula of P (word sequence) is as follows:

[0073] P(word sequence word M )=P(word M |word1word2…word M-1 );

[0074] Wherein, M represents the Mth word sequence. It can be understood that the word sequence can be a piece of discourse material in the first discourse set, or the word sequence can be one or more words in a piece of discourse material. M |word1word2…word M-1 ) represents P(word sequence word M ) depends on the M-1 word sequences before the Mth word sequence.

[0075] In the embodiment of the present application, the n-gram model can be a 3-gram model, which means n=3. The value of n determines the value of M. Specifically, M-1=n. When n=3, P(word sequence word M ) depends on the 3 word sequences before the Mth word sequence.

[0076] S1022: Determine the linguistic corpus whose probability value is greater than a preset threshold as the target corpus.

[0077] The larger the probability value, the higher the semantic smoothness. Therefore, a preset threshold can be set for the probability value. Only when the probability value is greater than the preset threshold can the corpus be determined as the target corpus. The specific value of the preset threshold can be determined according to the actual situation. For example, the preset threshold can be 0.7.

[0078] By checking the language smoothness based on the preset language model, it is possible to prevent semantically unsmooth speech materials from entering the second speech set, thereby avoiding subsequent impact on the training effect of the language synthesis module.

[0079] Figure 5A schematic diagram of the process of randomly extracting the second speech set provided in an embodiment of the present application.

[0080] like Figure 5 As shown, before step S103, the method of the embodiment of the present application may further include the following steps:

[0081] S401: Randomly sampling the target corpus in the second speech set according to a preset sampling ratio.

[0082] Among them, the preset sampling ratio can be determined based on the total number of target corpora in the second speech set and the actual required sampling accuracy. For example, the preset sampling ratio can be 0.05. Then, when extracting, the number of extracted target corpora and the total number of target corpora in the second speech set account for 0.05.

[0083] S402: Determine whether the extracted target corpus contains a preset grammatical defect.

[0084] Among them, the preset grammatical defects may include improper word order, improper collocation, incomplete or redundant components, chaotic structure, unclear meaning, illogicality, inconsistency, etc. For example, the extracted target corpus may be "Due to the good performance of this product, it has received favorable comments from many consumers". In this target corpus, "get" lacks a subject, so the target corpus has a grammatical defect of incomplete components.

[0085] S403: If the extracted target corpus contains a preset grammatical defect, the extracted target corpus is removed from the second speech set, and S401 is executed again. If the extracted target corpus does not contain a preset grammatical defect, S404 is executed.

[0086] Removing target corpus with grammatical defects from the second speech set can prevent it from affecting the training effect of the speech synthesis model.

[0087] S404: Determine whether the target corpus extracted N times in succession does not contain a preset grammatical defect.

[0088] The number N can be determined by the total number of target corpora in the second speech set and the actual required spot check accuracy, and is preferably three times.

[0089] S405: If the target corpus extracted N times in succession does not contain the preset grammatical defect, then the extraction is terminated. If the target corpus extracted N times in succession contains the preset grammatical defect, then S401 is executed.

[0090] Among them, if the target corpus extracted N times in succession does not contain the preset grammatical defects, it can be said that the target corpus in the second speech set does not contain the preset grammatical defects. Such a second speech set can ensure the accuracy of the language synthesis module training.

[0091] In addition, during the random sampling inspection process, some target corpora with problems such as swallowing of sounds and failure to distinguish between the Chinese Pinyin initials nl can be screened out. At this time, these target corpora can also be removed from the second speech set.

[0092] It should be supplemented that, in the inspection process of the target corpus in the second speech set in the embodiment of the present application, a smoothness check may be performed first and then a random sampling check may be performed, or a random sampling check may be performed first and then a smoothness check may be performed, or one of the smoothness check and the random sampling check may be selected for inspection.

[0093] Figure 6 A schematic diagram of the normalization process provided in an embodiment of the present application.

[0094] like Figure 6 As shown, step S104 may include the following steps:

[0095] S1041: Convert the initial audio data set and the public data set into a matrix form to obtain a first matrix.

[0096] Among them, some programs for processing audio data can be used to convert the initial audio data set and the public data set, such as FFmpeg.

[0097] In some implementations, the initial audio data set and the public data set may be matrix-converted respectively, and then the converted matrices may be merged to form a first matrix. Alternatively, the initial audio data set and the public data set may be merged, and the merged data set may be matrix-converted to obtain a first matrix, which is not specifically limited in this application.

[0098] S1042: Determine a plurality of matrix parameters of the first matrix, where the matrix parameters include a median, an average and / or a mode of the first matrix.

[0099] Since the first matrix is ​​derived from the initial audio data set and the public data set, the parameters of the first matrix can accurately express the mathematical properties of the initial audio data set and the public data set. In addition, the matrix parameters can be one or more of the median, the average and the mode, which is not specifically limited in this application.

[0100] S1043: using each matrix parameter, respectively normalize the first part of the audio data to obtain a normalization result corresponding to each matrix parameter. The first part of the audio data includes part of the data in the initial audio data set and part of the data in the public data set.

[0101] The first part of the audio data may be part of the data in the initial audio data set, part of the data in the public data set, or a combination of the part of the data in the initial audio data set and part of the data in the public data set. Normalizing the first part of the audio data may reduce the time of the normalization process.

[0102] S1044: Based on the audition effects of the normalized results corresponding to the matrix parameters, determine the normalized parameters from the multiple matrix parameters, the normalized parameters being the matrix parameters corresponding to the optimal audition effect.

[0103] Multiple normalization results are used to determine the normalization degree of the first part of the audio data when normalization processing is performed separately using each matrix parameter. In this way, a normalized result with the best normalization degree can be determined from multiple normalization results. For example: the normalized results can be auditioned to obtain a normalized result with a better listening effect. The parameter corresponding to the normalized result with a better effect can then be determined as a normalization parameter. In this way, the normalization parameter can be applied to the normalization process of the initial audio data set and the public data set, so that the normalized results of the initial audio data set and the public data set can be made clearer in terms of listening.

[0104] S1045: Use normalization parameters to normalize the initial audio data set and the public data set to obtain a target audio data set.

[0105] In this embodiment, the min-max normalization method may be used for normalization processing. For example, if the normalization parameter is determined to be the mean value in step S1044, the mean normalization may be performed using the following formula.

[0106]

[0107] Among them, x is the normalization result, value is the current value in the data set, u is the normalization parameter, max is the maximum value of the data set, and min is the minimum value of the data set.

[0108] It should be noted that the process of normalizing the initial audio data set and the public data set can be referred to as the process of post-processing the audio data. After normalization, the amplitudes of the initial audio data set and the public data set are adjusted to within a preset range to obtain a target audio data set. The target audio data set can be used as a training corpus in the training process of the speech synthesis model, which can improve the training effect of the speech synthesis model.

[0109] In some implementations, before executing steps S1041-S1045 for normalization processing, the embodiment of the present application may further include other post-processing steps, such as adding a clipping process before step S1041. Specifically, a certain period of silent audio is added before and after each audio data in the initial audio data set. Exemplarily, the duration of the silent audio may be 500 milliseconds. In this way, the fault tolerance of the initial audio data set can be increased, and the connection between the various audio data can be facilitated.

[0110] In the process of editing the initial audio data set, microphone noise and the like can also be eliminated.

[0111] In addition, a text verification process can be added before step S1041. For example, the speech material in the second speech set can be compared with the audio data in the initial audio data set to determine whether the speech material and the audio data are consistent. If inconsistency occurs at this time, it can be modified based on the audio data dialogue material. In this way, when the speech material and audio data are applied to the speech synthesis model training process, the accuracy of the training can be guaranteed.

[0112] It is understandable that the above post-processing processes can all be applied to preset public data sets.

[0113] It should be noted that the target audio data set may also be in wav format.

[0114] Figure 7 A schematic diagram of the structure of the audio generation device provided in the embodiment of the present application, such as Figure 7 As shown, the device may include the following modules:

[0115] The acquisition module 501 is used to acquire a first speech set, the first speech set includes multiple speech literals, the speech literals include first speech literals and second speech literals, the first speech literals are collected from the target business scenario, and the second speech literals are composed of multiple language elements with different sentence components.

[0116] The smoothness checking module 502 is used to check the smoothness of the speech corpus in the first speech set based on a preset language model, determine the target corpus, and generate a second speech set including the target corpus, wherein the smoothness of the target corpus is greater than a preset threshold.

[0117] The recording module 503 is used to record the second speech set in a preset recording environment to obtain an initial audio data set, wherein the initial audio data set includes audio data corresponding to each speech material in the second speech set. The control parameters of the recording environment include noise level and / or reverberation level.

[0118] The normalization module 504 is used to perform normalization processing on the initial audio data set and the preset public data set to obtain the target audio data set. The normalization processing is used to adjust the amplitude of the initial audio data set and the public data set to within a preset range.

[0119] In some embodiments, the normalization module 504 is specifically used to convert the initial audio data set and the public data set into a matrix form to obtain a first matrix; determine multiple matrix parameters of the first matrix, the matrix parameters include the median, mean and / or mode of the first matrix; use each matrix parameter to normalize the first part of the audio data respectively to obtain a normalized result corresponding to each matrix parameter; the first part of the audio data includes part of the data in the initial audio data set and part of the data in the public data set; based on the audition effect of the normalization result corresponding to each matrix parameter, determine the normalization parameter from the multiple matrix parameters, the normalization parameter is the matrix parameter corresponding to the optimal audition effect; use the normalization parameter to normalize the initial audio data set and the public data set to obtain the target audio data set.

[0120] In some embodiments, the smoothness check module 502 is specifically used to input the speech corpus in the first speech set into a preset language model to obtain the probability value of each speech corpus, and the probability value is used to indicate the smoothness of the speech corpus; the speech corpus with a probability value greater than a preset threshold is determined as the target corpus.

[0121] In some embodiments, the device also includes a spot check module, which is specifically used to randomly extract the target corpus from the second speech set according to a preset spot check ratio; if the extracted target corpus contains preset grammatical defects, the extracted target corpus is removed from the second speech set and randomly extracted again; if the target corpus extracted for N consecutive times does not contain the preset grammatical defects, the extraction is terminated.

[0122] In some embodiments, the acquisition module 501 is specifically used to determine multiple first candidate sets, the first candidate sets include multiple language elements corresponding to sentence components, the sentence components of the language elements in different first candidate sets are different, and the language elements are determined based on the discourse material collected from the target business scenario; extract one or more language elements corresponding to the sentence components from the multiple first candidate sets; and combine the extracted language elements to obtain second discourse material.

[0123] In some embodiments, the acquisition module 501 is specifically used to obtain a second candidate set, which includes a plurality of pre-set first speech samples and a plurality of second speech samples determined based on speech materials collected from business scenarios; extracting at least one first speech sample and at least one second speech sample from the second candidate set; combining the extracted at least one first speech sample and at least one second speech sample to obtain the second speech material.

[0124] In some embodiments, the device also includes a recording format determination module, which is specifically used to determine the recording format of the initial audio data set before recording the second speech set, and the recording format includes the number of channels and / or sampling rate.

[0125] In some embodiments, the preset language model is an n-gram model.

[0126] In some embodiments, the recording module 503 is specifically used to record the second speech set by a specific speaker in a preset recording environment, and the control parameters of the recording environment include noise level and / or reverberation level.

[0127] From the above content, it can be seen that the audio generation device provided in the embodiment of the present application forms a first speech set based on the first speech material and the second speech material, and then performs a smoothness check on the first speech set to form a second speech set. The method also includes recording the second speech set to obtain an initial audio data set. The initial audio data set and the public data set are then normalized to obtain a target audio data set. In this way, the target audio data set includes audio corresponding to the target business scenario, and the recording quality is good, which can clearly correspond to the text content of the speech material in the second speech set, and can reflect the actual situation in the target business scenario with high accuracy.

[0128] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment of the method for determining a network resource reuse area provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0129] It is easy to understand that those skilled in the art can combine, split, reorganize, etc. the embodiments of the present application to obtain other embodiments based on the several embodiments provided in the present application, and these embodiments do not exceed the protection scope of the present application.

[0130] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above are only specific implementation methods of the embodiments of the present application and are not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the protection scope of the embodiments of the present application.

Claims

1. An audio generation method, characterized in that: The method comprises: Acquire a first speech set, the first speech set includes a plurality of speech literals, the speech literals include a first speech literal and a second speech literal, the first speech literal is collected from a target business scenario, and the second speech literal is composed of a plurality of language elements with different sentence components; Performing a smoothness check on the speech corpus in the first speech set based on a preset language model, determining a target corpus, and generating a second speech set including the target corpus, wherein the smoothness of the target corpus is greater than a preset threshold; Recording the second speech set in a preset recording environment to obtain an initial audio data set, wherein the initial audio data set includes audio data corresponding to each speech material in the second speech set; The initial audio data set and the preset public data set are normalized to obtain a target audio data set; the normalization is used to adjust the amplitudes of the initial audio data set and the public data set to within a preset range.

2. The audio generation method according to claim 1, characterized in that: The step of normalizing the initial audio data set and the preset public data set to obtain a target audio data set includes: Converting the initial audio data set and the public data set into a matrix form to obtain a first matrix; determining a plurality of matrix parameters of the first matrix, the matrix parameters comprising a median, a mean and / or a mode of the first matrix; Using each of the matrix parameters, respectively normalizing the first part of the audio data to obtain a normalized result corresponding to each of the matrix parameters; the first part of the audio data includes part of the data in the initial audio data set and part of the data in the public data set; Based on the audition effects of the normalized results corresponding to the matrix parameters, determine a normalized parameter from the plurality of matrix parameters, the normalized parameter being the matrix parameter corresponding to the optimal audition effect; The initial audio data set and the public data set are normalized using the normalization parameters to obtain the target audio data set.

3. The audio generation method according to claim 1, characterized in that: The step of checking the fluency of the speech corpus in the first speech set based on a preset language model, determining a target corpus, and generating a second speech set including the target corpus includes: Inputting the speech language in the first speech language set into a preset language model to obtain a probability value of each of the speech language, wherein the probability value is used to indicate the smoothness of the speech language; The discourse corpus whose probability value is greater than a preset threshold is determined as the target corpus.

4. The audio generation method according to claim 1, characterized in that: Also includes: Randomly sampling the target corpus in the second speech set according to a preset sampling ratio; If the extracted target corpus contains a preset grammatical defect, the extracted target corpus is removed from the second speech set, and random extraction is performed again; If the target corpus extracted N times in succession does not contain the preset grammatical defects, the extraction is terminated.

5. The audio generation method according to claim 1, characterized in that: The second language material is obtained by the following steps: Determine a plurality of first candidate sets, wherein the first candidate sets include a plurality of the language elements corresponding to sentence components, the language elements in different first candidate sets have different sentence components, and the language elements are determined based on discourse material collected from a target business scenario; Extracting one or more language elements corresponding to sentence components from the first candidate sets; The extracted language elements are combined to obtain the second discourse material.

6. The audio generation method according to claim 1, characterized in that: The second language material is obtained by the following steps: Acquire a second candidate set, where the second candidate set includes a plurality of preset first speech samples and a plurality of second speech samples determined based on speech materials collected from business scenarios; Extract at least one of the first speech samples and at least one of the second speech samples from the second candidate set; The extracted at least one first speech sample and at least one second speech sample are combined to obtain the second speech material.

7. The audio generation method according to claim 1, characterized in that: Also includes: Before recording the second speech set, a recording format of the initial audio data set is determined, where the recording format includes the number of channels and / or the sampling rate.

8. The audio generation method according to claim 1, characterized in that: The preset language model is an n-gram model.

9. The audio generation method according to claim 1, characterized in that: The step of recording the second speech set in a preset recording environment includes: The second speech set is recorded by a specific speaker in a preset recording environment, and the control parameters of the recording environment include noise level and / or reverberation level.

10. An audio generating device, characterized in that: The device comprises: An acquisition module, configured to acquire a first speech set, wherein the first speech set includes a plurality of speech literals, wherein the speech literals include a first speech literal and a second speech literal, wherein the first speech literal is collected from a target business scenario, and the second speech literal is composed of a plurality of language elements having different sentence components; A smoothness checking module, configured to check the smoothness of the speech corpus in the first speech set based on a preset language model, determine a target corpus, and generate a second speech set including the target corpus, wherein the smoothness of the target corpus is greater than a preset threshold; A recording module, configured to record the second speech set in a preset recording environment to obtain an initial audio data set, wherein the initial audio data set includes audio data corresponding to each speech material in the second speech set; A normalization module is used to normalize the initial audio data set and the preset public data set to obtain a target audio data set; the normalization process is used to adjust the amplitudes of the initial audio data set and the public data set to within a preset range.

Citation Information

Patent Citations

  • Method and device for dynamically loading language model based on user input scene

    CN103577386A

  • Speech synthesis method, electronic device and storage medium

    CN110534088A