Method and device for obtaining audio and video file size
By extracting the rhythm and phoneme feature of the target text, predicting the size of the speech synthesis audio file, solving the problem that the audio file size cannot be obtained in advance in the prior art, and achieving high accuracy and timely prediction.
Patent Information
- Application Number
- CN202210346097.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Existing voice synthesis technology cannot obtain the size of the audio file before playing, resulting in users only knowing the file size after playing, which has lag.
By extracting the pronunciation features and phoneme features of the target text, the size of the target audio file is predicted based on these features, and the file size prediction model is used for accurate prediction.
It realizes that the file size can be predicted before the target audio file is generated, which is timely and has high accuracy and accuracy of prediction results.
Smart Images

Figure CN114708848B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech synthesis technology, and in particular to a method and device for obtaining the size of an audio or video file. Background Art
[0002] Speech synthesis technology is widely used in all aspects of daily life. The most common speech synthesis method at present is streaming speech synthesis. However, this method can only inform the user of the corresponding size of the audio file after the entire audio file is completely synthesized and played. The size of the audio file cannot be obtained before playback, which has a certain lag. Summary of the invention
[0003] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method for obtaining the size of an audio or video file.
[0004] The present application also proposes a device for obtaining the size of an audio or video file.
[0005] The application also provides an electronic device.
[0006] The present application also proposes a non-transitory computer-readable storage medium.
[0007] The present application also proposes a computer program product.
[0008] According to the first aspect of the present application, the method for obtaining the size of an audio or video file includes:
[0009] Get the target text;
[0010] Extracting features from the target text to generate target prosodic features and target phoneme features;
[0011] Based on the target prosodic feature and the target phoneme feature, a target file size of a target audio file is acquired, where the target audio file is generated by performing speech synthesis on the target text.
[0012] According to the method for obtaining the size of audio and video files in the embodiment of the present application, by extracting the rhythmic features and phoneme features of the target text, and predicting the size information of the target audio file synthesized by the target text based on the extracted target rhythmic features and target phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction result are high.
[0013] According to an embodiment of the present application, the step of acquiring a target file size of a target audio file based on the target prosodic feature and the target phoneme feature includes:
[0014] Based on the target prosodic feature and the target phoneme feature, obtaining a first predicted file size of the target audio file;
[0015] The first predicted file size and the target residual value are summed to generate the target file size, wherein the target residual value is determined based on the sample file size and the size of the sample audio file corresponding to the predicted sample text, and the sample file size is the actual size of the sample audio file corresponding to the sample text.
[0016] According to one embodiment of the present application, the target residual value is determined by the following steps:
[0017] Obtaining a sample text, a sample audio file corresponding to the sample text, and a sample file size corresponding to the sample audio file, wherein the sample audio file is generated by performing speech synthesis on the sample text;
[0018] Extracting features from the sample text to generate sample prosodic features and sample phoneme features;
[0019] Based on the sample prosodic feature and the sample phoneme feature, obtaining a second predicted file size of the sample audio file;
[0020] The maximum absolute value of the difference between the second predicted file size and the sample file size is determined as the target residual value.
[0021] According to an embodiment of the present application, obtaining a first predicted file size of the target audio file based on the target prosody feature and the target phoneme feature includes:
[0022] Inputting the target prosodic feature and the target phoneme feature into a file size prediction model, and obtaining the first predicted file size output by the file size prediction model; wherein,
[0023] The file size prediction model is obtained by training with sample prosodic features and sample phoneme features as samples and sample file sizes corresponding to the sample prosodic features and the sample phoneme features as sample labels.
[0024] According to an embodiment of the present application, after obtaining the target file size of the target audio file, the method further includes:
[0025] Segmenting the target text based on the target prosodic features and phoneme features to generate multiple sentence sequences;
[0026] Performing speech synthesis on the sentence sequence to generate sentence speech;
[0027] The sentence speech and the target file size are output, and the sentence speech is concatenated to generate the target audio file.
[0028] According to an embodiment of the present application, the step of extracting features from the target text to generate target prosodic features and target phoneme features includes:
[0029] Converting the target text into a prosodic phoneme sequence, wherein the prosodic phoneme sequence includes a plurality of phonemes corresponding to the target text and prosodic identifiers located between adjacent phonemes;
[0030] Feature extraction is performed on the prosodic phoneme sequence to generate the target prosodic feature and the target phoneme feature.
[0031] According to one embodiment of the present application, the target prosodic features and phonemic features include: the length of the prosodic phoneme sequence, the number of Chinese pinyins in the prosodic phoneme sequence, the number of pause symbols in the prosodic phoneme sequence, the number of English phonemes in the prosodic phoneme sequence, the number of Chinese phonemes in the prosodic phoneme sequence, the number of Chinese initials in the prosodic phoneme sequence, the number of Chinese finals in the prosodic phoneme sequence, and at least one of the various categories of English phonemes in the prosodic phoneme sequence.
[0032] According to the second aspect of the present application, the device for obtaining the size of an audio or video file includes:
[0033] A first processing module, used for acquiring a target text;
[0034] A second processing module is used to extract features from the target text to generate target prosodic features and target phoneme features;
[0035] The third processing module is used to obtain a target file size of a target audio file based on the target prosodic feature and the target phoneme feature, where the target audio file is generated by performing speech synthesis on the target text.
[0036] According to the device for obtaining the size of audio and video files in the embodiment of the present application, by extracting the rhythmic features and phoneme features of the target text, and predicting the size information of the target audio file synthesized by the target text based on the extracted target rhythmic features and phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction result are high.
[0037] According to an electronic device of an embodiment of the third aspect of the present application, the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for obtaining the size of an audio or video file as described above is implemented.
[0038] According to the non-transitory computer-readable storage medium of the fourth aspect embodiment of the present application, a computer program is stored thereon, and when the computer program is executed by a processor, the method for obtaining the size of an audio or video file as described in any one of the above is implemented.
[0039] According to the computer program product of the fifth aspect of the present application, the computer program includes a computer program, which, when executed by a processor, implements any of the methods for obtaining the size of an audio or video file as described above.
[0040] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0041] By extracting the prosodic features and phoneme features of the target text, and predicting the size information of the target audio file synthesized from the target text based on the extracted target prosodic features and phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction results are high.
[0042] Furthermore, by converting the target text into a phoneme sequence and marking the phoneme sequence based on prosodic identifiers corresponding to at least two of the sentence-end information, intonation phrases, prosodic phrases, prosodic words and syllables to generate a prosodic phoneme sequence, a more refined prosodic representation can be provided, thereby facilitating the segmentation sophistication and accuracy in the subsequent segmentation process.
[0043] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 This is one of the flow charts of the method for obtaining the size of an audio or video file provided in an embodiment of the present application;
[0046] Figure 2This is the second flow chart of the method for obtaining the size of an audio or video file provided in an embodiment of the present application;
[0047] Figure 3 It is a structural schematic diagram of a device for obtaining the size of an audio or video file provided in an embodiment of the present application;
[0048] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following is a further detailed description of the implementation of the present application in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but cannot be used to limit the scope of the present application.
[0050] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0051] Combine the following Figure 1-Figure 2 A method for obtaining the size of an audio or video file according to an embodiment of the present application is described.
[0052] The executor of the method for obtaining the size of audio and video files may be an audio and video file size obtaining device, or may be a server, or may be a user's terminal, including but not limited to mobile phones, tablet computers, PCs, vehicle-mounted terminals, and household smart appliances.
[0053] like Figure 1 As shown, the method for obtaining the size of the audio and video file includes: step 110, step 120 and step 130.
[0054] Step 110: Obtain target text;
[0055] In this step, the target text is the text currently used for speech synthesis.
[0056] The target text may be a regular text of tens to hundreds of levels, or an ultra-long text of thousands or tens of thousands of levels.
[0057] The target text can be a local file stored in a database, or it can also be a file downloaded from the network. This application does not make any limitations.
[0058] Step 120: Extract features from the target text to generate target prosody features and target phoneme features;
[0059] In this step, the target prosody features are used to represent the prosody features of the target text, and the target phoneme features are used to represent the phoneme features of the target text.
[0060] Among them, the target prosody features and target phoneme features include, but are not limited to: phonemes and their corresponding tones, syllables, prosodic words, prosodic phrases, intonation phrases, silences, and pauses, etc.
[0061] A syllable is a speech unit in the speech flow and is also the speech unit that is most easily distinguishable by people's hearing. For example, a syllable can be each Chinese character in the target text.
[0062] A prosodic word is a group of syllables that are closely related and pronounced together in the actual speech flow.
[0063] A prosodic phrase is a medium-sized rhythmic chunk between a prosodic word and an intonation phrase. A prosodic phrase can include multiple prosodic words and function words, and the multiple prosodic words that make up the prosodic phrase sound like they share a single rhythm group.
[0064] An intonation phrase is a sentence formed by connecting multiple prosodic phrases according to a certain intonation pattern, and is used to represent a larger pause.
[0065] The end-of-sentence information is used to represent the end of each long sentence.
[0066] For example, for the target text "Shanghai is cloudy turning to sunny today, with southeast winds of level 3 to 4", each Chinese character such as "上", "海", and "市" is the syllable corresponding to this target text; words or phrases composed of words such as "上海市", "今天", and "阴转多云" are the prosodic phrases corresponding to this target text; and the sentence "上海市今天阴转多云" formed by the prosodic phrases "上海市", "今天", and "阴转多云" is the intonation phrase corresponding to this target text.
[0067] In some embodiments, step 120 may include:
[0068] Convert the target text into a prosody phoneme sequence, where the prosody phoneme sequence includes prosody identifiers located between adjacent phonemes and multiple phonemes corresponding to the target text;
[0069] Extract the features of the prosody phoneme sequence to generate target prosody features and phoneme features.
[0070] In this embodiment, the prosodic phoneme sequence is a sequence of prosodic features and phoneme features used to characterize the target text.
[0071] The prosodic phoneme sequence includes prosodic features and phoneme identifiers located between adjacent phonemes and a plurality of phonemes corresponding to the target text.
[0072] Among them, a phoneme can be a combination of one or more phonetic units divided according to the natural attributes of speech, and a phonetic unit can be the pinyin, initial consonant or final vowel corresponding to a Chinese character, or an English word, English phonetic symbol or English letter.
[0073] The prosodic identifier is an identifier used to characterize the prosodic features corresponding to each phoneme in the target text. The prosodic features include but are not limited to: tones, syllables, prosodic words, prosodic phrases, intonation phrases, silence, and pauses corresponding to the phonemes.
[0074] Among them, the fine-grainedness of the prosodic identifier used to characterize pauses is higher than the fine-grainedness of the prosodic identifier used to characterize intonation phrases, the fine-grainedness used to characterize intonation phrases is higher than the fine-grainedness used to characterize prosodic phrases, the fine-grainedness used to characterize prosodic phrases is higher than the fine-grainedness used to characterize prosodic words, and the fine-grainedness used to characterize prosodic words is higher than the fine-grainedness used to characterize syllables.
[0075] In actual implementation, different symbols can be used to represent prosodic features of different granularity levels.
[0076] For example, for the target text "Shanghai is cloudy to overcast with southeast wind level three to four today", it can be converted into a prosodic phoneme sequence: sil shang4#0hai3#0shi4#2jin1#0tian1#2yin1#0zhuan3#1duo1#0yun2#3dong1#0nan2#0feng1#2san1#0dao4#1si4#0ji2#4sil.
[0077] It can be understood that, for the prosodic phoneme sequence, the prosodic identifier may include: numbers, symbols and English character strings between adjacent phonemes; the phonemes may include the pinyin corresponding to each Chinese character.
[0078] Among them, sil in the prosodic phoneme sequence represents the silence at the beginning and end of a sentence, #0 represents a syllable, #1 represents a prosodic word, #2 represents a prosodic phrase, #3 represents an intonation phrase, and #4 represents the end of a sentence. The number after each phoneme represents the tone of the phoneme, such as shang4, which means that the tone of the pinyin "shang" is the fourth tone.
[0079] In some embodiments, converting the target text into a prosodic phoneme sequence may include:
[0080] Convert the target text into a phoneme sequence;
[0081] Obtain the end-of-sentence information, intonation phrases, prosodic phrases, prosodic words, and syllables of the phoneme sequence;
[0082] Based on at least two of the end-of-sentence information, intonation phrases, prosodic phrases, prosodic words, and syllables, mark the phoneme sequence to generate a prosodic phoneme sequence.
[0083] In this embodiment, a syllable is a speech unit in the speech flow and is also the speech unit that is most easily distinguishable by people's hearing. For example, a syllable can be each Chinese character in the target text.
[0084] A prosodic word is a group of syllables that are closely related and pronounced together in the actual speech flow.
[0085] A prosodic phrase is a medium-rhythm chunk between a prosodic word and an intonation phrase. A prosodic phrase can include multiple prosodic words and function words, and the multiple prosodic words that make up the prosodic phrase sound like they share a rhythm group.
[0086] An intonation phrase is a sentence formed by connecting multiple prosodic phrases according to a certain intonation pattern, and is used to represent a larger pause.
[0087] The end-of-sentence information is used to represent the end of each long sentence.
[0088] For example, for the target text "It is cloudy turning to overcast in Shanghai today, with southeast winds of force 3 to 4", each Chinese character such as "上", "海", and "市" is a syllable corresponding to this target text; words or phrases composed of words such as "上海市", "今天", and "阴转多云" are prosodic phrases corresponding to this target text; and the sentence "上海市今天阴转多云" composed of the prosodic phrases "上海市", "今天", and "阴转多云" is the intonation phrase corresponding to this target text.
[0089] After obtaining the end-of-sentence information, intonation phrases, prosodic phrases, prosodic words, and syllables and other information of the target text, mark the target text based on at least two of them to generate a prosodic sequence.
[0090] The applicant found during the R & D process that in related technologies, the prosody of a sentence is often represented by using punctuation marks in the sentence. For example, the sentence is segmented at the position of a comma or a full stop in the sentence to obtain multiple clauses. On the one hand, this method cannot meet the segmentation of text without punctuation marks. On the other hand, it will also cause the two ends after segmentation to be unbalanced, and the segmentation effect is not good.
[0091] In the present application, at least two items of sentence-end information, intonation phrases, prosodic phrases, prosodic words and syllables are used to represent the rhythm of a sentence, and the target text is segmented based on this. The text will not be cut off in the middle of a whole word, so that the pauses and rhythm of the sentences obtained after segmentation are more natural.
[0092] The phoneme sequence is a sequence formed by connecting the phonemes (such as pinyin, tone or phonetic symbol) corresponding to each syllable in the target text.
[0093] For example, for the target text "Shanghai is cloudy to overcast with southeast wind level three to four today", it can be converted into a phoneme sequence: shang4 hai3 shi4 jin1 tian1 yin1 zhuan3 duo1 yun2 dong1 nan2 feng1san1dao4 si4 ji2.
[0094] The prosodic identifier is an identifier used to represent the prosodic feature corresponding to each phoneme in the target text, that is, the prosodic identifier is a symbol used to represent sentence-end information, intonation phrases, prosodic phrases, prosodic words, and syllables.
[0095] In the actual implementation process, the prosody identifier can be represented by a combination of special symbols and numbers or a specific letter combination, for example, "#0", "#1", "#2", "#3" and "#4" are used to represent the prosody identifier respectively, and different combinations represent different granularity levels.
[0096] For example, #0 represents a syllable, #1 represents a prosodic word, #2 represents a prosodic phrase, #3 represents an intonation phrase, and #4 represents the end of a sentence. In this embodiment, the granularity is from small to large as follows: #0<#1<#2<#3<#4.
[0097] After obtaining the phoneme sequence and prosodic identifier corresponding to the target text, the prosodic identifier is inserted into the corresponding position in the phoneme sequence, such as inserting the prosodic identifier #0 used to represent the syllable after the pinyin corresponding to each syllable in the phoneme sequence, and inserting the prosodic identifier #2 used to represent the prosodic phrase after each prosodic phrase in the phoneme sequence, thereby converting the phoneme sequence into a prosodic phoneme sequence.
[0098] For example, the phoneme sequence "shang4 hai3shi4jin1 tian1 yin1 zhuan3 duo1 yun2 dong1 nan2 feng1 san1 dao4 si4 ji2" is marked with #0", "#1", "#2", "#3" and "#4" respectively, so as to generate the prosodic phoneme sequence: sil shang4#0hai3#0shi4#2jin1#0tian1#2yin1#0zhuan3#1duo1#0yun2#3dong1#0nan2#0feng1#2san1#0dao4#1si4#0ji2#4sil.
[0099] Among them, sil represents the silence at the beginning and end of a sentence.
[0100] In this embodiment, by converting the target text into a phoneme sequence and marking the phoneme sequence based on prosodic identifiers corresponding to at least two of sentence-end information, intonation phrases, prosodic phrases, prosodic words and syllables to generate a prosodic phoneme sequence, a more refined prosodic representation can be provided, thereby facilitating the segmentation sophistication and accuracy in the subsequent segmentation process.
[0101] After obtaining the prosodic phoneme sequence, the prosodic features and phoneme features in the prosodic phoneme sequence are extracted to generate target prosodic features and target phoneme features.
[0102] In some embodiments, the target prosodic features and target phoneme features may include: the length of the prosodic phoneme sequence, the number of Chinese pinyins in the prosodic phoneme sequence, the number of pause symbols in the prosodic phoneme sequence, the number of English phonemes in the prosodic phoneme sequence, the number of Chinese phonemes in the prosodic phoneme sequence, the number of Chinese initials in the prosodic phoneme sequence, the number of Chinese finals in the prosodic phoneme sequence, and at least one of the various categories of English phonemes in the prosodic phoneme sequence.
[0103] The length of the prosodic phoneme sequence may be the number of phonemes in the prosodic phoneme sequence.
[0104] Step 130: Obtain a target file size of a target audio file based on the target prosodic feature and the target phoneme feature.
[0105] In this step, the target audio file is an audio file generated by performing speech synthesis on the entire target text.
[0106] It can be understood that, for an audio file, the target audio file is the audio file; for a video file, the target audio file is the audio file included in the video file.
[0107] The target file size is the predicted file size of the target audio file.
[0108] The target file size may be file volume information, or may be voice length information of the third voice information, which is not limited in this application.
[0109] In some embodiments, step 130 may include:
[0110] Based on the target prosodic feature and the target phoneme feature, obtaining a first predicted file size of the target audio file;
[0111] The target residual value and the first predicted file size are summed to generate the target file size.
[0112] In this embodiment, the first predicted file size is an initial file size value of the speech synthesized from the target text that is predicted based on the target prosodic features and the target phoneme features and is not corrected.
[0113] The target residual value is used to correct the first predicted file size to improve the accuracy of the target file size finally generated.
[0114] The target residual value is determined based on the sample file size and the size of the sample audio file corresponding to the predicted sample text, where the sample file size is the actual size of the sample audio file corresponding to the sample text.
[0115] The target file size is the file size value of the speech synthesized from the target text after prediction based on the target prosodic features and the target phoneme features and correction. It can be understood that the accuracy of the target file size is higher than the first predicted file size.
[0116] The target residual value is a predetermined value, for example, the target residual value may be the maximum absolute value of the residual value.
[0117] In this embodiment, the first predicted file size is corrected by performing supplementary residual processing on the first predicted file size, thereby improving the accuracy of the target file size finally generated.
[0118] In the actual implementation process, a neural network model can be used to predict the first predicted file size.
[0119] The following uses a neural network model as an example of a file size prediction model to illustrate the method for generating the first predicted file size in this embodiment.
[0120] In some embodiments, step 130 may include:
[0121] The target prosodic feature and the target phoneme feature are input into a file size prediction model, and a first predicted file size output by the file size prediction model is obtained.
[0122] In this embodiment, the file size prediction model may be a pre-trained neural network model.
[0123] The file size prediction model is used to predict the file size value of speech synthesized from the text based on the prosodic features and phoneme features of the text.
[0124] The training process of the file size prediction model is as follows: taking sample prosodic features and sample phoneme features as samples, and taking sample file sizes corresponding to the sample prosodic features and sample phoneme features as sample labels, the file size prediction model is trained.
[0125] Among them, the sample prosodic features and sample phoneme features are generated by extracting the prosodic features and phoneme features of the sample text. The extraction method of the sample prosodic features and sample phoneme features is similar to the extraction method of the target prosodic features and target phoneme features mentioned above, and will not be repeated here.
[0126] The sample file size corresponding to the sample prosodic feature and the sample phoneme feature is the actual size value of the sample audio file generated by performing speech synthesis on the sample text.
[0127] In actual application, the target prosodic features and target phoneme features are input into the trained file size prediction model, and the file size prediction model can output the initial file size value corresponding to the speech generated by speech synthesis of the target text corresponding to the target prosodic features and target phoneme features, that is, the first predicted file size.
[0128] After obtaining the first predicted file size, the sum of the first predicted file size and the target residual value is calculated to generate the target file size.
[0129] In this embodiment, by adopting a pre-trained model to obtain the first predicted file size, the computational efficiency in the actual application process can be improved.
[0130] In addition, the target prosodic features and target phoneme features corresponding to each target text in the actual application process can be used as training samples for subsequent training of the file size prediction model. As the volume of training samples increases, the intelligence of the file size prediction model will continue to improve, and the final predicted results will be more accurate.
[0131] The following describes a method for determining the target residual value through a specific embodiment.
[0132] In some embodiments, the target residual value is determined by the following steps:
[0133] Obtaining a sample text, a sample file size corresponding to a sample audio file, and a sample audio file corresponding to the sample text, wherein the sample audio file is generated by performing speech synthesis on the sample text;
[0134] Extract features from the sample text to generate sample prosodic features and sample phoneme features;
[0135] Based on the sample prosodic feature and the sample phoneme feature, obtaining a second predicted file size of the sample audio file;
[0136] The maximum absolute value of the difference between the second predicted file size and the sample file size is determined as the target residual value.
[0137] In this embodiment, the sample text may be a regular text of tens to hundreds of levels, or an ultra-long text of thousands or tens of thousands of levels.
[0138] The sample audio file is the audio file finally generated by performing speech synthesis on the sample text.
[0139] The sample file size is the actual size value of the sample audio file or the actual audio duration.
[0140] For example, a speech synthesis system may be used to calculate the actual wav file size or audio duration of a sample audio file corresponding to the sample text.
[0141] The second predicted file size is the predicted, uncorrected size value or audio duration of the sample audio file.
[0142] It should be noted that the method for generating the second predicted file size should be consistent with the method for generating the first predicted file size.
[0143] In the actual execution process, feature extraction can be performed on the sample text to generate sample prosodic features and sample phoneme features, and the sample prosodic features and sample phoneme features can be input into the file size prediction model to obtain the second predicted file size output by the file size prediction model.
[0144] Then, the maximum absolute value of the difference between the second predicted file size and the sample file size is calculated as the target residual value.
[0145] It is understandable that during the execution process, the sample prosodic features and the sample phoneme features may be predicted multiple times to obtain multiple second predicted file sizes. The difference between each second predicted file size and the sample file size is calculated to obtain multiple candidate difference values; then the absolute value of the minimum non-positive value is selected from the multiple candidate difference values and determined as the target residual value to improve the accuracy of the target residual value.
[0146] According to the method for obtaining the size of audio and video files provided in the embodiment of the present application, by extracting the rhythmic features and phoneme features of the target text, and predicting the size information of the target audio file synthesized by the target text based on the extracted target rhythmic features and target phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction result are high.
[0147] like Figure 2 As shown, according to some embodiments of the present application, after step 130, the method may further include:
[0148] Segment the target text based on the target prosodic features and the target phoneme features to generate multiple sentence sequences;
[0149] Perform speech synthesis on the sentence sequence to generate sentence speech;
[0150] Output the sentence-by-sentence speech and target file size, and concatenate the sentence-by-sentence speech to generate the target audio file.
[0151] In this embodiment, each sentence sequence includes at least one phoneme, wherein the phoneme can be a Chinese phoneme or an English phoneme.
[0152] The target text is segmented based on at least one feature of syllables, prosodic words, prosodic phrases, and intonation phrases in the target prosodic features to obtain at least two sentence sequences.
[0153] For example, for the target text "Shanghai will be cloudy and turn to overcast with southeasterly winds of level 3 to 4 today", it can be first converted into a prosodic phoneme sequence: sil shang4#0hai3#0shi4#2jin1#0tian1#2yin1#0zhuan3#1duo1#0yun2#3dong1#0nan2#0feng1#2san1#0dao4#1si4#0ji2#4sil;
[0154] Then, segmentation is performed at #3, so that the rhythmic phoneme sequence can be segmented into the following multiple sentence sequences:
[0155] Clause sequence 1: sil shang4#0hai3#0shi4#2jin1#0tian1#2yin1#0zhuan3#1duo1#0yun2#3;
[0156] Clause sequence 2: dong1#0nan2#0feng1#2san1#0dao4#1si4#0ji2#4sil.
[0157] Performing speech synthesis on the sentence sequence that is at the front of the segmentation order among the multiple sentence sequences to generate sentence speech corresponding to the sentence sequence;
[0158] Output the sentence speech and target file size corresponding to the sentence sequence, and synthesize the subsequent sentence sequence.
[0159] For example, for the sample text: Please search for the details on the APP, which can be converted into a sample prosodic phoneme sequence: sil xiang2#0xi4#1nei4#0rong2#2ma2#0fan5#2zai4#1AE1 P#0shang4#1sou1#0xun2#0xia4#4sil;
[0160] Then, features are extracted from the sample prosodic phoneme sequence, and the extracted sample prosodic features and sample phoneme features include but are not limited to: the length of the sample prosodic phoneme sequence; the number of Chinese pinyin occurrences in the sample prosodic phoneme sequence, the number of pause symbols (#0#1#2#3sil) in the sample prosodic phoneme sequence, the number of English phonemes in the sample prosodic phoneme sequence, the number of Chinese phonemes in the sample prosodic phoneme sequence, the number of Chinese initials in the sample prosodic phoneme sequence, the number of Chinese finals in the sample prosodic phoneme sequence, and the number of each category of English phonemes (Vowels, Diphthongs, R colored vowels, Stops, Affricates, Fricatives, Nasals, Liquids, Semivowels) in the sample prosodic phoneme sequence.
[0161] After the training data is prepared, the WAV file size prediction model based on the ElasticNet regression model can be trained.
[0162] The sample prosodic features and sample phoneme features obtained above are input into the wav file size prediction model, and the target output of the training process is the actual wav file byte number of the sample audio file.
[0163] Specifically, cross-validation can be used to select the best performing model parameters, and then the ElasticNet regression model can be trained with the selected parameters.
[0164] Then, the target residual value is calculated, such as using the sample prosodic features and the sample phoneme features as inputs of the model to obtain a second predicted file size.
[0165] The absolute value of the minimum non-positive value of the second prediction file size minus the sample file size is calculated as the maximum residual value.
[0166] In actual application, the client initiates a request, such as obtaining the target text: Shanghai will be overcast today with southeasterly winds of level 3 to 4.
[0167] In response to the request, the system extracts target prosodic features and target phoneme features from the target text requested by the client.
[0168] The extracted target prosodic features and target phoneme features are input into the model as described above to obtain a first predicted file size.
[0169] Then, the residual is added to the first predicted file size, and the generated target file size is the sum of the first predicted file size and the target residual value.
[0170] The generated target file size is used as the predicted value of the wav file size.
[0171] Write the predicted WAV file size into the WAV file header.
[0172] The target text requested by the client is then segmented to generate multiple sentence sequences, for example:
[0173] The first sentence sequence: Shanghai will be overcast today;
[0174] Second sentence sequence: Southeast wind level three to four.
[0175] The audio of the first sentence sequence "Shanghai will be cloudy today" is synthesized to generate the first sentence voice, write it into a wav file, and return it to the client.
[0176] Then, the audio after the first sentence sequence is synthesized in order and written into a wav file until all requests are synthesized. For example, the audio of "southeast wind level 3 to 4" is synthesized and written into a wav file, and the process ends.
[0177] For example, when the file size is expressed as duration, the sample text "can be controlled" can be converted into a rhythmic phoneme sequence: silk e2#0y i3#1k ong4#0zh i4#3sil eos, and the duration of the rhythm and phonemes (number of Mel-spectrogram frames) is predicted: 3 1 3 1 1 6 2 2 7 2 4 5 11 4 12, and the total duration of the phonemes is used as the sample file size.
[0178] In the subsequent model training process, the model can be set to 1 layer of 256-dimensional embedding layer, followed by 4 layers of 1-D convolutional neural network with 256 channels, followed by layer norm, followed by dropout, and then a fully connected layer with an output dimension of 1.
[0179] Then the phoneme duration sequence d is converted to the log domain, where d'=log(d+1);
[0180] The loss function may include the MSE loss of the phoneme duration sequence and the MAE loss of the average total duration of each phoneme.
[0181] Then the Adam optimizer is used to iteratively optimize the model.
[0182] In the process of calculating the target residual value, the total number of predicted Mel spectrum frames can be obtained based on the above model as the second prediction file size, and then the maximum residual value is calculated.
[0183] In actual application, the client initiates a request, such as obtaining the target text: Shanghai will be overcast today with southeasterly winds of level 3 to 4.
[0184] In response to the request, the system extracts target prosodic features and target phoneme features from the target text requested by the client.
[0185] The extracted target prosodic features and target phoneme features are input into the model as described above to obtain a first predicted file size.
[0186] Then, the residual is added to the first predicted file size, and the generated target file size is the sum of the first predicted file size and the target residual value.
[0187] It should be noted that in this embodiment, the number of Mel spectrum frames is calculated, and the audio duration (total number of Mel spectrum frames) is converted into the WAV file size according to the Mel spectrum frame shift and the sampling frequency of the WAV file 16000, the sampling bit number 16, and the number of channels 1:
[0188] The wav file size = ((number of mel spectrum frames x mel spectrum frame shift / 16000)*16000*16*1 / 8+44) bytes.
[0189] According to the method for obtaining the size of audio and video files provided in the embodiment of the present application, by extracting the rhythm and phoneme features of the target text, and predicting the size information of the target audio file synthesized by the target text based on the extracted target rhythm features and target phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction result are high.
[0190] The following is a description of an apparatus for obtaining the size of an audio or video file provided in an embodiment of the present application. The apparatus for obtaining the size of an audio or video file described below and the method for obtaining the size of an audio or video file described above can refer to each other.
[0191] like Figure 3 As shown, the device for obtaining the size of the audio and video file includes: a first processing module 310 , a second processing module 320 and a third processing module 330 .
[0192] A first processing module 310 is used to obtain a target text;
[0193] The second processing module 320 is used to extract the characteristics of the target text and generate target prosodic characteristics and target phoneme characteristics;
[0194] The third processing module 330 is used to obtain a target file size of a target audio file based on the target prosodic feature and the target phoneme feature, where the target audio file is generated by performing speech synthesis on the target text.
[0195] According to the device for obtaining the size of audio and video files provided in the embodiment of the present application, by extracting the rhythmic features and phoneme features of the target text, and predicting the size information of the target audio file synthesized by the target text based on the extracted target rhythmic features and target phoneme features, the size value of the target file can be predicted before the target audio file is generated, which has a certain degree of timeliness; and the accuracy and precision of the prediction result are high.
[0196] In some embodiments, the third processing module 330 is configured to:
[0197] Based on the target prosodic feature and the target phoneme feature, obtaining a first predicted file size of the target audio file;
[0198] The first predicted file size and the target residual value are summed to generate a target file size, where the target residual value is determined based on the sample file size and the size of the sample audio file corresponding to the predicted sample text, and the sample file size is the actual size of the sample audio file corresponding to the sample text.
[0199] In some embodiments, the target residual value is determined by the following steps:
[0200] Obtaining a sample text, a sample audio file corresponding to the sample text, and a sample file size corresponding to the sample audio file, wherein the sample audio file is generated by performing speech synthesis on the sample text;
[0201] Extract features from the sample text to generate sample prosodic features and sample phoneme features;
[0202] Based on the sample prosodic feature and the sample phoneme feature, obtaining a second predicted file size of the sample audio file;
[0203] The maximum absolute value of the difference between the second predicted file size and the sample file size is determined as the target residual value.
[0204] In some embodiments, the third processing module 330 is configured to:
[0205] Inputting the target prosodic features and the target phoneme features into the file size prediction model, obtaining a first predicted file size output by the file size prediction model; wherein,
[0206] The file size prediction model is obtained by training with sample prosodic features and sample phoneme features as samples and sample file sizes corresponding to the sample prosodic features and sample phoneme features as sample labels.
[0207] In some embodiments, the apparatus may further include:
[0208] A fourth processing module is used for segmenting the target text based on the target prosodic features and the phoneme features to generate a plurality of sentence sequences after generating the target file size of the target audio file;
[0209] Perform speech synthesis on the sentence sequence to generate sentence speech;
[0210] Output the sentence-by-sentence speech and target file size, and concatenate the sentence-by-sentence speech to generate the target audio file.
[0211] In some embodiments, the second processing module 320 is further configured to:
[0212] Converting the target text into a prosodic phoneme sequence, the prosodic phoneme sequence comprising a plurality of phonemes corresponding to the target text and a prosodic identifier located between adjacent phonemes;
[0213] Feature extraction is performed on the prosodic phoneme sequence to generate target prosodic features and phoneme features.
[0214] In some embodiments, the second processing module 320 is further configured to:
[0215] Convert the target text into a phoneme sequence;
[0216] Obtain syllables, prosodic words, prosodic phrases, intonation phrases and sentence-end information of phoneme sequences;
[0217] The phoneme sequence is marked based on at least two of syllables, prosodic words, prosodic phrases, intonation phrases and sentence end information to generate a prosodic phoneme sequence.
[0218] In some embodiments, the target prosodic features and phoneme features include: the length of the prosodic phoneme sequence, the number of Chinese pinyins in the prosodic phoneme sequence, the number of pause symbols in the prosodic phoneme sequence, the number of English phonemes in the prosodic phoneme sequence, the number of Chinese phonemes in the prosodic phoneme sequence, the number of Chinese initials in the prosodic phoneme sequence, the number of Chinese finals in the prosodic phoneme sequence, and at least one of each category of English phonemes in the prosodic phoneme sequence.
[0219] Figure 4An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the method for obtaining the size of the audio and video file, the method comprising: obtaining the target text; extracting the features of the target text, generating the target prosodic features and the target phoneme features; based on the generated target prosodic features and the target phoneme features, generating the target file size of the target audio file, the target audio file being generated by performing speech synthesis on the target text.
[0220] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0221] Furthermore, the present application also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for obtaining the audio and video file size provided by the above-mentioned method embodiments, and the method includes: obtaining a target text; extracting features of the target text to generate target prosodic features and target phoneme features; based on the generated target prosodic features and target phoneme features, generating a target file size of a target audio file, the target audio file being generated by performing speech synthesis on the target text.
[0222] On the other hand, an embodiment of the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the method for obtaining the size of an audio or video file provided in the above embodiments, the method comprising: obtaining a target text; extracting features of the target text to generate target prosodic features and target phoneme features; based on the generated target prosodic features and target phoneme features, generating a target file size of a target audio file, the target audio file being generated by performing speech synthesis on the target text.
[0223] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0224] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0225] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0226] The above implementation modes are only used to illustrate the present application, but not to limit the present application. Although the present application is described in detail with reference to the embodiments, a person skilled in the art should understand that various combinations, modifications or equivalent substitutions of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application, and should be included in the scope of the claims of the present application.
Claims
1. A method for obtaining the size of an audio or video file, characterized in that: include: Get the target text; Extracting features from the target text to generate target prosodic features and target phoneme features; Based on the target prosodic feature and the target phoneme feature, obtaining a target file size of a target audio file, wherein the target audio file is generated by performing speech synthesis on the target text; Wherein, obtaining a target file size of a target audio file based on the target prosodic feature and the target phoneme feature includes: Based on the target prosodic feature and the target phoneme feature, obtaining a first predicted file size of the target audio file; The first predicted file size and the target residual value are summed to generate the target file size, wherein the target residual value is determined based on the sample file size and the size of the sample audio file corresponding to the predicted sample text, and the sample file size is the actual size of the sample audio file corresponding to the sample text.
2. The method for obtaining the size of an audio or video file according to claim 1, characterized in that: The target residual value is determined by the following steps: Obtaining a sample text, a sample audio file corresponding to the sample text, and a sample file size corresponding to the sample audio file, wherein the sample audio file is generated by performing speech synthesis on the sample text; Extracting features from the sample text to generate sample prosodic features and sample phoneme features; Based on the sample prosodic feature and the sample phoneme feature, obtaining a second predicted file size of the sample audio file; The absolute value of the difference between the second predicted file size and the sample file size is determined as the target residual value.
3. The method for obtaining the size of an audio or video file according to claim 1, characterized in that: The step of obtaining a first predicted file size of the target audio file based on the target prosody feature and the target phoneme feature includes: Inputting the target prosodic feature and the target phoneme feature into a file size prediction model, and obtaining the first predicted file size output by the file size prediction model; wherein, The file size prediction model is obtained by training with sample prosodic features and sample phoneme features as samples and sample file sizes corresponding to the sample prosodic features and the sample phoneme features as sample labels.
4. The method for obtaining the size of an audio or video file according to claim 1, characterized in that: After obtaining the target file size of the target audio file, the method further includes: Segmenting the target text based on the target prosodic features and the target phoneme features to generate a plurality of sentence sequences; Performing speech synthesis on the sentence sequence to generate sentence speech; The sentence speech and the target file size are output, and the sentence speech is concatenated to generate the target audio file.
5. The method for obtaining the size of an audio or video file according to any one of claims 1 to 4, characterized in that: The step of extracting features from the target text to generate target prosodic features and target phoneme features includes: Converting the target text into a prosodic phoneme sequence, wherein the prosodic phoneme sequence includes a plurality of phonemes corresponding to the target text and prosodic identifiers located between adjacent phonemes; Feature extraction is performed on the prosodic phoneme sequence to generate the target prosodic feature and the target phoneme feature.
6. The method for obtaining the size of an audio or video file according to claim 5, characterized in that: The target prosodic features and the target phoneme features include: the length of the prosodic phoneme sequence, the number of Chinese pinyins in the prosodic phoneme sequence, the number of pause symbols in the prosodic phoneme sequence, the number of English phonemes in the prosodic phoneme sequence, the number of Chinese phonemes in the prosodic phoneme sequence, the number of Chinese initials in the prosodic phoneme sequence, the number of Chinese finals in the prosodic phoneme sequence, and at least one of the various categories of English phonemes in the prosodic phoneme sequence.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for obtaining the size of the audio and video file as described in any one of claims 1 to 6 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for obtaining the size of an audio or video file as described in any one of claims 1 to 6 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for obtaining the size of an audio or video file as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Speech synthesis method and device, medium and electronic equipment
CN114242035A
Speech synthesis method, speech synthesis apparatus, smart terminal and storage medium
WO2021134591A1