A data processing method and device, computer equipment and storage medium
Patent Information
- Application Number
- CN202211134478.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-09-16
AI Technical Summary
[0002]目前的歌曲填词市场上已存在各类歌唱合成模型、系统和方案,但共存的问题是数据需求大,且基本都需要以曲谱信息为特征建模,这需要耗费大量的资源成本,需要专业性极强的懂得乐理的人进行人工标注,这种方式即使耗费大量的人力财力依然很难做到大规模的数据积累
[0016] This application embodiment can acquire text data to be processed and convert it into first text data; acquire a target musical score and determine second text data based on the target musical score and the first text data; determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In this way, text data can be added to the target musical score more efficiently and flexibly to generate target audio data while saving data costs, thus improving the accuracy and efficiency of data processing.
Smart Images

Figure CN117765898B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Technology
[0002] Currently, there are various singing synthesis models, systems, and solutions in the song lyrics market. However, the common problem is that they require a large amount of data and basically all require modeling based on musical score information. This requires a lot of resources and highly professional music theory experts to manually annotate the lyrics. Even with a lot of manpower and financial resources, it is still difficult to achieve large-scale data accumulation. Summary of the Invention
[0003] This application provides a data processing method, apparatus, computer equipment, and storage medium that can add text data to a target musical score more efficiently and flexibly to generate target audio data while saving data costs, thereby improving the accuracy and efficiency of data processing.
[0004] In a first aspect, embodiments of this application provide a data processing method, including:
[0005] Acquire the text data to be processed, and convert the text data to be processed into first text data;
[0006] Obtain the target musical score, and determine the second text data based on the target musical score and the first text data;
[0007] Determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme;
[0008] Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0009] Secondly, embodiments of this application provide a data processing apparatus, including:
[0010] The first acquisition unit is used to acquire text data to be processed and convert the text data to be processed into first text data.
[0011] The second acquisition unit is used to acquire the target musical score and determine the second text data based on the target musical score and the first text data;
[0012] The first determining unit is used to determine the phoneme of each keyword in the second text data and the time information of each phoneme, and to determine the volume information based on the time information of each phoneme.
[0013] The second determining unit is used to acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0014] Thirdly, embodiments of this application provide a computer device, the computer device including: a processor and a memory, the processor being configured to execute the method described in the first aspect above.
[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program instructions that, when executed, implement the method described in the first aspect above.
[0016] This application embodiment can acquire text data to be processed and convert it into first text data; acquire a target musical score and determine second text data based on the target musical score and the first text data; determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In this way, text data can be added to the target musical score more efficiently and flexibly to generate target audio data while saving data costs, thus improving the accuracy and efficiency of data processing. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0019] Figure 2 This is a flowchart illustrating another data processing method provided in an embodiment of this application;
[0020] Figure 3 This is a flowchart illustrating another data processing method provided in an embodiment of this application;
[0021] Figure 4 These are schematic diagrams illustrating two methods for obtaining baseband information;
[0022] Figure 5 This is a flowchart illustrating another data processing method provided in an embodiment of this application;
[0023] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] This application proposes a data processing scheme that converts the text data to be processed into first text data, determines the second text data of the target score based on the target score and the first text data, determines the phonemes and time information of each keyword in the second text data, determines the volume information based on the time information of each keyword's phonemes, and determines the target audio data based on the fundamental frequency information, phoneme time information, and volume information of the target score. This scheme requires a small amount of data and does not require score modeling, which greatly saves costs and can generate audio data more efficiently and accurately, thus improving the accuracy and efficiency of data processing.
[0027] This application provides a data processing method that can be applied to scenarios where lyrics are added to musical scores. In some embodiments, the data processing method can also be applied to other scenarios where text data is added to audio data.
[0028] The data processing method provided in this application embodiment can be applied to a data processing device. In some embodiments, the data processing device is applied to a computer device. In some embodiments, the computer device may include, but is not limited to, smart terminal devices such as smartphones, tablets, laptops, desktop computers, in-vehicle smart terminals, and smartwatches.
[0029] The data processing method provided in the embodiments of this application will be illustrated below with reference to the accompanying drawings.
[0030] Please see details. Figure 1 , Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The data processing method in this embodiment can be executed by a data processing device, which is located in a computer device, as explained above.
[0031] S101: Obtain the text data to be processed and convert it into the first text data.
[0032] In the embodiments of the present application, a computer device may acquire text data to be processed, and convert the text data to be processed into first text data. In some embodiments, the text data to be processed may include one or more of characters selected from characters (such as Chinese characters), letters, numbers, punctuation marks, etc., and the first text data may include text and / or punctuation marks of a specified type.
[0033] In one embodiment, when converting the text data to be processed into first text data, the computer device may acquire text content of an unspecified type from the text data to be processed; convert the text content of the unspecified type into text content of the specified type, and determine the specified-type text content obtained after conversion as the first text data. In some embodiments, the specified type may include, but is not limited to, one or more preset text types, for example, the specified type is Chinese characters.
[0034] For example, assuming that the specified type is Chinese characters and the text data to be processed is "Wo love ni", the computer device may acquire that the non-Chinese character text content in the text data to be processed is "love", then may convert the non-Chinese character text content "love" into the Chinese character text content "Ai", and determine the text content "Wo Ai ni" obtained after conversion as the first text data.
[0035] In one embodiment, when converting the unspecified-type text content into the specified-type text content, the computer device may acquire keywords of the unspecified-type text content, and acquire the association degree between each keyword and the text data to be processed; if the association degree is less than an association degree threshold, delete the keyword; if the association degree is greater than or equal to the association degree threshold, convert the keyword into specified-type text content.
[0036] In one embodiment, when acquiring the association degree between each keyword and the text data to be processed, the computer device may input each keyword and the text data to be processed into a pre-trained association prediction model, and obtain the association degree between each keyword and the text data to be processed through prediction. In some embodiments, the pre-trained association prediction model may be obtained through training by a neural network model.
[0037] For example, assuming that the specified type is Chinese characters, the text data to be processed is "123wo[K1] love ni[K2]", and the relevance threshold is 60%, then the keywords "1, 2, 3, wo[K1], love, ni[K2]" can be obtained. If it is predicted that the relevance between the keyword "1" and the text data to be processed is 20%, the relevance between the keyword "2" and the text data to be processed is 22%, the relevance between the keyword "3" and the text data to be processed is 21%, the relevance between the keyword "wo[K1]" and the text data to be processed is 70%, the relevance between the keyword "love" and the text data to be processed is 80%, and the relevance between the keyword "ni[K2]" and the text data to be processed is 75%, then the keywords "1, 2, 3" can be deleted, and the keywords "wo[K1] love ni[K2]" are converted into the Chinese character text content "Wo Ai Ni" (I love you).
[0038] In one embodiment, when converting text content of a non-specified type into text content of a specified type, the computer device may use a pre-trained text conversion model to convert the text content of the non-specified type into the text content of the specified type. Wherein, the pre-trained text conversion model may be obtained through training by a neural network model.
[0039] In one embodiment, when converting text data to be processed into first text data, the computer device may obtain text content of a specified type and text content of a non-specified type in the text data to be processed, delete the text content of the non-specified type, and determine the text content of the specified type as the first text data.
[0040] The present application converts the text data to be processed into first text data of a specified type, which facilitates subsequent more efficient and better addition of second text data to a target music score.
[0041] S102: Acquire a target music score, and determine second text data according to the target music score and the first text data.
[0042] In the embodiments of the present application, a computer device may acquire a target music score, and determine second text data according to the target music score and the first text data. In some embodiments, the target music score may be audio that does not include text. For example, a song includes lyrics and melody, and the target music score may be the melody of a song.
[0043] In one embodiment, when determining second text data according to the target music score and the first text data, the computer device may acquire note information of the target music score; determine a conversion strategy according to the note information and the first text data, and convert the first text data into the second text data according to the conversion strategy. In some embodiments, the note information includes but is not limited to note duration, number of notes, etc.
[0044] This application helps to more flexibly and effectively determine the second text data that matches the target musical score by using the target musical score and the first text data, thus preparing for the subsequent generation of target audio data.
[0045] S103: Determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme.
[0046] In this embodiment of the application, the computer device can determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme.
[0047] In one embodiment, when determining the phoneme of each keyword in the second text data and the time information of each phoneme, the computer device can determine the time information of each note in the target musical score based on the note information of the target musical score; and determine the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note.
[0048] S104: Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0049] In this embodiment of the application, the computer device can obtain the fundamental frequency information of the target musical score and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0050] This application embodiment can acquire text data to be processed and convert it into first text data; acquire a target musical score and determine second text data based on the target musical score and the first text data; determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In this way, text data can be added to the target musical score more efficiently and flexibly while saving data costs, improving the accuracy and efficiency of data processing.
[0051] Please see details. Figure 2 , Figure 2 This is a flowchart illustrating another data processing method provided in an embodiment of this application. The data processing method of this application embodiment can be executed by a data processing device, wherein the data processing device is disposed in a computer device, and the specific explanation of the computer device is as described above. This application embodiment describes how to determine second text data based on a target musical score and text data to be processed. Specifically, the method of this application embodiment includes the following steps.
[0052] S201: Obtain the text data to be processed and convert it into the first text data.
[0053] In this embodiment of the application, the computer device can acquire text data to be processed and convert the text data to be processed into first text data.
[0054] S202: Obtain the note information of the target musical score.
[0055] In this embodiment of the application, the computer device can acquire the note information of the target musical score.
[0056] S203: Determine the conversion strategy based on the note information and the first text data, and convert the first text data into the second text data according to the conversion strategy.
[0057] In this embodiment, the computer device can determine a conversion strategy based on the note information and the first text data, and convert the first text data into second text data according to the conversion strategy. In some embodiments, the conversion strategy can be one of a plurality of preset conversion strategies.
[0058] In one embodiment, the note information may include the number of notes. When the computer device determines a conversion strategy based on the note information and the first text data, and converts the first text data into second text data according to the conversion strategy, it may determine one or more text length thresholds based on the number of notes in the target score, compare the first text data with one or more text length thresholds based on the text length of the first text data, determine a conversion strategy based on the comparison result, and convert the first text data into second text data according to the conversion strategy.
[0059] In one embodiment, when a computer device compares first text data with one or more text length thresholds, if it detects that the text length of the first text data in the comparison result is greater than or equal to the first text length threshold, it can select text of the first text length threshold from the first text data and use the selected text of the first text length threshold as second text data.
[0060] For example, assuming the target musical score has 14 notes, the first text length threshold is 14, and each character occupies one note, then the output of 14 words after word replacement is the second text data. If the text length of the first text data is greater than or equal to 14 characters, then 14 characters can be selected from the first text data to determine the second text data.
[0061] In one embodiment, when a computer device compares first text data with one or more text length thresholds, if it detects that the text length of the first text data in the comparison result is less than or equal to a second text length threshold, it can determine the number of notes occupied by each character in the first text data based on the text length of the first text data, determine third text data based on the number of notes occupied by each character in the first text data, and if the number of notes in the third text data is less than the second text length threshold, it can add specified text to the third text data so that the text length (i.e., the number of notes) of the third text data after adding the specified text matches the number of notes in the target musical score, and determine the third text data after adding the specified text as the second text data, wherein the second text length threshold is less than the first text length threshold.
[0062] In one embodiment, when a computer device adds specified text to third text data, it can add the specified text at any position in the third text data, without any specific limitation.
[0063] For example, assuming the target musical score has 14 notes and the second text length threshold is 7, if the length of the first text data is less than or equal to 7 characters, then each character in the first text data occupies 2 notes. The third text data is determined based on the number of notes occupied by each character in the first text data (2). If the length of the notes in the third text data is less than 7, then specified text such as "la" or "ah" can be added to the third text data so that the length of the third text data after adding the specified text is 14, and the third text data after adding the specified text is determined as the second text data.
[0064] In one embodiment, when a computer device compares first text data with one or more text length thresholds, if it detects that the text length of the first text data in the comparison result is greater than a second text length threshold and less than the first text length threshold, it can add specified text to the first text data according to the number of notes in the target musical score, so that the text length of the first text data after adding the specified text matches the number of notes in the target musical score, and determine the first text data after adding the specified text as the second text data.
[0065] For example, assuming the target musical score has 14 notes and the second text length threshold is 7, if the text length of the first text data in the comparison result is detected to be greater than 7 and less than 14, then a specified text such as "la" or "ah" can be added to the first text data according to the number of notes of the target musical score of 14, so that the text length of the first text data after adding the specified text is 14, and the first text data after adding the specified text is determined to be the second text data.
[0066] In this embodiment, the first text data is converted into second text data based on the note information of the target score and the first text data, which helps to determine the second text data that is more compatible with the target score.
[0067] S204: Determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme.
[0068] In this embodiment of the application, the computer device can determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme.
[0069] S205: Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0070] In this embodiment of the application, the computer device can obtain the fundamental frequency information of the target musical score and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0071] This application embodiment converts the text data to be processed into first text data to obtain the note information of the target musical score. Based on the note information and the first text data, a conversion strategy is determined, and the first text data is converted into second text data according to the conversion strategy. This helps to determine the second text data that better matches the target musical score. By determining the phoneme of each keyword in the second text data and the time information of each phoneme, and determining the volume information based on the time information of each phoneme, the fundamental frequency information of the target musical score is obtained. Based on the fundamental frequency information, the time information of each phoneme, and the volume information, the target audio data is determined. This saves data costs and adds text data to the target musical score more efficiently and flexibly, improving the accuracy and efficiency of data processing.
[0072] Please see details. Figure 3 , Figure 3 This is a flowchart illustrating another data processing method provided in this application embodiment. The data processing method of this application embodiment can be executed by a data processing device, wherein the data processing device is disposed in a computer device, and the specific explanation of the computer device is as described above. This application embodiment describes how to determine target audio data based on target musical score and second text data. Specifically, the method of this application embodiment includes the following steps.
[0073] S301: Obtain the text data to be processed and convert it into the first text data.
[0074] In this embodiment of the application, the computer device can acquire text data to be processed and convert the text data to be processed into first text data.
[0075] S302: Obtain the target musical score and determine the second text data based on the target musical score and the first text data.
[0076] In this embodiment of the application, the computer device can acquire the target musical score and determine the second text data based on the target musical score and the first text data.
[0077] S303: Determine the time information of each note in the target score based on the note information of the target score.
[0078] In this embodiment of the application, the computer device can determine the time information of each note in the target musical score based on the note information of the target musical score.
[0079] S304: Based on the time information of each note, determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme.
[0080] In this embodiment, the computer device can determine the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note. In some embodiments, a syllable can be the basic unit of pronunciation, and the pronunciation of any word can be broken down into individual syllables. In some embodiments, a phoneme can be the smallest unit of speech distinguished from the perspective of sound quality, and can be divided into two categories based on pronunciation characteristics: vowel (also called vowel) phonemes and consonant (also called consonant) phonemes.
[0081] In one embodiment, when a computer device determines the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note, it can convert each keyword in the second text data into a syllable; determine the time information of each syllable based on the time information of each note and the syllables of each keyword in the second text data; split each syllable into phonemes, and determine the time information of each phoneme based on the time information of each syllable.
[0082] In some embodiments, syllables with single vowels do not need to be split; the syllable time is the same as the phoneme time. For syllables composed of an initial consonant and a final vowel, only the initial consonant time needs to be obtained to determine the remaining time of the final vowel. Obtaining the initial consonant time can include several methods: one is based on accumulated experience and language rules, dividing initial consonants into multiple categories, each corresponding to a different time; another is based on statistical analysis of the initial consonant time distribution of the target timbre, dividing initial consonants into multiple categories according to the time distribution, each corresponding to a different time. Simultaneously, it should be considered that when the total duration of a syllable is less than a duration threshold, the duration of certain initial consonants can be shortened.
[0083] In one embodiment, when determining volume information based on the time information of each phoneme, the computer device can use a pre-trained volume prediction model to predict the volume information at the current moment based on phoneme information, location information, and speaker information. In some embodiments, the location information may be the position of the current audio time frame (5ms) within the current phoneme. In some embodiments, the speaker information includes the target timbre. In some embodiments, the volume prediction model may be obtained by training a neural network model.
[0084] This application, by introducing volume information, helps to improve the consistency and stability of the output target audio data in terms of effect, volume, etc.
[0085] S305: Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0086] In this embodiment, the computer device can acquire the fundamental frequency information of the target musical score and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In some embodiments, the frequency of the fundamental tone is the fundamental frequency, which determines the pitch of the entire note.
[0087] In one embodiment, when a computer device acquires the fundamental frequency information of a target musical score, it can use the notes of the target musical score as a basis to obtain a more natural fundamental frequency curve by using a pre-trained anthropomorphic voice fundamental frequency generation algorithm. The anthropomorphic voice fundamental frequency generation algorithm can be obtained through neural network training, which will not be described in detail here.
[0088] In one embodiment, when acquiring the fundamental frequency information of the target musical score, the computer device can use the actual fundamental frequency of the target musical score as a basis. That is, the fundamental frequency is extracted from the actual vocals in the target musical score, manually aligned and repaired to ensure continuity, and then, based on the obtained time information and pronunciation rules of individual phonemes, the fundamental frequency curve corresponding to the time of phonemes without a fundamental frequency is zeroed out to obtain a new fundamental frequency curve. The specific method for acquiring the fundamental frequency is as follows: Figure 4 As shown, Figure 4 These are schematic diagrams illustrating two methods for obtaining baseband information.
[0089] In one embodiment, when a computer device determines target audio data based on fundamental frequency information, time information of each phoneme, and volume information, it can input the fundamental frequency information, time information of each phoneme, and volume information into a pre-trained audio synthesis model to obtain Mel-spectral features; convert the Mel-spectral features into audio data; and adjust the audio data to the target audio data based on audio adjustment data. In some embodiments, the audio adjustment data may include, but is not limited to, timbre, volume, speed adjustment, pitch shifting, equalizer adjustment, reverb, accompaniment, and audio-accompaniment volume ratio.
[0090] In one embodiment, when a computer device adjusts audio data to target audio data based on audio adjustment data, it can adjust different timbres to a fixed timbre. The timbre can be adjusted using an equalizer. A corresponding fixed accompaniment can also be added to the audio data. Furthermore, adjustments such as volume adjustment, speed adjustment, pitch adjustment, equalizer adjustment, adding reverb, adding accompaniment, and adjusting the volume ratio of the audio and accompaniment can be made according to the user's personal preferences.
[0091] In one embodiment, before inputting fundamental frequency information, time information of each phoneme, and volume information into a pre-trained audio synthesis model to obtain Mel-spectrum features, the computer device can train the audio synthesis model using sample data through a neural network model. Optionally, the sample data includes sample fundamental frequency data, sample phoneme information, sample volume information, and sample audio data. The sample data can be input into a preset neural network model for training to obtain predicted audio data. The predicted audio data is then compared with the sample audio data, and a loss function value is determined based on the comparison result. If the loss function value is greater than a threshold, the model parameters are adjusted according to the loss function value, and the sample data is input into the neural network model after the model parameters are adjusted for retraining. When the loss function value obtained from the retraining is less than the threshold, the audio synthesis model is determined to be obtained.
[0092] This application converts the text data to be processed into first text data, determines second text data based on the target musical score and the first text data, determines the time information of each note in the target musical score based on the note information, determines the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note, determines the volume information based on the time information of each phoneme, obtains the fundamental frequency information of the target musical score, and determines the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. The proposed solution requires less data and does not require musical score modeling, which greatly saves costs. It can more accurately generate the text data of the target musical score to synthesize the target audio data, improving the accuracy and efficiency of data processing.
[0093] Please see Figure 5 , Figure 5 This is a flowchart illustrating another data processing method provided in an embodiment of this application, as shown below. Figure 5As shown, the front-end module of the computer device processes the user-input text data into the first text data in the required format. Through a series of word substitution strategies, the syllable information of each character in the first text data in the given target musical score is determined. Then, the time information of each phoneme is determined through the phoneme timing module. The fundamental frequency information of the target musical score is generated using the fundamental frequency generation algorithm. The volume information of the target musical score is generated using the volume prediction model. The fundamental frequency information, the time information of each phoneme, and the volume information are then input into the trained audio synthesis model to infer and predict the Mel spectrum features. The audio data is then synthesized through a pre-trained vocoder. The sound effects post-processing module performs effects, equalization, reverb, accompaniment, etc. Finally, the target audio data, such as singing audio, is output through the user terminal of the computer device.
[0094] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Specifically, the device is disposed in a computer device, and the device includes: a first acquisition unit 601, a second acquisition unit 602, a first determination unit 603, and a second determination unit 604;
[0095] The first acquisition unit 601 is used to acquire text data to be processed and convert the text data to be processed into first text data.
[0096] The second acquisition unit 602 is used to acquire the target musical score and determine the second text data based on the target musical score and the first text data;
[0097] The first determining unit 603 is used to determine the phoneme of each keyword in the second text data and the time information of each phoneme, and to determine the volume information based on the time information of each phoneme.
[0098] The second determining unit 604 is used to acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0099] Furthermore, when the first acquisition unit 601 converts the text data to be processed into first text data, it is specifically used for:
[0100] Obtain the text content of non-specified type from the text data to be processed;
[0101] The text content of the non-specified type is converted into text content of the specified type, and the converted text content of the specified type is determined as the first text data.
[0102] Furthermore, when the first acquisition unit 601 converts the non-specified type text content into specified type text content, it is specifically used for:
[0103] Obtain the keywords of the text content of the non-specified type, and obtain the correlation degree between each keyword and the text data to be processed;
[0104] If the relevance is less than the relevance threshold, then the keyword is deleted;
[0105] If the relevance is greater than or equal to the relevance threshold, then the keyword is converted into text content of the specified type.
[0106] Furthermore, when the second acquisition unit 602 determines the second text data based on the target musical score and the first text data, it is specifically used for:
[0107] Obtain the note information of the target musical score;
[0108] A conversion strategy is determined based on the note information and the first text data, and the first text data is converted into the second text data according to the conversion strategy.
[0109] Furthermore, when the first determining unit 603 determines the phoneme of each keyword in the second text data and the time information of each phoneme, it is specifically used for:
[0110] The time information of each note in the target musical score is determined based on the note information of the target musical score;
[0111] Based on the time information of each note, the phoneme of each keyword in the second text data and the time information of each phoneme are determined.
[0112] Furthermore, when the first determining unit 603 determines the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note, it is specifically used for:
[0113] Convert each keyword in the second text data into a syllable;
[0114] Based on the time information of each note and the syllables of each keyword in the second text data, determine the time information of each syllable;
[0115] Each syllable is broken down into phonemes, and the time information of each phoneme is determined based on the time information of each syllable.
[0116] Furthermore, when the second determining unit 604 determines the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information, it is specifically used for:
[0117] The fundamental frequency information, the time information and volume information of each phoneme are input into a pre-trained audio synthesis model to obtain Mel spectrum features;
[0118] The Mel spectrum features are converted into audio data, and the audio data is adjusted to the target audio data based on the audio adjustment data.
[0119] This application embodiment can acquire text data to be processed and convert it into first text data; acquire a target musical score and determine second text data based on the target musical score and the first text data; determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In this way, text data can be added to the target musical score more efficiently and flexibly to generate target audio data while saving data costs, thus improving the accuracy and efficiency of data processing.
[0120] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Specifically, the computer device includes: a memory 701 and a processor 702.
[0121] In one embodiment, the computer device further includes a data interface 703 for transmitting data information between the computer device and other devices.
[0122] The memory 701 may include volatile memory; the memory 701 may also include non-volatile memory; the memory 701 may also include a combination of the above types of memory. The processor 702 may be a central processing unit (CPU). The processor 702 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), or any combination thereof.
[0123] The memory 701 is used to store programs, and the processor 702 can call the programs stored in the memory 701 to perform the following steps:
[0124] Acquire the text data to be processed, and convert the text data to be processed into first text data;
[0125] Obtain the target musical score, and determine the second text data based on the target musical score and the first text data;
[0126] Determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme;
[0127] Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
[0128] Furthermore, when the processor 702 converts the text data to be processed into the first text data, it is specifically used for:
[0129] Obtain the text content of non-specified type from the text data to be processed;
[0130] The text content of the non-specified type is converted into text content of the specified type, and the converted text content of the specified type is determined as the first text data.
[0131] Furthermore, when the processor 702 converts the non-specified type text content into specified type text content, it specifically performs the following:
[0132] Obtain the keywords of the text content of the non-specified type, and obtain the correlation degree between each keyword and the text data to be processed;
[0133] If the relevance is less than the relevance threshold, then the keyword is deleted;
[0134] If the relevance is greater than or equal to the relevance threshold, then the keyword is converted into text content of the specified type.
[0135] Furthermore, when the processor 702 determines the second text data based on the target musical score and the first text data, it is specifically used for:
[0136] Obtain the note information of the target musical score;
[0137] A conversion strategy is determined based on the note information and the first text data, and the first text data is converted into the second text data according to the conversion strategy.
[0138] Furthermore, when the processor 702 determines the phoneme of each keyword in the second text data and the time information of each phoneme, it is specifically used for:
[0139] The time information of each note in the target musical score is determined based on the note information of the target musical score;
[0140] Based on the time information of each note, the phoneme of each keyword in the second text data and the time information of each phoneme are determined.
[0141] Furthermore, when the processor 702 determines the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note, it is specifically used for:
[0142] Convert each keyword in the second text data into a syllable;
[0143] Based on the time information of each note and the syllables of each keyword in the second text data, determine the time information of each syllable;
[0144] Each syllable is broken down into phonemes, and the time information of each phoneme is determined based on the time information of each syllable.
[0145] Furthermore, when the processor 702 determines the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information, it is specifically used for:
[0146] The fundamental frequency information, the time information and volume information of each phoneme are input into a pre-trained audio synthesis model to obtain Mel spectrum features;
[0147] The Mel spectrum features are converted into audio data, and the audio data is adjusted to the target audio data based on the audio adjustment data.
[0148] This application embodiment can acquire text data to be processed and convert it into first text data; acquire a target musical score and determine second text data based on the target musical score and the first text data; determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information. In this way, text data can be added to the target musical score more efficiently and flexibly to generate target audio data while saving data costs, thus improving the accuracy and efficiency of data processing.
[0149] Embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements this application. Figure 1 or Figure 2 or Figure 3 or Figure 5 The method described in the corresponding embodiment can also be implemented. Figure 6 The apparatus described in the embodiments corresponding to this application will not be repeated here.
[0150] The computer-readable storage medium can be an internal storage unit of the device described in any of the foregoing embodiments, such as the device's hard drive or memory. The computer-readable storage medium can also be an external storage device of the device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the device. Further, the computer-readable storage medium may include both internal and external storage units of the device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0151] Embodiments of this application also provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0152] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0153] The above-disclosed embodiments are merely some of the embodiments of this application, and should not be construed as limiting the scope of this application. Those skilled in the art can understand that implementing all or part of the above embodiments and making equivalent changes in accordance with the claims of this application still fall within the scope of this invention.
Claims
1. A data processing method, characterized by, include: Acquire the text data to be processed, and convert the text data to be processed into first text data; Obtain the target musical score and obtain the note information of the target musical score, the note information including the number of notes; The process involves determining a conversion strategy based on the musical note information and the text length of the first text data, and converting the first text data into second text data according to the conversion strategy. This includes: determining a first text length threshold and a second text length threshold based on the musical note information, wherein the first text length threshold is greater than the second text length threshold; comparing the first text data with the first text length threshold and the second text length threshold based on the text length of the first text data; selecting text of the first text length threshold as the second text data when the text length of the first text data is greater than or equal to the first text length threshold; and determining a conversion strategy based on the text length of the first text data when the text length of the first text data is less than or equal to the second text length threshold. The number of musical notes occupied by each character in the first text data is used to determine the third text data. If the number of musical notes in the third text data is less than the second text length threshold, then specified text is added to the third text data so that the text length of the third text data after adding the specified text matches the number of musical notes in the target score, and the third text data after adding the specified text is determined to be the second text data. If the text length of the first text data is greater than the second text length threshold but less than the first text length threshold, then specified text is added to the first text data according to the number of musical notes in the target score so that the text length of the third text data after adding the specified text matches the number of musical notes in the target score, and the first text data after adding the specified text is determined to be the second text data. Determine the phoneme of each keyword in the second text data and the time information of each phoneme, and determine the volume information based on the time information of each phoneme; Obtain the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
2. The method of claim 1, wherein, The step of converting the text data to be processed into first text data includes: Obtain the text content of non-specified type from the text data to be processed; The text content of the non-specified type is converted into text content of the specified type, and the converted text content of the specified type is determined as the first text data.
3. The method of claim 2, wherein, The step of converting the non-specified text content into specified text content includes: Obtain the keywords of the text content of the non-specified type, and obtain the correlation degree between each keyword and the text data to be processed; If the relevance is less than the relevance threshold, then the keyword is deleted; If the relevance is greater than or equal to the relevance threshold, then the keyword is converted into text content of the specified type.
4. The method of claim 2, wherein, The step of determining the phoneme of each keyword in the second text data and the time information of each phoneme includes: The time information of each note in the target musical score is determined based on the note information of the target musical score; Based on the time information of each note, the phoneme of each keyword in the second text data and the time information of each phoneme are determined.
5. The method of claim 4, wherein, The step of determining the phoneme of each keyword in the second text data and the time information of each phoneme based on the time information of each note includes: Convert each keyword in the second text data into a syllable; Based on the time information of each note and the syllables of each keyword in the second text data, determine the time information of each syllable; Each syllable is broken down into phonemes, and the time information of each phoneme is determined based on the time information of each syllable.
6. The method of claim 1, wherein, The step of determining the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information includes: The fundamental frequency information, the time information and volume information of each phoneme are input into a pre-trained audio synthesis model to obtain Mel spectrum features; The Mel spectrum features are converted into audio data, and the audio data is adjusted to target audio data according to the audio adjustment data, which includes: timbre, volume, speed, pitch, equalizer adjustment, reverb, accompaniment, and audio-accompaniment volume ratio.
7. A data processing apparatus, characterized by, include: The first acquisition unit is used to acquire text data to be processed and convert the text data to be processed into first text data. The second acquisition unit is used to acquire a target musical score and acquire the note information of the target musical score, the note information including the number of notes; determine a conversion strategy based on the note information and the text length of the first text data, and convert the first text data into second text data according to the conversion strategy, including: determining a first text length threshold and a second text length threshold based on the note information, the first text length threshold being greater than the second text length threshold; comparing the first text data with the first text length threshold and the second text length threshold based on the text length of the first text data; when the text length of the first text data is greater than or equal to the first text length threshold, selecting text of the first text length threshold as the second text data; when the text length of the first text data is less than or equal to the second text length threshold, selecting text of the first text data with .... When the length threshold is reached, the number of musical notes occupied by each character in the first text data is determined based on the text length of the first text data. The third text data is then determined based on the number of musical notes occupied by each character in the first text data. If the number of musical notes in the third text data is less than the second text length threshold, a specified character is added to the third text data so that the text length of the third text data after adding the specified character matches the number of musical notes in the target score, and the third text data after adding the specified character is determined as the second text data. If the text length of the first text data is greater than the second text length threshold but less than the first text length threshold, a specified character is added to the first text data based on the number of musical notes in the target score so that the text length of the third text data after adding the specified character matches the number of musical notes in the target score, and the first text data after adding the specified character is determined as the second text data. The first determining unit is used to determine the phoneme of each keyword in the second text data and the time information of each phoneme, and to determine the volume information based on the time information of each phoneme. The second determining unit is used to acquire the fundamental frequency information of the target musical score, and determine the target audio data based on the fundamental frequency information, the time information of each phoneme, and the volume information.
8. A computer device, comprising: The device includes a processor and a memory interconnected thereto, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed, implement the method as described in any one of claims 1-6.
10. A computer program product, characterised in that, The computer program product includes computer instructions that, when executed, implement the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Song synthesizing method, device and equipment and storage medium
CN108877766A
Audio detection method, device, electronic equipment and readable storage medium
CN111312231A
Audio synthesis method and device, computer readable storage medium and electronic equipment
CN113838443A
Text data quality determination method and device
CN115048907A
Singing synthesis method and device, electronic equipment and storage medium
CN117116234A